Normalized Telegram record

Role confusion helps explain prompt injection

Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.

Signal contract

  • Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.
  • Tool output cannot be treated as inert data simply because the transport labels it as a tool message; agents still need isolation and policy enforcement outside the model.
  • Novelty: structural. Verification: source-inspected.

Source ledger

Publishable sources attached to this record.

2 public sources
#SourceRolePublic status
1github.comrepoprimary receiptsource_urls
2role-confusion.github.ioreposupporting receiptsource_urls

Observation

Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.

Why it matters

Tool output cannot be treated as inert data simply because the transport labels it as a tool message; agents still need isolation and policy enforcement outside the model.

Entities

Role Confusion

Provenance

This public record is an English normalization of QWG AI Telegram message 2560. The complete original-language post remains the canonical raw message.

Open the original Telegram record