Normalized Telegram record
Role confusion helps explain prompt injection
Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.
Signal contract
- Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.
- Tool output cannot be treated as inert data simply because the transport labels it as a tool message; agents still need isolation and policy enforcement outside the model.
- Novelty: structural. Verification: source-inspected.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | github.comrepo | primary receipt | source_urls |
| 2 | role-confusion.github.iorepo | supporting receipt | source_urls |
Observation
Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.
Why it matters
Tool output cannot be treated as inert data simply because the transport labels it as a tool message; agents still need isolation and policy enforcement outside the model.
Entities
Role Confusion
Provenance
This public record is an English normalization of QWG AI Telegram message 2560. The complete original-language post remains the canonical raw message.