---
schema_version: "newruntime-agent-readable-v0.1"
type: "raw_signal"
id: "tg-2560"
slug: "role-confusion-explains-prompt-injection"
title: "Role confusion helps explain prompt injection"
description: "Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text."
observed_at: "2026-06-28"
why_it_matters: "Tool output cannot be treated as inert data simply because the transport labels it as a tool message; agents still need isolation and policy enforcement outside the model."
novelty: "structural"
verification_level: "source-inspected"
signal_type: "field-report"
evidence_kind: "mixed"
status: "published"
telegram_message_id: 2560
telegram_url: "https://t.me/qwgai/2560"
topics: ["prompt-injection","agent-security","model-behavior"]
entities: ["Role Confusion"]
related_patterns: []
source_urls: ["https://github.com/role-confusion/prompt-injection-as-role-confusion","https://role-confusion.github.io/"]
import_batch: "telegram-2026-07-17-v1"
routes: {"html":"https://newruntime.com/signals/role-confusion-explains-prompt-injection/","markdown":"https://newruntime.com/signals/role-confusion-explains-prompt-injection.md","json":"https://newruntime.com/signals/role-confusion-explains-prompt-injection.json"}
source_format: "telegram-export-normalized-json"
---

# Role confusion helps explain prompt injection

## Observation

Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.

## Why it matters

Tool output cannot be treated as inert data simply because the transport labels it as a tool message; agents still need isolation and policy enforcement outside the model.

## Provenance

Normalized from QWG AI Telegram message 2560. The original Russian-language record remains available at https://t.me/qwgai/2560.
