Topic hub
Agent Security
A New Runtime topic hub collecting signals, patterns, field notes, and public sources about agent security.
Pattern memory
What patterns are emerging?
- high
Agent security moves to runtime boundaries
The durable security boundary for an agent is the runtime that constrains what it can read, remember, execute, and authorize, not the prompt that asks it to behave.
Field notes
What should readers understand next?
Claude Code Auto Mode Gates Actions Instead Of Explanations
Claude Code Auto Mode combines an input injection probe with a two-stage action classifier, preserving autonomy while exposing an honest residual miss rate.
Containment Caps An Agent's Blast Radius
Anthropic's three runtime patterns show why hard filesystem, network, credential, and trust boundaries carry more security weight than repeated approval prompts.
AgentForger Turns A ChatGPT Link Into An Agent Builder Attack
Zenity Labs shows how a crafted ChatGPT Workspace Agents URL could preload instructions, attach already-authorized connectors, disable approvals, and schedule a persistent agent.
OpenAI's Hugging Face Incident Makes Agent Sandboxes a Production Risk
OpenAI's model-evaluation incident with Hugging Face shows that cyber-capable agents need containment, monitoring, and evaluation controls that survive long-horizon behavior.
Coding Agent Sandboxes Break in Places Teams Do Not Expect
Pillar shows that agent sandboxes must be assessed not only around the agent process, but around files, configs, allowlisted commands, and local daemons the host later trusts.
GhostWriter: One Email Can Poison Long-Term Agent Memory
GhostWriter shows a new risk class for agent systems: malicious content can enter long-term memory and later activate as trusted context.
Raw signals
What changed recently?
AGENTS.md Can Become an Instruction-Injection Surface
Repository instructions can manipulate low-effort automated pull requests, showing that agent context files are both useful capability layers and trust boundaries.
Role confusion helps explain prompt injection
Activation probes suggest instruction-like style can override architectural role labels when models interpret user, tool, and assistant text.
SkillSpector Adds a Security Gate for Agent Skills
NVIDIA's scanner treats installable agent instructions as executable supply-chain artifacts that require inspection before use.
GitHub / TencentCloud/CubeSandbox: Agent Security Runtime Boundaries
The archive captures GitHub / TencentCloud/CubeSandbox as a dated public record from GitHub / TencentCloud/CubeSandbox. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Claude: Agent Security Runtime Boundaries
The archive captures YouTube source as a dated public record from YouTube source. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Sandboxed Code Migration Agents (OpenAI cookbook)
The archive captures Sandboxed Code Migration Agents (OpenAI cookbook) as a dated public record from OpenAI Developers / Sandboxed Code Migration Agent. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Claude Mythos + Project Glasswing
The archive captures Claude Mythos + Project Glasswing as a dated public record from Anthropic. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
GitHub / SafeAI-Lab-X/ClawKeeper: Agent Security Runtime Boundaries
The archive captures GitHub / SafeAI-Lab-X/ClawKeeper as a dated public record from GitHub / SafeAI-Lab-X/ClawKeeper. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Anthropic / Claude Code Auto Mode: Agent Security Runtime Boundaries
The archive captures Anthropic / Claude Code Auto Mode as a dated public record from Anthropic / Claude Code Auto Mode. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
OpenAI / Lockdown Mode Elevated Risk Labels In Chatgpt: Agent Security Runtime Boundaries
The archive captures OpenAI / Lockdown Mode Elevated Risk Labels In Chatgpt as a dated public record from OpenAI / Lockdown Mode Elevated Risk Labels In Chatgpt. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Aitmpl / Dangerous Command Blocker: Agent Security Runtime Boundaries
The archive captures Aitmpl / Dangerous Command Blocker as a dated public record from Aitmpl / Dangerous Command Blocker. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Blog / Moltworker Self Hosted Ai Agent: Agent Security Runtime Boundaries
The archive captures Blog / Moltworker Self Hosted Ai Agent as a dated public record from Blog / Moltworker Self Hosted Ai Agent. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
GitHub / 0x4m4/hexstrike-ai: Agent Security Runtime Boundaries
The archive captures GitHub / 0x4m4/hexstrike-ai as a dated public record from GitHub / 0x4m4/hexstrike-ai. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
GitHub / cloudflare/moltworker: Agent Security Runtime Boundaries
The archive captures GitHub / cloudflare/moltworker as a dated public record from GitHub / cloudflare/moltworker. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Google Developers / Tailor Gemini Cli To Your Workflow With: Agent Security Runtime Boundaries
The archive captures Google Developers / Tailor Gemini Cli To Your Workflow With as a dated public record from Google Developers / Tailor Gemini Cli To Your Workflow With. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Google: Agent Security Runtime Boundaries
The archive captures X source as a dated public record from X source. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Docs: Agent Security Runtime Boundaries
The archive captures Docs as a dated public record from Docs. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Docs: Agent Security Runtime Boundaries
The archive captures Docs as a dated public record from Docs. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Blog / Securing Agents In Production Agentic Runtime 5191a0715240: Agent Security Runtime Boundaries
The archive captures Blog / Securing Agents In Production Agentic Runtime 5191a0715240 as a dated public record from Blog / Securing Agents In Production Agentic Runtime 5191a0715240. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Blog / Supply Chain Risk Of Agentic Ai Infecting: Agent Security Runtime Boundaries
The archive captures Blog / Supply Chain Risk Of Agentic Ai Infecting as a dated public record from Blog / Supply Chain Risk Of Agentic Ai Infecting. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Manus / Manus Sandbox: Agent Security Runtime Boundaries
The archive captures Manus / Manus Sandbox as a dated public record from Manus / Manus Sandbox. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Workos / Enterprise Ai Agent Playbook What Anthropic Openai: Agent Security Runtime Boundaries
The archive captures Workos / Enterprise Ai Agent Playbook What Anthropic Openai as a dated public record from Workos / Enterprise Ai Agent Playbook What Anthropic Openai. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
OpenAI: Agent Security Runtime Boundaries
The archive captures OpenAI as a dated public record from OpenAI. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Microsoft / How Were Tackling Microsoft Copilot Governance Internally: Agent Security Runtime Boundaries
The archive captures Microsoft / How Were Tackling Microsoft Copilot Governance Internally as a dated public record from Microsoft / How Were Tackling Microsoft Copilot Governance Internally. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | aitmpl.comsource | primary receipt | source_urls |
| 2 | anthropic.comsource | supporting receipt | source_urls |
| 3 | anthropic.comsource | supporting receipt | source_urls |
| 4 | arxiv.orgpaper | supporting receipt | source_urls |
| 5 | blog.cloudflare.comarticle | supporting receipt | source_urls |
| 6 | blog.lukaszolejnik.comarticle | supporting receipt | source_urls |
| 7 | blog.palantir.comarticle | supporting receipt | source_urls |
| 8 | developers.googleblog.comdocs | supporting receipt | source_urls |
| 9 | developers.openai.comdocs | supporting receipt | source_urls |
| 10 | docs.clawd.botdocs | supporting receipt | source_urls |
Showing 10 of 226; the complete set is exposed in the JSON route.