---
schema_version: "newruntime-topic-hub-v0.1"
type: "topic_hub"
slug: "agent-security"
title: "Agent Security - New Runtime"
description: "A New Runtime topic hub collecting signals, patterns, field notes, and public sources about agent security."
answer: ["Agent Security is tracked here as an evidence-linked topic, not as a static glossary entry.","The page connects raw observations to pattern hypotheses, longer analysis, and public sources.","Use it as the canonical landing page before drilling into individual records."]
search_intents: ["agent security","agent security AI agents","agent security software"]
status: "tracked"
last_updated: "2026-07-22"
counts: {"total":6,"signals":3,"patterns":0,"posts":3,"atlas":0,"sources":8}
routes: {"html":"https://newruntime.com/topics/agent-security/","markdown":"https://newruntime.com/topics/agent-security.md","json":"https://newruntime.com/topics/agent-security.json"}
top_sources: ["https://arxiv.org/abs/2607.06595","https://github.com/nvidia/skillspector","https://github.com/role-confusion/prompt-injection-as-role-confusion","https://huggingface.co/blog/security-incident-july-2026","https://openai.com/index/hugging-face-model-evaluation-security-incident/","https://role-confusion.github.io/","https://www.pillar.security/blog/the-week-of-sandbox-escapes","https://x.com/mitchellh/status/2067970516951150721"]
---

# Agent Security - New Runtime

A New Runtime topic hub collecting signals, patterns, field notes, and public sources about agent security.

## Short Answer

- Agent Security is tracked here as an evidence-linked topic, not as a static glossary entry.
- The page connects raw observations to pattern hypotheses, longer analysis, and public sources.
- Use it as the canonical landing page before drilling into individual records.

## Patterns


## Field Notes

- [OpenAI's Hugging Face Incident Makes Agent Sandboxes a Production Risk](https://newruntime.com/posts/openai-hugging-face-security-incident/): OpenAI's model-evaluation incident with Hugging Face shows that cyber-capable agents need containment, monitoring, and evaluation controls that survive long-horizon behavior.
- [Coding Agent Sandboxes Break in Places Teams Do Not Expect](https://newruntime.com/posts/pillar-sandbox-escapes/): Pillar shows that agent sandboxes must be assessed not only around the agent process, but around files, configs, allowlisted commands, and local daemons the host later trusts.
- [GhostWriter: One Email Can Poison Long-Term Agent Memory](https://newruntime.com/posts/ghostwriter-agent-memory-poisoning/): GhostWriter shows a new risk class for agent systems: malicious content can enter long-term memory and later activate as trusted context.

## Recent Raw Signals

- 2026-06-28: [AGENTS.md Can Become an Instruction-Injection Surface](https://newruntime.com/signals/agents-md-can-become-an-instruction-injection-surface/)
- 2026-06-28: [Role confusion helps explain prompt injection](https://newruntime.com/signals/role-confusion-explains-prompt-injection/)
- 2026-06-02: [SkillSpector Adds a Security Gate for Agent Skills](https://newruntime.com/signals/skillspector-adds-a-security-gate-for-agent-skills/)

## Public Sources

- https://arxiv.org/abs/2607.06595
- https://github.com/nvidia/skillspector
- https://github.com/role-confusion/prompt-injection-as-role-confusion
- https://huggingface.co/blog/security-incident-july-2026
- https://openai.com/index/hugging-face-model-evaluation-security-incident/
- https://role-confusion.github.io/
- https://www.pillar.security/blog/the-week-of-sandbox-escapes
- https://x.com/mitchellh/status/2067970516951150721
