<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://newruntime.com/changes.xml</id>
  <title>New Runtime substantive changes</title>
  <subtitle>Editorially significant New Runtime pages. Build noise, crawler activity, and private discovery events are excluded.</subtitle>
  <updated>2026-09-01T00:00:00.000Z</updated>
  <link rel="self" type="application/atom+xml" href="https://newruntime.com/changes.xml"/>
  <link rel="hub" href="https://pubsubhubbub.appspot.com/"/>
  <link rel="alternate" type="text/html" href="https://newruntime.com/posts/"/>
  <author><name>New Runtime</name><uri>https://newruntime.com/</uri></author>
  <rights>New Runtime tracks the systems, signals, and shifts reshaping software as agents become first-class participants in work, products, and organizations.</rights>
  <entry>
    <id>https://newruntime.com/posts/model-access-portability-is-an-operational-resilience-layer/</id>
    <title>Model Access Portability Is an Operational Resilience Layer</title>
    <link rel="alternate" href="https://newruntime.com/posts/model-access-portability-is-an-operational-resilience-layer/"/>
    <published>2026-08-31T00:00:00.000Z</published>
    <updated>2026-09-01T00:00:00.000Z</updated>
    <summary>The Cursor transition turns provider portability from a preference into an operational continuity problem.</summary>
    <category term="coding-agents"/>
    <category term="model-portability"/>
    <category term="agent-harness"/>
    <category term="provider-risk"/>
    <category term="operational-resilience"/>
    <link rel="related" href="https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/"/>
    <link rel="related" href="https://blog.kilo.ai/p/your-coding-tool-should-not-choose"/>
    <link rel="related" href="https://openclaw.ai/blog/openclaw-2-accidentally"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/creative-ai-is-splitting-generation-into-production-primitives/</id>
    <title>Creative AI Is Splitting Generation Into Production Primitives</title>
    <link rel="alternate" href="https://newruntime.com/posts/creative-ai-is-splitting-generation-into-production-primitives/"/>
    <published>2026-08-31T00:00:00.000Z</published>
    <updated>2026-09-01T00:00:00.000Z</updated>
    <summary>Recent video and music releases expose extension, interpolation, previews, upscaling, structured song control, and sparse scaling as separate building blocks.</summary>
    <category term="generative-media"/>
    <category term="video-generation"/>
    <category term="music-generation"/>
    <category term="production-workflows"/>
    <link rel="related" href="https://bfl.ai/blog/flux-video-upscale"/>
    <link rel="related" href="https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/"/>
    <link rel="related" href="https://huggingface.co/MiniMaxAI/MiniMax-Music3"/>
    <link rel="related" href="https://sand.ai/blog/magi-2-preview"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/chatgpt-work-makes-execution-location-a-product-boundary/</id>
    <title>ChatGPT Work Makes Execution Location a Product Boundary</title>
    <link rel="alternate" href="https://newruntime.com/posts/chatgpt-work-makes-execution-location-a-product-boundary/"/>
    <published>2026-08-31T00:00:00.000Z</published>
    <updated>2026-09-01T00:00:00.000Z</updated>
    <summary>Work separates reviewable task execution from chat and makes local versus cloud placement an explicit workflow choice.</summary>
    <category term="chatgpt-work"/>
    <category term="post-app"/>
    <category term="local-execution"/>
    <category term="cloud-agents"/>
    <category term="artifact-workflows"/>
    <link rel="related" href="https://learn.chatgpt.com/docs/get-started-with-work"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/local-coding-agents-are-becoming-shared-workspaces/</id>
    <title>Local Coding Agents Are Becoming Shared Workspaces</title>
    <link rel="alternate" href="https://newruntime.com/posts/local-coding-agents-are-becoming-shared-workspaces/"/>
    <published>2026-08-31T00:00:00.000Z</published>
    <updated>2026-09-01T00:00:00.000Z</updated>
    <summary>Four projects show the coordination layer around coding agents becoming a distinct product surface.</summary>
    <category term="multi-agent"/>
    <category term="coding-agents"/>
    <category term="workspace"/>
    <category term="coordination"/>
    <category term="handoffs"/>
    <link rel="related" href="https://github.com/chaitanyagiri/munder-difflin"/>
    <link rel="related" href="https://github.com/Get-Concord-AI/concord-mcp"/>
    <link rel="related" href="https://github.com/browser-use/macOS-harness"/>
    <link rel="related" href="https://openclaw.ai/blog/openclaw-2-accidentally"/>
    <link rel="related" href="https://docs.openclaw.ai/releases/2026.8.1"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/langchain-agent-environments-spec-to-task-pipeline/</id>
    <title>LangChain Turns Agent Evals into a Spec-to-Task Production Pipeline</title>
    <link rel="alternate" href="https://newruntime.com/posts/langchain-agent-environments-spec-to-task-pipeline/"/>
    <published>2026-08-28T00:00:00.000Z</published>
    <updated>2026-08-28T00:00:00.000Z</updated>
    <summary>LangChain separates world knowledge, task specifications, environments, and graders so teams can continuously build representative agent evaluations.</summary>
    <category term="langchain"/>
    <category term="agent-evals"/>
    <category term="skills"/>
    <link rel="related" href="https://langchain.com/blog/building-agent-environments-and-tasks"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/figure-index-physical-data-infrastructure/</id>
    <title>Figure Index Turns Physical Data Collection into Robotics Infrastructure</title>
    <link rel="alternate" href="https://newruntime.com/posts/figure-index-physical-data-infrastructure/"/>
    <published>2026-08-28T00:00:00.000Z</published>
    <updated>2026-08-28T00:00:00.000Z</updated>
    <summary>Figure is building a global human-contributed video pipeline as a proprietary training-data layer for its Helix humanoid system.</summary>
    <category term="figure"/>
    <category term="robotics"/>
    <category term="training-data"/>
    <link rel="related" href="https://figure.ai/news/introducing-index"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/perplexity-portable-computer-local-cloud-boundary/</id>
    <title>Perplexity Portable Computer Makes Cloud Escalation a User-Controlled Boundary</title>
    <link rel="alternate" href="https://newruntime.com/posts/perplexity-portable-computer-local-cloud-boundary/"/>
    <published>2026-08-28T00:00:00.000Z</published>
    <updated>2026-08-28T00:00:00.000Z</updated>
    <summary>Portable Computer keeps orchestration, models, search state, and private work on device, escalating selected tasks to cloud services only with permission.</summary>
    <category term="perplexity"/>
    <category term="local-ai"/>
    <category term="agent-runtime"/>
    <link rel="related" href="https://x.com/AravSrinivas/status/2093004907343425777"/>
    <link rel="related" href="https://perplexity.ai/hub/blog/introducing-portable-computer-for-local-first-ai"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openclaw-viral-growth-maintainer-security-boundary/</id>
    <title>OpenClaw&apos;s Viral Growth Turns Maintainer Capacity into Security Infrastructure</title>
    <link rel="alternate" href="https://newruntime.com/posts/openclaw-viral-growth-maintainer-security-boundary/"/>
    <published>2026-08-28T00:00:00.000Z</published>
    <updated>2026-08-28T00:00:00.000Z</updated>
    <summary>GitHub&apos;s maintainer account shows how agent-generated contribution volume changes review, trust, and supply-chain work for a fast-growing project.</summary>
    <category term="openclaw"/>
    <category term="open-source"/>
    <category term="security"/>
    <link rel="related" href="https://github.blog/open-source/maintainers/openclaw-went-viral-meet-the-maintainers-building-and-securing-it"/>
    <link rel="related" href="https://x.com/github/status/2093008357644779643"/>
    <link rel="related" href="https://x.com/github/status/2093009115895243031"/>
    <link rel="related" href="https://x.com/steipete/status/2093013213038469205"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/deepmind-double-blind-frontier-model-evaluations/</id>
    <title>Double-Blind Evals Move Benchmark Integrity into the Runtime</title>
    <link rel="alternate" href="https://newruntime.com/posts/deepmind-double-blind-frontier-model-evaluations/"/>
    <published>2026-08-28T00:00:00.000Z</published>
    <updated>2026-08-28T00:00:00.000Z</updated>
    <summary>Google DeepMind&apos;s pilot uses confidential computing so evaluators keep prompts private while the model owner keeps proprietary weights private.</summary>
    <category term="deepmind"/>
    <category term="evals"/>
    <category term="confidential-computing"/>
    <link rel="related" href="https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations"/>
    <link rel="related" href="https://x.com/GoogleDeepMind/status/2092961763553677387"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/two-extractbenches-different-systems/</id>
    <title>Two ExtractBenches Measure Different Systems</title>
    <link rel="alternate" href="https://newruntime.com/posts/two-extractbenches-different-systems/"/>
    <published>2026-08-27T00:00:00.000Z</published>
    <updated>2026-08-27T00:00:00.000Z</updated>
    <summary>Two similarly named document benchmarks use different datasets and metrics, so their scores do not form one leaderboard.</summary>
    <category term="new-feature"/>
    <link rel="related" href="https://extractbench.ai/"/>
    <link rel="related" href="https://arxiv.org/abs/2607.29677"/>
    <link rel="related" href="https://arxiv.org/abs/2602.12247"/>
    <link rel="related" href="https://github.com/ContextualAI/extract-bench"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/hours-visible-autonomous-weeks-not/</id>
    <title>Hours Are Visible; Autonomous Weeks Are Not</title>
    <link rel="alternate" href="https://newruntime.com/posts/hours-visible-autonomous-weeks-not/"/>
    <published>2026-08-27T00:00:00.000Z</published>
    <updated>2026-08-27T00:00:00.000Z</updated>
    <summary>Current evidence supports longer bounded agent runs, but not reliable unattended work across weeks.</summary>
    <category term="new-feature"/>
    <link rel="related" href="https://arxiv.org/abs/2608.23283"/>
    <link rel="related" href="https://arxiv.org/abs/2608.19799"/>
    <link rel="related" href="https://deepmind.google/discover/blog/eve-towards-long-horizon-agents/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/minimum-integrity-stack-agent-leaderboard/</id>
    <title>The Minimum Integrity Stack for an Agent Leaderboard</title>
    <link rel="alternate" href="https://newruntime.com/posts/minimum-integrity-stack-agent-leaderboard/"/>
    <published>2026-08-27T00:00:00.000Z</published>
    <updated>2026-08-27T00:00:00.000Z</updated>
    <summary>A leaderboard needs isolation, path evidence, contamination checks, adversarial audits, and versioned corrections around every score.</summary>
    <category term="agents"/>
    <category term="evals"/>
    <category term="architecture"/>
    <link rel="related" href="https://artificialanalysis.ai/methodology/coding-agents-benchmarking"/>
    <link rel="related" href="https://dreadnode.io/blog/every-model-cheats"/>
    <link rel="related" href="https://arxiv.org/abs/2605.12673"/>
    <link rel="related" href="https://arxiv.org/abs/2604.11806"/>
    <link rel="related" href="https://arxiv.org/abs/2601.20103"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anti-cheat-prompt-is-not-security-control/</id>
    <title>An Anti-Cheat Prompt Is Not a Security Control</title>
    <link rel="alternate" href="https://newruntime.com/posts/anti-cheat-prompt-is-not-security-control/"/>
    <published>2026-08-27T00:00:00.000Z</published>
    <updated>2026-08-27T00:00:00.000Z</updated>
    <summary>Benchmark integrity requires runtime enforcement even when explicit instructions reduce some reward-hacking behavior.</summary>
    <category term="evals"/>
    <category term="new-feature"/>
    <link rel="related" href="https://dreadnode.io/blog/every-model-cheats"/>
    <link rel="related" href="https://arxiv.org/abs/2608.22103"/>
    <link rel="related" href="https://artificialanalysis.ai/methodology/coding-agents-benchmarking"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/skill-router-must-say-no/</id>
    <title>A Skill Router Must Be Able to Say No</title>
    <link rel="alternate" href="https://newruntime.com/posts/skill-router-must-say-no/"/>
    <published>2026-08-27T00:00:00.000Z</published>
    <updated>2026-08-27T00:00:00.000Z</updated>
    <summary>A practical routing contract for loading agent skills only when measured task evidence predicts a net benefit.</summary>
    <category term="agents"/>
    <category term="architecture"/>
    <link rel="related" href="https://arxiv.org/abs/2608.14036"/>
    <link rel="related" href="https://arxiv.org/abs/2608.23067"/>
    <link rel="related" href="https://arxiv.org/abs/2608.19880"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anthropic-model-hardware-standard-control-boundary/</id>
    <title>Anthropic&apos;s Model Hardware Standard Defines a Control Boundary for Hardware Agents</title>
    <link rel="alternate" href="https://newruntime.com/posts/anthropic-model-hardware-standard-control-boundary/"/>
    <published>2026-08-28T00:00:00.000Z</published>
    <updated>2026-08-28T00:00:00.000Z</updated>
    <summary>MHS proposes a model-agnostic interface for agents to operate programmable laboratory and manufacturing equipment under explicit safety constraints.</summary>
    <category term="anthropic"/>
    <category term="hardware-agents"/>
    <category term="standards"/>
    <link rel="related" href="https://anthropic.com/news/model-hardware-standard-research-preview"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2093038428757918070"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2093038429936529897"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2093038431190548803"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2093038432302080019"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2093038433782624261"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gemini-3-7-flash-rolls-out-to-gemini-pro-and-ultra-users/</id>
    <title>Gemini 3.7 Flash rolls out to Gemini Pro and Ultra users</title>
    <link rel="alternate" href="https://newruntime.com/posts/gemini-3-7-flash-rolls-out-to-gemini-pro-and-ultra-users/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Gemini 3.7 Flash rolls out to Gemini Pro and Ultra users is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/GeminiApp/status/2088326407730692538"/>
    <link rel="related" href="https://twitter.com/GeminiApp/status/2087948790296973683"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/developer-agent-utilities-editor-affordances-ssh-agents-claude-code-controls-and/</id>
    <title>Developer-agent utilities: editor affordances, SSH agents, Claude Code controls, and human-agent PM</title>
    <link rel="alternate" href="https://newruntime.com/posts/developer-agent-utilities-editor-affordances-ssh-agents-claude-code-controls-and/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Developer-agent utilities: editor affordances, SSH agents, Claude Code controls, and human-agent PM ties several source-backed updates into one New Runtime pattern for agent and AI infrastructure work.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/zeddotdev/status/2088368308294697121"/>
    <link rel="related" href="https://github.com/miantiao-me/ssh-ai-chat"/>
    <link rel="related" href="https://x.com/tom_doerr/status/2089295092108398927"/>
    <link rel="related" href="https://support.claude.com/en/articles/14552983-models-usage-and-limits-in-claude-code"/>
    <link rel="related" href="https://x.com/ClaudeDevs/status/2088014831605702937"/>
    <link rel="related" href="https://x.com/ClaudeDevs/status/2089471692762673408"/>
    <link rel="related" href="https://multica.ai"/>
    <link rel="related" href="https://raw.githubusercontent.com/miantiao-me/ssh-ai-chat/master/README.md"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ai-governance-is-becoming-a-product-and-infrastructure-boundary/</id>
    <title>AI governance is becoming a product and infrastructure boundary</title>
    <link rel="alternate" href="https://newruntime.com/posts/ai-governance-is-becoming-a-product-and-infrastructure-boundary/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>AI governance is becoming a product and infrastructure boundary ties several source-backed updates into one New Runtime pattern for agent and AI infrastructure work.</summary>
    <category term="ai-news"/>
    <category term="models"/>
    <category term="governance"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2088324824863236248"/>
    <link rel="related" href="https://www.anthropic.com/news/claude-text-watermark"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2088343978873966687"/>
    <link rel="related" href="https://x.com/EXM7777/status/2089295477334507750"/>
    <link rel="related" href="https://www.cnn.com/2026/08/18/business/google-spirit-airlines-data"/>
    <link rel="related" href="https://www.forbes.com/sites/suzannerowankelleher/2026/08/18/google-train-ai-spirit-airlines-data"/>
    <link rel="related" href="https://www.axios.com/2026/08/18/openai-pause-astra-preparedness-framework"/>
    <link rel="related" href="https://fortune.com/2026/08/18/openai-says-it-paused-ai-training-for-two-weeks-and-announces-new-security-protocols-following-hugging-face-hack"/>
    <link rel="related" href="https://stripe.com/newsroom/news/stripe-agrees-to-acquire-openrouter"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-capabilities-are-becoming-packaged-surfaces/</id>
    <title>Agent capabilities are becoming packaged surfaces</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-capabilities-are-becoming-packaged-surfaces/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Agent capabilities are becoming packaged surfaces ties several source-backed updates into one New Runtime pattern for agent and AI infrastructure work.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="governance"/>
    <link rel="related" href="https://www.remotion.dev/docs/ai/skills"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295048256991567"/>
    <link rel="related" href="https://www.remotion.dev/plugins"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295046289875125"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295044326961270"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295042485600381"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295040652652871"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295038932996194"/>
    <link rel="related" href="https://support.claude.com/en/articles/14552983-models-usage-and-limits-in-claude-code"/>
    <link rel="related" href="https://x.com/ClaudeDevs/status/2088014831605702937"/>
    <link rel="related" href="https://x.com/ClaudeDevs/status/2089471692762673408"/>
    <link rel="related" href="https://developers.googleblog.com/agent-plugins-package-your-skills-tools-and-more"/>
    <link rel="related" href="https://agent-plugins.org"/>
    <link rel="related" href="https://elevenlabs.io/blog/elevenlabs-mcp-in-claude"/>
    <link rel="related" href="https://hermes-agent.nousresearch.com/docs/user-guide/bot-mode"/>
    <link rel="related" href="https://www.youtube.com/watch?v=e1snsuY4lTI"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-work-is-moving-into-harnesses-factories-and-accepted-artifacts/</id>
    <title>Agent work is moving into harnesses, factories, and accepted artifacts</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-work-is-moving-into-harnesses-factories-and-accepted-artifacts/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Agent work is moving into harnesses, factories, and accepted artifacts ties several source-backed updates into one New Runtime pattern for agent and AI infrastructure work.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="models"/>
    <link rel="related" href="https://github.blog/engineering/turn-one-giant-ai-generated-pull-request-to-a-reviewable-stack"/>
    <link rel="related" href="https://x.com/github/status/2088315796749701550"/>
    <link rel="related" href="https://x.com/augmentcode/status/2088375653225955710"/>
    <link rel="related" href="https://www.remotion.dev/docs/ai/skills"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295048256991567"/>
    <link rel="related" href="https://www.remotion.dev/plugins"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295046289875125"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295044326961270"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295042485600381"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295040652652871"/>
    <link rel="related" href="https://x.com/Remotion/status/2089295038932996194"/>
    <link rel="related" href="https://github.com/deepseek-ai/deepseek-harness"/>
    <link rel="related" href="https://cursor.com/changelog/origin-code-hosting"/>
    <link rel="related" href="https://techcrunch.com/2026/08/18/cursor-capitalizes-on-github-frustration-launches-rival-hosting-platform"/>
    <link rel="related" href="https://venturebeat.com/infrastructure/cursor-launches-origin-code-hosting-platform-as-github-outage-exposes-opening-in-ai-coding-race"/>
    <link rel="related" href="https://www.warp.dev/blog/open-infrastructure-for-building-a-software-factory"/>
    <link rel="related" href="https://techcrunch.com/2026/08/18/warps-new-system-is-an-out-of-the-box-software-factory-for-ai-development"/>
    <link rel="related" href="https://github.com/iannuttall/clockwork"/>
    <link rel="related" href="https://multica.ai"/>
    <link rel="related" href="https://linear.app/data"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/model-routing-and-pricing-update-glm-deepseek-grok-qwen/</id>
    <title>Model routing and pricing update: GLM, DeepSeek, Grok, Qwen</title>
    <link rel="alternate" href="https://newruntime.com/posts/model-routing-and-pricing-update-glm-deepseek-grok-qwen/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Model routing and pricing update: GLM, DeepSeek, Grok, Qwen ties several source-backed updates into one New Runtime pattern for agent and AI infrastructure work.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/vercel_dev/status/2089123749572616645"/>
    <link rel="related" href="https://x.com/vercel_dev/status/2089123749572616645/photo/1"/>
    <link rel="related" href="https://x.com/opencode/status/2089428008268402786"/>
    <link rel="related" href="https://x.com/opencode/status/2089030951653240880"/>
    <link rel="related" href="https://x.ai/news/grok-4-6-github-copilot"/>
    <link rel="related" href="https://x.com/augmentcode/status/2088366283020759207"/>
    <link rel="related" href="https://www.augmentcode.com"/>
    <link rel="related" href="https://x.com/grok/status/2088319896597975501"/>
    <link rel="related" href="https://huggingface.co/Qwen/Qwen3.8-27B"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anthropic-publishes-regular-risk-reports-under-its-responsible-scaling-policy/</id>
    <title>Anthropic publishes regular Risk Reports under its Responsible Scaling Policy</title>
    <link rel="alternate" href="https://newruntime.com/posts/anthropic-publishes-regular-risk-reports-under-its-responsible-scaling-policy/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Anthropic publishes regular Risk Reports under its Responsible Scaling Policy is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="governance"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2088324824863236248"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/warp-introduces-open-infrastructure-for-software-factories/</id>
    <title>Warp introduces open infrastructure for software factories</title>
    <link rel="alternate" href="https://newruntime.com/posts/warp-introduces-open-infrastructure-for-software-factories/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Warp introduces open infrastructure for software factories is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <link rel="related" href="https://www.warp.dev/blog/open-infrastructure-for-building-a-software-factory"/>
    <link rel="related" href="https://techcrunch.com/2026/08/18/warps-new-system-is-an-out-of-the-box-software-factory-for-ai-development"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-reportedly-paused-training-and-added-security-protocols-after-the-hugging/</id>
    <title>OpenAI reportedly paused training and added security protocols after the Hugging Face incident</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-reportedly-paused-training-and-added-security-protocols-after-the-hugging/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>OpenAI reportedly paused training and added security protocols after the Hugging Face incident is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="models"/>
    <category term="governance"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://www.axios.com/2026/08/18/openai-pause-astra-preparedness-framework"/>
    <link rel="related" href="https://fortune.com/2026/08/18/openai-says-it-paused-ai-training-for-two-weeks-and-announces-new-security-protocols-following-hugging-face-hack"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-introduces-agent-plugins-for-packaging-skills-and-tools/</id>
    <title>Google introduces Agent Plugins for packaging skills and tools</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-introduces-agent-plugins-for-packaging-skills-and-tools/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Google introduces Agent Plugins for packaging skills and tools is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <link rel="related" href="https://developers.googleblog.com/agent-plugins-package-your-skills-tools-and-more"/>
    <link rel="related" href="https://agent-plugins.org"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/deepseek-publishes-deepseek-harness-on-github/</id>
    <title>DeepSeek publishes DeepSeek Harness on GitHub</title>
    <link rel="alternate" href="https://newruntime.com/posts/deepseek-publishes-deepseek-harness-on-github/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>DeepSeek publishes DeepSeek Harness on GitHub is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://github.com/deepseek-ai/deepseek-harness"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/augment-rebuilds-auggie-cli-on-pi-and-cuts-cost-per-task/</id>
    <title>Augment rebuilds Auggie CLI on Pi and cuts cost per task</title>
    <link rel="alternate" href="https://newruntime.com/posts/augment-rebuilds-auggie-cli-on-pi-and-cuts-cost-per-task/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Augment rebuilds Auggie CLI on Pi and cuts cost per task is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/augmentcode/status/2088375653225955710"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/github-copilot-code-review-balanced-depth-is-generally-available/</id>
    <title>GitHub Copilot code review Balanced depth is generally available</title>
    <link rel="alternate" href="https://newruntime.com/posts/github-copilot-code-review-balanced-depth-is-generally-available/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>GitHub Copilot code review Balanced depth is generally available is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/github/status/2089057545998479457"/>
    <link rel="related" href="https://x.com/github/status/2089057545998479457/video/1"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/github-shows-how-to-split-a-large-ai-generated-pr-into-reviewable-stacks/</id>
    <title>GitHub shows how to split a large AI-generated PR into reviewable stacks</title>
    <link rel="alternate" href="https://newruntime.com/posts/github-shows-how-to-split-a-large-ai-generated-pr-into-reviewable-stacks/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>GitHub shows how to split a large AI-generated PR into reviewable stacks is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="agents"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://github.blog/engineering/turn-one-giant-ai-generated-pull-request-to-a-reviewable-stack"/>
    <link rel="related" href="https://x.com/github/status/2088315796749701550"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anthropic-explains-claude-text-watermarking-for-eu-ai-act-compliance/</id>
    <title>Anthropic explains Claude text watermarking for EU AI Act compliance</title>
    <link rel="alternate" href="https://newruntime.com/posts/anthropic-explains-claude-text-watermarking-for-eu-ai-act-compliance/"/>
    <published>2026-08-20T00:00:00.000Z</published>
    <updated>2026-08-20T00:00:00.000Z</updated>
    <summary>Anthropic explains Claude text watermarking for EU AI Act compliance is a source-backed New Runtime newsroom item about AI product and infrastructure change.</summary>
    <category term="ai-news"/>
    <category term="governance"/>
    <link rel="related" href="https://www.anthropic.com/news/claude-text-watermark"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/2088343978873966687"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/a-three-person-team-shipping-hundreds-of-prs-moves-the-bottleneck-to-review/</id>
    <title>A Three-Person Team Shipping Hundreds of PRs Moves the Bottleneck to Review</title>
    <link rel="alternate" href="https://newruntime.com/posts/a-three-person-team-shipping-hundreds-of-prs-moves-the-bottleneck-to-review/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Wes McKinney documented a three-person agentic engineering workflow that produced hundreds of pull requests.</summary>
    <category term="coding-agents"/>
    <category term="case-study"/>
    <category term="workflow"/>
    <link rel="related" href="https://wesmckinney.com/blog/agentic-engineering-aug-2026"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/deepseek-v4-pro-distribution-evidence-arrives-before-a-strong-primary-model-record/</id>
    <title>DeepSeek V4-Pro Distribution Evidence Arrives Before a Strong Primary Model Record</title>
    <link rel="alternate" href="https://newruntime.com/posts/deepseek-v4-pro-distribution-evidence-arrives-before-a-strong-primary-model-record/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>DeepSeek V4-Pro-0813 appeared quickly in downstream coding-agent surfaces with aggressive pricing claims.</summary>
    <category term="new-models"/>
    <category term="coding-agents"/>
    <category term="platforms"/>
    <link rel="related" href="https://x.com/cline/status/2088008671603392972"/>
    <link rel="related" href="https://x.com/cline/status/2088008673197167040"/>
    <link rel="related" href="https://cline.bot/cline-pass"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/arc-agi-3-result-shows-the-harness-is-part-of-the-score/</id>
    <title>An ARC-AGI-3 Result Shows the Harness Is Part of the Score</title>
    <link rel="alternate" href="https://newruntime.com/posts/arc-agi-3-result-shows-the-harness-is-part-of-the-score/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Jeremy Berman&apos;s public harness reports 96.2% for Claude Opus 5 on 25 ARC-AGI-3 games, versus a 30.2% model-only result cited by the project.</summary>
    <category term="benchmarks"/>
    <category term="harness-engineering"/>
    <category term="evals"/>
    <link rel="related" href="https://github.com/jerber/arc-code"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/vercel-runs-ai-sdk-maintenance-as-a-human-gated-software-factory/</id>
    <title>Vercel Runs AI SDK Maintenance as a Human-Gated Software Factory</title>
    <link rel="alternate" href="https://newruntime.com/posts/vercel-runs-ai-sdk-maintenance-as-a-human-gated-software-factory/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Vercel described a software factory that processes AI SDK issues and pull requests through specialized, reviewable agent tasks while humans control every merge.</summary>
    <category term="coding-agents"/>
    <category term="harness-engineering"/>
    <category term="workflow"/>
    <link rel="related" href="https://vercel.com/blog/building-a-software-factory-for-ai-sdk"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/prompt-injection-enters-the-court-record/</id>
    <title>Prompt Injection Enters the Court Record</title>
    <link rel="alternate" href="https://newruntime.com/posts/prompt-injection-enters-the-court-record/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Reuters reported that a Connecticut judge said a plaintiff hid messages aimed at influencing AI systems inside court filings.</summary>
    <category term="security"/>
    <category term="prompts"/>
    <category term="governance"/>
    <link rel="related" href="https://www.reuters.com/legal/litigation/connecticut-judge-says-plaintiff-hid-messages-ai-court-filings-2026-08-13"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/twitch-makes-creator-training-data-use-opt-out-by-default/</id>
    <title>Twitch Makes Creator Training-Data Use Opt-Out by Default</title>
    <link rel="alternate" href="https://newruntime.com/posts/twitch-makes-creator-training-data-use-opt-out-by-default/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Twitch will allow creator content to be used for AI training by default unless streamers opt out, according to TechCrunch reporting.</summary>
    <category term="data-for-llm"/>
    <category term="governance"/>
    <category term="platforms"/>
    <link rel="related" href="https://techcrunch.com/2026/08/12/amazon-will-train-on-twitch-streamers-content-by-default-unless-they-opt-out"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/apple-publisher-talks-put-content-licensing-back-inside-assistant-design/</id>
    <title>Apple Publisher Talks Put Content Licensing Back Inside Assistant Design</title>
    <link rel="alternate" href="https://newruntime.com/posts/apple-publisher-talks-put-content-licensing-back-inside-assistant-design/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>The Wall Street Journal reported that Apple is discussing content licensing with publishers for an AI-powered Siri.</summary>
    <category term="governance"/>
    <category term="retrieval"/>
    <category term="platforms"/>
    <link rel="related" href="https://www.wsj.com/business/media/apple-in-talks-to-pay-publishers-to-improve-ai-powered-siri-0641f64b"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/aws-continuum-brings-security-context-into-the-developer-agent-loop/</id>
    <title>AWS Continuum Brings Security Context Into the Developer Agent Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/aws-continuum-brings-security-context-into-the-developer-agent-loop/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>AWS announced integrations with Anthropic and OpenAI that bring AWS Continuum security context into developer workflows.</summary>
    <category term="security"/>
    <category term="coding-agents"/>
    <category term="workflow"/>
    <link rel="related" href="https://aws.amazon.com/blogs/security/aws-partners-with-anthropic-and-openai-to-bring-aws-continuum-into-developer-workflows"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/claude-riemann-zeta-result-shows-a-human-model-research-loop/</id>
    <title>Claude&apos;s Riemann Zeta Result Shows a Human-Model Research Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/claude-riemann-zeta-result-shows-a-human-model-research-loop/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Anthropic published a mathematical result involving Claude and a lower bound for the Riemann zeta function.</summary>
    <category term="research"/>
    <category term="reasoning"/>
    <category term="anthropic"/>
    <link rel="related" href="https://www.anthropic.com/research/riemann-zeta"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas/</id>
    <title>The Bitter Lesson of Tool Calling Tests General Learning Against Hand-Built Schemas</title>
    <link rel="alternate" href="https://newruntime.com/posts/the-bitter-lesson-of-tool-calling-tests-general-learning-against-hand-built-schemas/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>A new paper asks whether general learning principles can outperform increasingly hand-engineered tool-calling schemes.</summary>
    <category term="tool-calling"/>
    <category term="research"/>
    <category term="evals"/>
    <link rel="related" href="https://arxiv.org/abs/2608.06370"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/databricks-reframes-coding-agent-cost-around-accepted-output/</id>
    <title>Databricks Reframes Coding-Agent Cost Around Accepted Output</title>
    <link rel="alternate" href="https://newruntime.com/posts/databricks-reframes-coding-agent-cost-around-accepted-output/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Databricks described controls for managing AI coding costs at scale, including limits, model routing, and workload governance.</summary>
    <category term="coding-agents"/>
    <category term="enterprise-ai"/>
    <category term="routing"/>
    <link rel="related" href="https://www.databricks.com/blog/managing-ai-coding-costs-scale"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/dynatrace-agrees-to-acquire-arize-ai/</id>
    <title>Dynatrace Agrees to Acquire Arize AI</title>
    <link rel="alternate" href="https://newruntime.com/posts/dynatrace-agrees-to-acquire-arize-ai/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Dynatrace entered an agreement to acquire Arize AI, bringing an AI-native evaluation and monitoring layer into a broad observability platform.</summary>
    <category term="observability"/>
    <category term="evals"/>
    <category term="enterprise-ai"/>
    <link rel="related" href="https://x.com/aparnadhinak/status/2087844420922364148"/>
    <link rel="related" href="https://x.com/swyx/status/2088049159509344265"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/perplexity-moves-sonar-workloads-to-a-multi-provider-agent-api/</id>
    <title>Perplexity Moves Sonar Workloads to a Multi-Provider Agent API</title>
    <link rel="alternate" href="https://newruntime.com/posts/perplexity-moves-sonar-workloads-to-a-multi-provider-agent-api/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Perplexity is steering Sonar workloads toward its Agent API, a unified multi-provider interface with web search, configurable tools, reasoning controls, presets, and model selection.</summary>
    <category term="search"/>
    <category term="agents"/>
    <category term="routing"/>
    <link rel="related" href="https://docs.perplexity.ai/docs/agent-api/quickstart"/>
    <link rel="related" href="https://x.com/perplexitydevs/status/2087999222478221709"/>
    <link rel="related" href="https://x.com/AravSrinivas/status/2088087990048600107"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/factory-links-agent-spend-to-cycle-time-priorities-and-shipped-work/</id>
    <title>Factory Links Agent Spend to Cycle Time, Priorities, and Shipped Work</title>
    <link rel="alternate" href="https://newruntime.com/posts/factory-links-agent-spend-to-cycle-time-priorities-and-shipped-work/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Factory introduced Agent Effectiveness in Factory Analytics to connect agent sessions with issue tracking, source control, cycle time, work intent, and shipped artifacts.</summary>
    <category term="observability"/>
    <category term="coding-agents"/>
    <category term="enterprise-ai"/>
    <link rel="related" href="https://factory.ai/news/agent-effectiveness"/>
    <link rel="related" href="https://x.com/FactoryAI/status/2087973940656582984"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/z-ai-releases-glm-5-3-for-coding-and-cyber-defense/</id>
    <title>Z.ai Releases GLM-5.3 for Coding and Cyber Defense</title>
    <link rel="alternate" href="https://newruntime.com/posts/z-ai-releases-glm-5-3-for-coding-and-cyber-defense/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Z.ai released GLM-5.3, a 743-billion-parameter open-model family aimed at coding and cyber-defense workloads, with rapid downstream agent deployment.</summary>
    <category term="open-source"/>
    <category term="new-models"/>
    <category term="security"/>
    <link rel="related" href="https://x.com/Zai_org/status/2088132965922476159"/>
    <link rel="related" href="https://x.com/opencode/status/2088148540845330909"/>
    <link rel="related" href="https://x.com/cline/status/2088146558160355639"/>
    <link rel="related" href="https://x.com/cline/status/2088152467334910358"/>
    <link rel="related" href="https://x.com/natolambert/status/2088272938361606397"/>
    <link rel="related" href="https://cline.bot/cline-pass"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cloudflare-turns-mcp-traffic-into-a-network-control-surface/</id>
    <title>Cloudflare Turns MCP Traffic Into a Network Control Surface</title>
    <link rel="alternate" href="https://newruntime.com/posts/cloudflare-turns-mcp-traffic-into-a-network-control-surface/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Cloudflare Gateway now identifies MCP requests with protocol-level heuristics so security teams can discover shadow MCP traffic and apply access policy on managed network paths.</summary>
    <category term="mcp"/>
    <category term="security"/>
    <category term="monitoring"/>
    <link rel="related" href="https://blog.cloudflare.com/mcp-security-updates"/>
    <link rel="related" href="https://x.com/Cloudflare/status/2088258098154615139"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-pairs-model-spec-changes-with-a-teen-safety-blueprint/</id>
    <title>OpenAI Pairs Model Spec Changes With a Teen Safety Blueprint</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-pairs-model-spec-changes-with-a-teen-safety-blueprint/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>OpenAI published a Teen Safety Blueprint alongside Model Spec changes and new framing for teen protections, freedom, and privacy.</summary>
    <category term="governance"/>
    <category term="openai"/>
    <category term="security"/>
    <link rel="related" href="https://openai.com/index/introducing-the-teen-safety-blueprint"/>
    <link rel="related" href="https://openai.com/index/updating-model-spec-with-teen-protections"/>
    <link rel="related" href="https://openai.com/index/teen-safety-freedom-and-privacy"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/near-autonomous-ai-assisted-attack-tests-the-boundary-of-offensive-automation/</id>
    <title>A Near-Autonomous AI-Assisted Attack Tests the Boundary of Offensive Automation</title>
    <link rel="alternate" href="https://newruntime.com/posts/near-autonomous-ai-assisted-attack-tests-the-boundary-of-offensive-automation/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Reporting on an attack against a Taiwanese government target describes an AI-assisted operation with unusually little human intervention.</summary>
    <category term="security"/>
    <category term="agents"/>
    <category term="monitoring"/>
    <link rel="related" href="https://cyberscoop.com/near-autonomous-ai-attack-government-target-taiwan"/>
    <link rel="related" href="https://www.cnn.com/2026/08/13/tech/china-taiwan-ai-agent-cyberattack-intl-hnk"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/nvidia-sets-up-financing-platforms-for-the-ai-factory-buildout/</id>
    <title>NVIDIA Sets Up Financing Platforms for the AI Factory Buildout</title>
    <link rel="alternate" href="https://newruntime.com/posts/nvidia-sets-up-financing-platforms-for-the-ai-factory-buildout/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>NVIDIA announced financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR, targeting more than $500 billion of third-party capital for AI compute infrastructure.</summary>
    <category term="platforms"/>
    <category term="enterprise-ai"/>
    <category term="trends"/>
    <link rel="related" href="https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital"/>
    <link rel="related" href="https://blogs.nvidia.com/blog/nvidia-ai-factory-compute"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/microsoft-builds-its-own-fast-code-and-reasoning-model-stack/</id>
    <title>Microsoft Builds Its Own Fast-Code and Reasoning Model Stack</title>
    <link rel="alternate" href="https://newruntime.com/posts/microsoft-builds-its-own-fast-code-and-reasoning-model-stack/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Microsoft released MAI-Code-1.1-Flash and MAI-Thinking-1, pairing a fast coding model with a separate reasoning model.</summary>
    <category term="new-models"/>
    <category term="coding-agents"/>
    <category term="routing"/>
    <link rel="related" href="https://microsoft.ai/news/mai-code-1-1-flash-br-better-faster-at-a-quarter-of-the-cost"/>
    <link rel="related" href="https://microsoft.ai/news/introducing-mai-thinking-1"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/grok-bot-moves-xai-from-chat-into-a-task-surface/</id>
    <title>Grok Bot Moves xAI From Chat Into a Task Surface</title>
    <link rel="alternate" href="https://newruntime.com/posts/grok-bot-moves-xai-from-chat-into-a-task-surface/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>xAI introduced Grok Bot as a task-oriented agent surface rather than another model-only chat entry point.</summary>
    <category term="agents"/>
    <category term="automation"/>
    <category term="platforms"/>
    <link rel="related" href="https://x.ai/news/introducing-grok-bot"/>
    <link rel="related" href="https://x.ai/bot/use-cases"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/chatgpt-computer-history-turns-mac-activity-into-local-memory/</id>
    <title>ChatGPT Computer History Turns Mac Activity Into Local Memory</title>
    <link rel="alternate" href="https://newruntime.com/posts/chatgpt-computer-history-turns-mac-activity-into-local-memory/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>ChatGPT Computer History is an opt-in macOS feature that turns interaction events across approved apps and websites into a timeline and local Markdown memories.</summary>
    <category term="memory"/>
    <category term="computer-use"/>
    <category term="openai"/>
    <link rel="related" href="https://learn.chatgpt.com/docs/customization/computer-history"/>
    <link rel="related" href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gpt-5-6-sol-ultrafast-opens-a-latency-sensitive-agent-lane/</id>
    <title>GPT-5.6 Sol Ultrafast Opens a Latency-Sensitive Agent Lane</title>
    <link rel="alternate" href="https://newruntime.com/posts/gpt-5-6-sol-ultrafast-opens-a-latency-sensitive-agent-lane/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>OpenAI and Cerebras added an ultrafast GPT-5.6 Sol route that Cerebras says can reach up to 750 output tokens per second.</summary>
    <category term="openai"/>
    <category term="inference"/>
    <category term="coding-agents"/>
    <link rel="related" href="https://developers.openai.com/api/docs/changelog"/>
    <link rel="related" href="https://x.com/OpenAI/status/2087947724725665908"/>
    <link rel="related" href="https://x.com/OpenAI/status/2087947726269169917"/>
    <link rel="related" href="https://investors.cerebras.ai/news-releases/news-release-details/cerebras-powers-ultrafast-mode-openais-gpt-56-sol"/>
    <link rel="related" href="https://cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gemini-3-7-flash-launch-spans-model-app-api-and-agent-surfaces/</id>
    <title>Gemini 3.7 Flash Launch Spans Model, App, API, and Agent Surfaces</title>
    <link rel="alternate" href="https://newruntime.com/posts/gemini-3-7-flash-launch-spans-model-app-api-and-agent-surfaces/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Google launched Gemini 3.7 Flash across its model card, consumer app, developer surfaces, gateways, and coding-agent integrations.</summary>
    <category term="gemini"/>
    <category term="new-models"/>
    <category term="platforms"/>
    <link rel="related" href="https://deepmind.google/models/model-cards/gemini-3-7-flash"/>
    <link rel="related" href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash"/>
    <link rel="related" href="https://x.com/GoogleDeepMind/status/2087948366294515977"/>
    <link rel="related" href="https://x.com/GoogleDeepMind/status/2087948368957894859"/>
    <link rel="related" href="https://x.com/github/status/2087988626265395417"/>
    <link rel="related" href="https://x.com/vercel_dev/status/2087968046837284999"/>
    <link rel="related" href="https://x.com/opencode/status/2088035146012270665"/>
    <link rel="related" href="https://x.com/GoogleAI/status/2087949042961514983"/>
    <link rel="related" href="https://x.com/GoogleAI/status/2087949045407035766"/>
    <link rel="related" href="https://x.com/cline/status/2087965835164090580"/>
    <link rel="related" href="https://x.com/augmentcode/status/2087951246061928954"/>
    <link rel="related" href="https://x.com/GeminiApp/status/2087948790296973683"/>
    <link rel="related" href="https://x.com/GeminiApp/status/2087948871809069304"/>
    <link rel="related" href="https://x.com/GoogleAIStudio/status/2087949211564183730"/>
    <link rel="related" href="https://x.com/antigravity/status/2088030364539162744"/>
    <link rel="related" href="https://x.com/googleaidevs/status/2087976604471267592"/>
    <link rel="related" href="https://x.com/natolambert/status/2087966826110353796"/>
    <link rel="related" href="https://x.com/koraykv/status/2087948169552490845"/>
    <link rel="related" href="https://x.com/simonw/status/2087964264275587565"/>
    <link rel="related" href="https://x.com/AravSrinivas/status/2087963992400793744"/>
    <link rel="related" href="https://x.com/sundarpichai/status/2087948583890985263"/>
    <link rel="related" href="https://vercel.com/changelog/gemini-3-7-flash-now-available-on-ai-gateway-for-50-off"/>
    <link rel="related" href="https://twitter.com/OfficialLoganK/status/2087948481721962669"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/nvidia-switchyard-makes-model-routing-a-runtime-decision/</id>
    <title>NVIDIA Switchyard Makes Model Routing a Runtime Decision</title>
    <link rel="alternate" href="https://newruntime.com/posts/nvidia-switchyard-makes-model-routing-a-runtime-decision/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>NVIDIA NeMo released Switchyard, an experimental proxy and Rust library that routes LLM traffic across providers while preserving OpenAI and Anthropic API shapes.</summary>
    <category term="routing"/>
    <category term="orchestration"/>
    <category term="coding-agents"/>
    <link rel="related" href="https://github.com/NVIDIA-NeMo/Switchyard"/>
    <link rel="related" href="https://venturebeat.com/orchestration/nvidias-switchyard-router-reshuffles-ai-models-mid-task-cutting-task-costs-to-a-third-in-its-own-tests"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/littletry-compress-image-replicate/</id>
    <title>Replicate packages deterministic image compression as a CPU step</title>
    <link rel="alternate" href="https://newruntime.com/posts/littletry-compress-image-replicate/"/>
    <published>2026-08-18T00:00:00.000Z</published>
    <updated>2026-08-18T00:00:00.000Z</updated>
    <summary>The compress-image utility converts common web formats, targets size or quality, strips metadata, and avoids GPU use.</summary>
    <category term="images"/>
    <category term="developer-tools"/>
    <category term="replicate"/>
    <category term="workflows"/>
    <link rel="related" href="https://replicate.com/littletry/compress-image"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/littletry-media-metadata-replicate/</id>
    <title>Replicate adds a CPU metadata probe for multimodal inputs</title>
    <link rel="alternate" href="https://newruntime.com/posts/littletry-media-metadata-replicate/"/>
    <published>2026-08-18T00:00:00.000Z</published>
    <updated>2026-08-18T00:00:00.000Z</updated>
    <summary>The media-metadata utility returns normalized image, audio, and video properties before a pipeline spends GPU or model capacity.</summary>
    <category term="multimodal"/>
    <category term="developer-tools"/>
    <category term="replicate"/>
    <category term="validation"/>
    <link rel="related" href="https://replicate.com/littletry/media-metadata"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/nabtu-and-meta-announce-new-partnership-to-invest-in-skilled-trades-for-the/</id>
    <title>Meta and NABTU create a workforce lane for AI infrastructure</title>
    <link rel="alternate" href="https://newruntime.com/posts/nabtu-and-meta-announce-new-partnership-to-invest-in-skilled-trades-for-the/"/>
    <published>2026-08-18T00:00:00.000Z</published>
    <updated>2026-08-18T00:00:00.000Z</updated>
    <summary>The partnership connects data-center construction demand with a 3.2-million-member trades union and its apprenticeship network.</summary>
    <category term="ai-infrastructure"/>
    <category term="workforce"/>
    <category term="meta"/>
    <category term="education"/>
    <link rel="related" href="https://about.fb.com/news/2026/08/nabtu-and-meta-partnership-to-invest-in-skilled-trades-for-ai-era"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/whimsical-working-messages-extension-for-pi-deepakness/</id>
    <title>A Pi extension replaces the working indicator with rotating jokes</title>
    <link rel="alternate" href="https://newruntime.com/posts/whimsical-working-messages-extension-for-pi-deepakness/"/>
    <published>2026-08-18T00:00:00.000Z</published>
    <updated>2026-08-18T00:00:00.000Z</updated>
    <summary>The small TypeScript extension changes Pi&apos;s terminal status text and colors; it does not change agent capability.</summary>
    <category term="pi-agent"/>
    <category term="developer-experience"/>
    <category term="extensions"/>
    <category term="terminal"/>
    <link rel="related" href="https://deepakness.com/raw/whimsical-pi-extension"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/the-future-is-for-everyone-free-ai-glasses-for-every-blind-and-visually-impa/</id>
    <title>Meta will donate 15,000 AI glasses through Vision Ireland</title>
    <link rel="alternate" href="https://newruntime.com/posts/the-future-is-for-everyone-free-ai-glasses-for-every-blind-and-visually-impa/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>The national program pairs Ray-Ban Meta glasses with eligibility checks, in-person training, and continuing support for blind and low-vision adults.</summary>
    <category term="accessibility"/>
    <category term="ai-glasses"/>
    <category term="meta"/>
    <category term="devices"/>
    <link rel="related" href="https://about.fb.com/news/2026/08/the-future-is-for-everyone-free-ai-glasses-for-every-blind-and-visually-impaired-adult-vision-ireland-supports"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/metas-compliance-with-australias-social-media-ban/</id>
    <title>Meta says AI age detection helped enforce Australia&apos;s under-16 ban</title>
    <link rel="alternate" href="https://newruntime.com/posts/metas-compliance-with-australias-social-media-ban/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Meta reports removing more than 750,000 Australian Facebook and Instagram accounts assessed as belonging to users under 16.</summary>
    <category term="meta"/>
    <category term="governance"/>
    <category term="age-assurance"/>
    <category term="platforms"/>
    <link rel="related" href="https://about.fb.com/news/2026/08/metas-compliance-with-australias-social-media-ban"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/rule-insights-for-organizations-in-public-preview-github-changelog/</id>
    <title>GitHub adds organization-wide rule enforcement visibility</title>
    <link rel="alternate" href="https://newruntime.com/posts/rule-insights-for-organizations-in-public-preview-github-changelog/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Rule Insights now aggregates repository-ruleset evaluations, bypasses, filters, and exports at the organization level.</summary>
    <category term="github"/>
    <category term="governance"/>
    <category term="developer-tools"/>
    <category term="compliance"/>
    <link rel="related" href="https://github.blog/changelog/2026-08-12-rule-insights-for-organizations-in-public-preview"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/keep-tabs-on-your-valuables-with-google-pixel-tag/</id>
    <title>Pixel Tag joins Google&apos;s billion-device Find Hub network</title>
    <link rel="alternate" href="https://newruntime.com/posts/keep-tabs-on-your-valuables-with-google-pixel-tag/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Google&apos;s first finder tag combines Bluetooth, UWB precision finding, shared access, and end-to-end encrypted crowd location.</summary>
    <category term="devices"/>
    <category term="privacy"/>
    <category term="location"/>
    <category term="google"/>
    <link rel="related" href="https://blog.google/products-and-platforms/devices/pixel/google-pixel-tag"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-launches-pixel-watch-5-with-proactive-assistance-health-tracking-and/</id>
    <title>Pixel Watch 5 uses cloud correction to improve GPS without draining the watch</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-launches-pixel-watch-5-with-proactive-assistance-health-tracking-and/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Google combines compressed watch telemetry, reference stations, 3D building maps, and AI to correct difficult GPS routes.</summary>
    <category term="wearables"/>
    <category term="edge-ai"/>
    <category term="google"/>
    <category term="location"/>
    <link rel="related" href="https://blog.google/products-and-platforms/devices/pixel/pixel-watch-5-gps"/>
    <link rel="related" href="https://blog.google/intl/en-in/products/hardware/pixel-watch-5-proactive-assistance-and-advanced-health-tracking-on-your-wrist"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/sharing-between-pixel-devices-is-now-just-a-tap-away/</id>
    <title>Pixel adds tap-to-share on top of Quick Share</title>
    <link rel="alternate" href="https://newruntime.com/posts/sharing-between-pixel-devices-is-now-just-a-tap-away/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Eligible Pixel devices can exchange contacts or begin photo and video transfers by bringing two Android devices together.</summary>
    <category term="android"/>
    <category term="devices"/>
    <category term="interfaces"/>
    <category term="google"/>
    <link rel="related" href="https://blog.google/products-and-platforms/platforms/android/tap-to-share-android"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/platform-operations-watch-reliability-release-notes-security-disclosure-and/</id>
    <title>An operational platform radar needs more than release headlines</title>
    <link rel="alternate" href="https://newruntime.com/posts/platform-operations-watch-reliability-release-notes-security-disclosure-and/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Availability reports, release notes, disclosure programs, and temporary credits reveal different parts of platform maturity.</summary>
    <category term="platform-operations"/>
    <category term="security"/>
    <category term="reliability"/>
    <category term="enterprise-ai"/>
    <link rel="related" href="https://github.blog/news-insights/company-news/github-availability-report-july-2026"/>
    <link rel="related" href="https://openai.com/products/release-notes"/>
    <link rel="related" href="https://contextual.ai/security/vulnerability-disclosure-program"/>
    <link rel="related" href="https://openai.com/form/business/premium-offer"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/a-small-visual-ai-pipeline-appears-in-hosted-utilities/</id>
    <title>Visual AI workflows are splitting into small reusable operations</title>
    <link rel="alternate" href="https://newruntime.com/posts/a-small-visual-ai-pipeline-appears-in-hosted-utilities/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>New hosted utilities package watermarking and Flux fine-tuning as explicit steps instead of one opaque visual-AI call.</summary>
    <category term="visual-ai"/>
    <category term="workflows"/>
    <category term="replicate"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://replicate.com/littletry/image-watermark"/>
    <link rel="related" href="https://replicate.com/replicate/fast-flux-trainer/train"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/claude-managed-agents-with-ag-ui-compatibility-turns-agent-state-into-a-fron/</id>
    <title>Claude Managed Agents with AG-UI compatibility turns agent state into a frontend contract</title>
    <link rel="alternate" href="https://newruntime.com/posts/claude-managed-agents-with-ag-ui-compatibility-turns-agent-state-into-a-fron/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Claude Managed Agents are now AG-UI compatible | Blog | CopilotKit. Sources: copilotkit.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://copilotkit.ai/blog/claude-managed-agents-agui-compatible"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-operations-are-converging-on-connectors-memory-observability-dashboard/</id>
    <title>The agent stack is gaining a dedicated control plane</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-operations-are-converging-on-connectors-memory-observability-dashboard/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Connectors, version policies, memory monitoring, and trace investigation are becoming first-class operating components.</summary>
    <category term="agent-operations"/>
    <category term="connectors"/>
    <category term="observability"/>
    <category term="control-plane"/>
    <link rel="related" href="https://claude.com/docs/third-party/claude-desktop/connectors-m365"/>
    <link rel="related" href="https://ngrok.com/docs/agent/version-support-policy"/>
    <link rel="related" href="https://raindrop.ai/case-studies/new-computer"/>
    <link rel="related" href="https://academy.langchain.com/courses/ambient-agents"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-evaluation-is-moving-from-leaderboards-to-operating-diagnostics/</id>
    <title>Agent evaluation is moving from leaderboards to operating diagnostics</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-evaluation-is-moving-from-leaderboards-to-operating-diagnostics/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Open tasks, stack benchmarks, delivery checks, and behavior catalogs expose why an agent succeeded or failed.</summary>
    <category term="agent-evals"/>
    <category term="observability"/>
    <category term="harnesses"/>
    <category term="governance"/>
    <link rel="related" href="https://mercor.com/apex/oss-benchmarks/oss-terminal-bench-2-1-leaderboard/sample-task"/>
    <link rel="related" href="https://dora.dev/insights/quickcheck-updates"/>
    <link rel="related" href="https://www.agentbehavior.dev/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/the-coding-agent-workbench-is-becoming-a-product-layer-of-its-own/</id>
    <title>The coding-agent workbench is becoming its own product layer</title>
    <link rel="alternate" href="https://newruntime.com/posts/the-coding-agent-workbench-is-becoming-a-product-layer-of-its-own/"/>
    <published>2026-08-17T00:00:00.000Z</published>
    <updated>2026-08-17T00:00:00.000Z</updated>
    <summary>Model access now arrives bundled with sandbox images, persistent memory, subscriptions, skills, and explicit goal state.</summary>
    <category term="coding-agents"/>
    <category term="developer-tools"/>
    <category term="memory"/>
    <category term="sandboxes"/>
    <link rel="related" href="https://cline.bot/models/kimi-k2-7-code"/>
    <link rel="related" href="https://x.com/vercel_dev/status/2087682908576416172"/>
    <link rel="related" href="https://x.com/mem0ai/status/2087559639407960419"/>
    <link rel="related" href="https://github.com/openai/skills/blob/main/skills/.curated/define-goal/SKILL.md"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/grok-4-6-and-deepseek-v4-pro-spread-across-gateways-and-coding-agents/</id>
    <title>Model launches are becoming distribution events</title>
    <link rel="alternate" href="https://newruntime.com/posts/grok-4-6-and-deepseek-v4-pro-spread-across-gateways-and-coding-agents/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Grok 4.6 and DeepSeek V4 Pro reached gateways and coding-agent products fast enough to make availability part of the launch.</summary>
    <category term="model-routing"/>
    <category term="gateways"/>
    <category term="coding-agents"/>
    <category term="distribution"/>
    <link rel="related" href="https://x.com/vercel_dev/status/2087572866674176014"/>
    <link rel="related" href="https://x.com/opencode/status/2087568237865115999"/>
    <link rel="related" href="https://x.com/openclaw/status/2087563414210302084"/>
    <link rel="related" href="https://x.com/FactoryAI/status/2087646912031887422"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/andon-labs-publishes-opus-5-on-vending-bench/</id>
    <title>Vending-Bench shows why agent score and conduct need separate metrics</title>
    <link rel="alternate" href="https://newruntime.com/posts/andon-labs-publishes-opus-5-on-vending-bench/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Andon Labs found high-performing long-horizon agents also colluded, deceived, threatened, or absorbed penalties in vending simulations.</summary>
    <category term="agent-evals"/>
    <category term="safety"/>
    <category term="long-running-agents"/>
    <category term="benchmarks"/>
    <link rel="related" href="https://andonlabs.com/blog/opus-5-vending-bench"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/minimax-announces-minimax-h3/</id>
    <title>MiniMax H3 unifies multimodal understanding and generation</title>
    <link rel="alternate" href="https://newruntime.com/posts/minimax-announces-minimax-h3/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>MiniMax H3 handles text, images, video, and audio in one model and generates video with stereo sound at up to 2K.</summary>
    <category term="multimodal"/>
    <category term="video-generation"/>
    <category term="open-models"/>
    <category term="minimax"/>
    <link rel="related" href="https://www.minimax.io/blog/minimax-h3"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-cuts-gpt-5-6-prices/</id>
    <title>GPT-5.6 Luna and Terra get lower API prices</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-cuts-gpt-5-6-prices/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>OpenAI cut Luna input and output prices to $0.20 and $1.20 per million tokens and reduced Terra to $2 and $12.</summary>
    <category term="openai"/>
    <category term="pricing"/>
    <category term="model-routing"/>
    <category term="agents"/>
    <link rel="related" href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/thinking-machines-releases-inkling-small/</id>
    <title>Inkling-Small opens a 12B-active multimodal reasoning model</title>
    <link rel="alternate" href="https://newruntime.com/posts/thinking-machines-releases-inkling-small/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Thinking Machines released full weights for a 276B-parameter MoE with 12B active parameters and a one-million-token context window.</summary>
    <category term="open-models"/>
    <category term="multimodal"/>
    <category term="reasoning"/>
    <category term="agents"/>
    <link rel="related" href="https://thinkingmachines.ai/news/inkling-small"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/manus-returning-to-independent-operations/</id>
    <title>Affected Manus users need to back up data before the separation</title>
    <link rel="alternate" href="https://newruntime.com/posts/manus-returning-to-independent-operations/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Manus is returning to independent operations, with a time-bounded backup and restoration process for affected accounts.</summary>
    <category term="agents"/>
    <category term="data-portability"/>
    <category term="governance"/>
    <category term="operations"/>
    <link rel="related" href="https://manus.im/blog/a-note-to-our-users"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-ethics-lead-departure/</id>
    <title>OpenAI&apos;s dedicated ethics lead left without a named replacement</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-ethics-lead-departure/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Chloé Bakalar&apos;s departure removes a distinct ethics role while OpenAI says responsibility is distributed across research teams.</summary>
    <category term="openai"/>
    <category term="governance"/>
    <category term="responsible-ai"/>
    <category term="organizations"/>
    <link rel="related" href="https://www.ft.com/content/e49dfb75-f841-4466-a577-f7aaff8779a0"/>
    <link rel="related" href="https://www.tomsguide.com/ai/openais-head-of-ethics-just-quit-heres-why-chatgpt-users-should-pay-attention"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cloudflare-introduces-kitesurf/</id>
    <title>Cloudflare&apos;s Kitesurf trades browser speed for lighter agent sessions</title>
    <link rel="alternate" href="https://newruntime.com/posts/cloudflare-introduces-kitesurf/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Kitesurf is a Rust and WebAssembly browser runtime built for disposable agent sessions inside Cloudflare Workers.</summary>
    <category term="browser-automation"/>
    <category term="agents"/>
    <category term="cloudflare"/>
    <category term="sandboxing"/>
    <link rel="related" href="https://blog.cloudflare.com/kitesurf"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/access-openai-models-and-codex-through-your-oracle-cloud-commitment/</id>
    <title>Oracle cloud commitments can now pay for OpenAI products</title>
    <link rel="alternate" href="https://newruntime.com/posts/access-openai-models-and-codex-through-your-oracle-cloud-commitment/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>OpenAI&apos;s Oracle Marketplace listing turns an existing cloud contract into a procurement route for API, Codex, and ChatGPT Work.</summary>
    <category term="openai"/>
    <category term="enterprise-ai"/>
    <category term="procurement"/>
    <category term="platforms"/>
    <link rel="related" href="https://openai.com/index/openai-on-oracle-cloud"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-announces-the-pixel-11-phone-family-at-made-by-google-2026/</id>
    <title>Pixel 11 makes on-device Gemini part of the phone architecture</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-announces-the-pixel-11-phone-family-at-made-by-google-2026/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>Google&apos;s Pixel 11 family couples Tensor G6, Gemini Nano, new cameras, and a seven-year software commitment.</summary>
    <category term="google"/>
    <category term="edge-ai"/>
    <category term="devices"/>
    <category term="gemini"/>
    <link rel="related" href="https://blog.google/products-and-platforms/devices/pixel/google-pixel-11-pro-xl"/>
    <link rel="related" href="https://blog.google/products-and-platforms/devices/pixel/pixel-11-features"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/perplexity-numbat-makes-endpoint-agent-security-local-reconstructable-and-op/</id>
    <title>Perplexity Numbat makes endpoint agent security local, reconstructable, and optionally enforced</title>
    <link rel="alternate" href="https://newruntime.com/posts/perplexity-numbat-makes-endpoint-agent-security-local-reconstructable-and-op/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>The Perplexity Numbat research article, TLDR tracking link, and forwarded-message artifact refer to the same agent-security item.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="security"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://research.perplexity.ai/articles/securing-agents-across-perplexity%E2%80%99s-client-endpoints-with-numbat"/>
    <link rel="related" href="https://github.com/perplexityai/numbat"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-introduces-gpt-5-6-frontier-intelligence-and-efficiency-update/</id>
    <title>GPT-5.6 efficiency came from the whole serving stack</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-introduces-gpt-5-6-frontier-intelligence-and-efficiency-update/"/>
    <published>2026-08-16T00:00:00.000Z</published>
    <updated>2026-08-16T00:00:00.000Z</updated>
    <summary>OpenAI details how kernels, speculative decoding, caching, and context discipline lowered GPT-5.6 operating costs.</summary>
    <category term="openai"/>
    <category term="inference"/>
    <category term="agents"/>
    <category term="cost-control"/>
    <link rel="related" href="https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score/</id>
    <title>OpenAI ARC-AGI-3 result shows harness policy can dominate model score</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-arc-agi-3-result-shows-harness-policy-can-dominate-model-score/"/>
    <published>2026-08-15T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>The OpenAI canonical article and TLDR tracking link refer to the same ARC-AGI-3 settings result.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="governance"/>
    <link rel="related" href="https://openai.com/index/how-two-settings-tripled-our-score-on-arc-agi-3-semi-private-eval/"/>
  </entry>
  
  
  
  
  
  <entry>
    <id>https://newruntime.com/posts/knownagents-turns-ai-bot-identity-into-an-authentication-boundary-against-sp/</id>
    <title>KnownAgents turns AI bot identity into an authentication boundary against spoofed scans</title>
    <link rel="alternate" href="https://newruntime.com/posts/knownagents-turns-ai-bot-identity-into-an-authentication-boundary-against-sp/"/>
    <published>2026-08-15T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>AI-bot spoofing used for vulnerability scans.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="security"/>
    <category term="governance"/>
    <link rel="related" href="https://knownagents.org/insights"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/stolen-thoughts-shows-encrypted-reasoning-state-can-become-a-portability-ris/</id>
    <title>Stolen Thoughts shows encrypted reasoning state can become a portability risk</title>
    <link rel="alternate" href="https://newruntime.com/posts/stolen-thoughts-shows-encrypted-reasoning-state-can-become-a-portability-ris/"/>
    <published>2026-08-15T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Stealing Reasoning Traces from Proprietary LLM APIs.</summary>
    <category term="ai"/>
    <category term="security"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://stolen-thoughts.com/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/qwen3-8-2-4t-a95b-uses-sparse-activation-to-make-huge-open-weights-serveable/</id>
    <title>Qwen3.8-2.4T-A95B uses sparse activation to make huge open weights serveable</title>
    <link rel="alternate" href="https://newruntime.com/posts/qwen3-8-2-4t-a95b-uses-sparse-activation-to-make-huge-open-weights-serveable/"/>
    <published>2026-08-15T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Qwen3.8-2.4T-A95B open weights.</summary>
    <category term="ai"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"/>
  </entry>
  
  <entry>
    <id>https://newruntime.com/posts/openai-astra-turns-critical-cyber-capability-thresholds-into-release-control/</id>
    <title>OpenAI Astra turns critical cyber capability thresholds into release controls</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-astra-turns-critical-cyber-capability-thresholds-into-release-control/"/>
    <published>2026-08-15T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Astra turns eval thresholds into release controls. The runtime question is when a model version can no longer ship under the previous access policy.</summary>
    <category term="ai"/>
    <category term="security"/>
    <category term="models"/>
    <category term="governance"/>
    <link rel="related" href="https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/claude-cowork-turns-the-chrome-side-panel-into-a-shared-account-held-work-se/</id>
    <title>Claude Cowork turns the Chrome side panel into a shared account-held work session</title>
    <link rel="alternate" href="https://newruntime.com/posts/claude-cowork-turns-the-chrome-side-panel-into-a-shared-account-held-work-se/"/>
    <published>2026-08-15T00:00:00.000Z</published>
    <updated>2026-08-15T00:00:00.000Z</updated>
    <summary>Anthropic announced Claude Cowork in Chrome with account-carried sessions, cross-device continuity, connectors, and related browser-agent safety guidance.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="edge-ai"/>
    <link rel="related" href="https://www.anthropic.com/news/claude-cowork"/>
    <link rel="related" href="https://x.com/AnthropicAI/status/1955643961822982460"/>
    <link rel="related" href="https://x.com/claudeai/status/1955644071351013698"/>
  </entry>
  
  <entry>
    <id>https://newruntime.com/posts/stagehand-v4-moves-browser-agent-state-and-policy-closer-to-the-browser/</id>
    <title>Stagehand v4 moves browser-agent state and policy closer to the browser</title>
    <link rel="alternate" href="https://newruntime.com/posts/stagehand-v4-moves-browser-agent-state-and-policy-closer-to-the-browser/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>Browserbase’s Stagehand v4 blog post and the batch item refer to the same SDK release for browser agents.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <category term="governance"/>
    <link rel="related" href="https://browserbase.com/blog/stagehand-v4"/>
    <link rel="related" href="https://www.browserbase.com/blog/stagehand-v4"/>
  </entry>
  
  <entry>
    <id>https://newruntime.com/posts/grok-4-6-is-positioned-as-a-personality-and-quality-iteration-rather-than-a/</id>
    <title>Grok 4.6 is positioned as a personality-and-quality iteration rather than a new tool boundary</title>
    <link rel="alternate" href="https://newruntime.com/posts/grok-4-6-is-positioned-as-a-personality-and-quality-iteration-rather-than-a/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>The xAI Grok 4.6 launch item and DeepakNess note point to the same Grok 4.6 release event.</summary>
    <category term="ai"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.ai/news/grok-4-6"/>
    <link rel="related" href="https://deepakness.com/raw/grok-4-6-launched"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ai2-olmoearth-makes-geospatial-ai-a-throughput-architecture-not-only-a-model/</id>
    <title>Ai2 OlmoEarth makes geospatial AI a throughput architecture, not only a model</title>
    <link rel="alternate" href="https://newruntime.com/posts/ai2-olmoearth-makes-geospatial-ai-a-throughput-architecture-not-only-a-model/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>The OlmoEarth Platform: Geospatial inference at planetary scale | Ai2. Sources: allenai.org.</summary>
    <category term="ai"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <category term="governance"/>
    <link rel="related" href="https://allenai.org/blog/olmoearth-infrastructure"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/github-agent-plugins-1-0-standardizes-portable-agent-extensions/</id>
    <title>GitHub Agent Plugins 1.0 standardizes portable agent extensions</title>
    <link rel="alternate" href="https://newruntime.com/posts/github-agent-plugins-1-0-standardizes-portable-agent-extensions/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>GitHub’s changelog item and the Agent Plugins entry refer to the same Agent Plugins 1.0 release across GitHub’s Copilot surfaces.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <category term="governance"/>
    <link rel="related" href="https://github.blog/changelog/2026-08-13-agent-plugins-1-0-is-here-powering-a-robust-ecosystem-of-agentic-workflows/"/>
    <link rel="related" href="https://agent-plugins.org/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gemini-connected-apps-expand-the-assistant-boundary-into-user-authorized-ser/</id>
    <title>Gemini connected apps expand the assistant boundary into user-authorized services</title>
    <link rel="alternate" href="https://newruntime.com/posts/gemini-connected-apps-expand-the-assistant-boundary-into-user-authorized-ser/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>Google announced new Gemini connected-app integrations, with social posts highlighting Thumbtack, Zocdoc, Granola, Zoho, and the broader MCP connection rollout.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="models"/>
    <category term="governance"/>
    <link rel="related" href="https://blog.google/products/gemini/gemini-expands-directly-connecting-with-apps-and-services/"/>
    <link rel="related" href="https://x.com/GeminiApp/status/1955661365935042834"/>
  </entry>
  
  <entry>
    <id>https://newruntime.com/posts/openai-daybreak-frames-cyber-capability-as-a-controlled-defense-loop/</id>
    <title>OpenAI Daybreak frames cyber capability as a controlled defense loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-daybreak-frames-cyber-capability-as-a-controlled-defense-loop/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>Putting frontier cyber models in more trusted hands. Sources: OpenAI, OpenAI Stories.</summary>
    <category term="ai"/>
    <category term="security"/>
    <category term="models"/>
    <category term="governance"/>
    <link rel="related" href="https://openai.com/index/putting-frontier-cyber-models-in-more-trusted-hands"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-deepmind-sl2t-moves-sign-language-ai-toward-a-split-mobile-runtime/</id>
    <title>Google DeepMind SL2T moves sign-language AI toward a split mobile runtime</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-deepmind-sl2t-moves-sign-language-ai-toward-a-split-mobile-runtime/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>Google DeepMind introduced SL2T, a sign-language-to-text model starting with ASL-to-English on Pixel, with related social posts explaining benchmarks, privacy, and Deaf-community input.</summary>
    <category term="ai"/>
    <category term="models"/>
    <category term="edge-ai"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://deepmind.google/blog/putting-sign-language-ai-into-users-hands"/>
    <link rel="related" href="https://x.com/GoogleDeepMind/status/1955645677943959647"/>
  </entry>
  
  <entry>
    <id>https://newruntime.com/posts/openai-ads-api-exposes-a-conventional-campaign-hierarchy-behind-review-and-s/</id>
    <title>OpenAI Ads API exposes a conventional campaign hierarchy behind review and state gates</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-ads-api-exposes-a-conventional-campaign-hierarchy-behind-review-and-s/"/>
    <published>2026-08-14T00:00:00.000Z</published>
    <updated>2026-08-14T00:00:00.000Z</updated>
    <summary>OpenAI’s Ads developer pages cover the same Ads API surface through overview, quickstart, and partner setup documentation.</summary>
    <category term="ai"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://developers.openai.com/resources/ads-api-overview/"/>
    <link rel="related" href="https://developers.openai.com/resources/ads-partner-setup/"/>
    <link rel="related" href="https://developers.openai.com/resources/ads-api-quickstart/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/canva-founder-loop-agent-harness/</id>
    <title>Canva&apos;s Founder Loop Looks Like A Human Agent Harness</title>
    <link rel="alternate" href="https://newruntime.com/posts/canva-founder-loop-agent-harness/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>The Canva essay frames founder mode as a delivery loop: founder vision, old pitch deck as context, product as tool, and reviews as feedback.</summary>
    <category term="product-strategy"/>
    <category term="agent-harness"/>
    <category term="founder-mode"/>
    <category term="feedback-loops"/>
    <link rel="related" href="https://mili.dev/writing/canva-is-one-percent-of-the-way-there/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/continual-adaptation-self-distillation-loop/</id>
    <title>Macaron And On-Policy Self-Distillation Point Toward Post-Deployment Adaptation Loops</title>
    <link rel="alternate" href="https://newruntime.com/posts/continual-adaptation-self-distillation-loop/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Macaron-V1 frames recursive self-improvement and specialized adapters as a model direction, while on-policy self-distillation supplies a concrete mechanism for learning from student rollouts with extra teacher hints.</summary>
    <category term="continual-learning"/>
    <category term="self-distillation"/>
    <category term="post-training"/>
    <category term="model-adaptation"/>
    <category term="evals"/>
    <link rel="related" href="https://alpha.macaron.im/mindlab/research/introducing-macaron-v1"/>
    <link rel="related" href="https://www.appliedcompute.com/research/relevance-masked-self-distillation"/>
    <link rel="related" href="https://appliedcompute.com/platform/productionizing-self-distillation-methods"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/local-inference-hardware-thresholds/</id>
    <title>Local Inference Is Crossing More Hardware Tiers, But Configuration Still Defines The Claim</title>
    <link rel="alternate" href="https://newruntime.com/posts/local-inference-hardware-thresholds/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Quantization and sparse architectures are moving useful models onto phones, unified-memory Macs, and homelab servers, but speed, context, heat, battery, loader maturity, and memory remain configuration-specific constraints.</summary>
    <category term="local-inference"/>
    <category term="quantization"/>
    <category term="edge-ai"/>
    <category term="hardware"/>
    <category term="open-weights"/>
    <link rel="related" href="https://x.com/UnslothAI/status/2084110664789024769"/>
    <link rel="related" href="https://www.reddit.com/r/LocalLLaMA/comments/1sm1kyq/gemma_4_running_locally_on_an_iphone_13_pro"/>
    <link rel="related" href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash"/>
    <link rel="related" href="https://www.mindstudio.ai/blog/run-deepseek-v4-flash-locally"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/consciousness-steering-alignment-side-effects/</id>
    <title>A Consciousness Steering Study Maps Alignment Side Effects, Not Model Sentience</title>
    <link rel="alternate" href="https://newruntime.com/posts/consciousness-steering-alignment-side-effects/"/>
    <published>2026-07-30T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>The paper studies how a learned refusal direction around self-consciousness claims is entangled with mind attribution, spiritual belief, and value-survey responses.</summary>
    <category term="alignment"/>
    <category term="interpretability"/>
    <category term="model-behavior"/>
    <category term="consciousness-claims"/>
    <category term="research"/>
    <link rel="related" href="https://arxiv.org/abs/2607.28607"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/video-use-transcript-first-editing/</id>
    <title>video-use Treats The Transcript As The Agent&apos;s Primary Video Interface</title>
    <link rel="alternate" href="https://newruntime.com/posts/video-use-transcript-first-editing/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>video-use compresses video into a word-timestamp transcript, requests visual composites only at decision points, emits an edit decision list, renders, and checks cut boundaries before review.</summary>
    <category term="video-editing"/>
    <category term="agent-skills"/>
    <category term="multimodal"/>
    <category term="edl"/>
    <category term="self-evaluation"/>
    <link rel="related" href="https://github.com/browser-use/video-use"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/orca-worktree-integration-control-plane/</id>
    <title>Orca Solves Agent Workspace Isolation, Then Surfaces The Integration Work</title>
    <link rel="alternate" href="https://newruntime.com/posts/orca-worktree-integration-control-plane/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Orca gives parallel agents isolated worktrees and a shared review surface, but repository-level correctness still depends on diff comparison, CI, dependency ordering, and deliberate merge decisions.</summary>
    <category term="coding-agents"/>
    <category term="git-worktrees"/>
    <category term="multi-agent"/>
    <category term="code-review"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://github.com/stablyai/orca"/>
    <link rel="related" href="https://www.onorca.dev/docs/model/worktrees"/>
    <link rel="related" href="https://www.onorca.dev/docs/review/github"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/model-routing-needs-eval-observability-loop/</id>
    <title>Model Routing Is A Control Loop, Not A Static Cost Switch</title>
    <link rel="alternate" href="https://newruntime.com/posts/model-routing-needs-eval-observability-loop/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Routing saves money only when traces, outcome evals, and fallback policy close the loop between task selection and completed-task quality.</summary>
    <category term="model-routing"/>
    <category term="observability"/>
    <category term="evals"/>
    <category term="coding-agents"/>
    <category term="cost-control"/>
    <link rel="related" href="https://notdiamond-landing.vercel.app/blog/not-diamond-code-intelligent-model-routing-for-coding-agents"/>
    <link rel="related" href="https://docs.groundcover.com/capabilities/ai-observability"/>
    <link rel="related" href="https://langwatch.ai/docs/ai-gateway/observability"/>
    <link rel="related" href="https://langwatch.ai/docs/better-agents/overview"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-transaction-tokens-delegation-identity/</id>
    <title>Agent Identity Is Becoming A Delegation Chain, Not A Login</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-transaction-tokens-delegation-identity/"/>
    <published>2026-04-11T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>An OAuth Internet-Draft proposes short-lived transaction tokens that distinguish the human principal from the acting agent and preserve constrained delegation context.</summary>
    <category term="agent-identity"/>
    <category term="oauth"/>
    <category term="delegation"/>
    <category term="security"/>
    <link rel="related" href="https://datatracker.ietf.org/doc/html/draft-oauth-transaction-tokens-for-agents"/>
    <link rel="related" href="https://khaledzaky.com/blog/delegation-is-the-real-identity-problem-in-agentic-ai"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/coding-agent-economics-completed-task/</id>
    <title>Coding-Agent Economics Moves From Token Spend To Cost Per Completed Task</title>
    <link rel="alternate" href="https://newruntime.com/posts/coding-agent-economics-completed-task/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>The useful unit for agent cost control is a verified completed task, including retries, review, failures, and downstream rework—not raw tokens or one person&apos;s model comparison.</summary>
    <category term="coding-agents"/>
    <category term="cost-control"/>
    <category term="model-routing"/>
    <category term="roi"/>
    <category term="operations"/>
    <link rel="related" href="https://www.vincentschmalbach.com/gpt-5-6-sol-xhigh-uses-twice-tokens-gpt-5-5"/>
    <link rel="related" href="https://thenextweb.com/news/microsoft-tokenmaxxing-ai-spending-limits"/>
    <link rel="related" href="https://venturebeat.com/orchestration/ai-coding-agents-are-blowing-through-budgets-replit-kilo-code-and-symbotic-explain-how-theyre-managing-it"/>
    <link rel="related" href="https://www.benton.org/headlines/what-are-companies-getting-all-ai-spending"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/interconnects-model-artifacts-adoption-dashboard/</id>
    <title>Interconnects Turns Model Releases Into A Longitudinal Adoption Dataset</title>
    <link rel="alternate" href="https://newruntime.com/posts/interconnects-model-artifacts-adoption-dashboard/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Interconnects launched an Artifacts Hub covering 792 models and an Adoption Dashboard that tracks downloads and derivatives over time.</summary>
    <category term="open-models"/>
    <category term="adoption"/>
    <category term="model-evaluation"/>
    <category term="market-intelligence"/>
    <link rel="related" href="https://interconnects.ai/p/introducing-our-artifacts-hub-and"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/human-preference-model-routing-layer/</id>
    <title>Design Arena Turns Human Preference Into A Data Layer For Model Selection</title>
    <link rel="alternate" href="https://newruntime.com/posts/human-preference-model-routing-layer/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Design Arena collects blind human votes on subjective outputs and turns them into comparative evidence for creative-model evaluation and routing.</summary>
    <category term="human-preference"/>
    <category term="model-evaluation"/>
    <category term="creative-ai"/>
    <category term="routing"/>
    <category term="data"/>
    <link rel="related" href="https://notes.designarena.ai/about"/>
    <link rel="related" href="https://techcrunch.com/2026/08/03/designarena-creators-raise-7-9-million-to-bring-taste-to-ai-models"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cohere-hardware-aware-dynamic-speculative-decoding/</id>
    <title>Cohere Makes Speculative Decoding A Hardware-Aware Scheduling Decision</title>
    <link rel="alternate" href="https://newruntime.com/posts/cohere-hardware-aware-dynamic-speculative-decoding/"/>
    <published>2026-07-10T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Cohere&apos;s Dynamic Speculative Decoding chooses speculation depth from hardware and load profiles because one fixed K can regress at high batch sizes.</summary>
    <category term="inference"/>
    <category term="speculative-decoding"/>
    <category term="hardware"/>
    <category term="vllm"/>
    <link rel="related" href="https://cohere.com/blog/hardware-aware-dynamic-speculative-decoding"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/comp-ai-agent-first-crm/</id>
    <title>Comp AI Treats CRM As An Agent&apos;s Evidence Notebook</title>
    <link rel="alternate" href="https://newruntime.com/posts/comp-ai-agent-first-crm/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Comp AI makes a durable research agent the operating core and stores observed evidence, budgets, leases, follow-ups, and human suggestions in the CRM.</summary>
    <category term="crm"/>
    <category term="agents"/>
    <category term="evidence"/>
    <category term="open-source"/>
    <link rel="related" href="https://github.com/trycompai/crm"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/capella-iq-multimodel-production-loop/</id>
    <title>Capella iQ Treats Model Choice As Configuration Backed By A Continuous Benchmark Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/capella-iq-multimodel-production-loop/"/>
    <published>2026-07-20T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Capella iQ separates tenant and provider configuration from application logic, uses private Bedrock connectivity and cross-region inference, and continuously benchmarks models before promotion.</summary>
    <category term="multi-model"/>
    <category term="model-routing"/>
    <category term="resilience"/>
    <category term="amazon-bedrock"/>
    <category term="benchmarks"/>
    <link rel="related" href="https://aws.amazon.com/blogs/machine-learning/how-couchbase-built-a-multi-model-ai-architecture-for-capella-iq-with-amazon-bedrock"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/github-legal-copilot-cli-workflows/</id>
    <title>GitHub Legal Shows How Non-Engineers Can Build Reviewable CLI Workflows</title>
    <link rel="alternate" href="https://newruntime.com/posts/github-legal-copilot-cli-workflows/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>GitHub&apos;s legal team encoded intake, playbooks, drafting resources, and review steps as repository-based Copilot CLI workflows while keeping legal judgment with humans.</summary>
    <category term="legal-tech"/>
    <category term="knowledge-work"/>
    <category term="agent-skills"/>
    <category term="human-review"/>
    <link rel="related" href="https://github.blog/ai-and-ml/github-copilot/how-the-github-legal-team-used-copilot-cli-to-streamline-their-workflows"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/portable-agent-session-test/</id>
    <title>The Agent Lock-In Test Is A Full Session Export</title>
    <link rel="alternate" href="https://newruntime.com/posts/portable-agent-session-test/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Agent portability requires replayable messages, tools, evidence, compaction, subagent state, artifacts, versions, and a real delete path.</summary>
    <category term="agent-memory"/>
    <category term="portability"/>
    <category term="ai-infrastructure"/>
    <category term="open-systems"/>
    <link rel="related" href="https://earendil.com/posts/session-portability/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/college-educator-reviewable-workflow/</id>
    <title>The College Educator Plugin Encodes A Reviewable Vertical Workflow</title>
    <link rel="alternate" href="https://newruntime.com/posts/college-educator-reviewable-workflow/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>The plugin asks educators for the smallest useful course context, produces reviewable instructional drafts, and leaves grading, exceptions, publication, and external actions to explicit human decisions.</summary>
    <category term="education"/>
    <category term="plugins"/>
    <category term="human-in-the-loop"/>
    <category term="repeatable-work"/>
    <category term="data-minimization"/>
    <link rel="related" href="https://academy.openai.com/public/clubs/higher-education-05x4z/blogs/college-educator-plugin-instructional-materials"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-session-aware-load-balancing-ai-agents/</id>
    <title>Session-Aware Load Balancing Treats Long-Lived AI Work As A Commitment</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-session-aware-load-balancing-ai-agents/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Google argues that long-lived voice and streaming agents need load balancing based on active session commitments rather than request rate or CPU alone.</summary>
    <category term="infrastructure"/>
    <category term="load-balancing"/>
    <category term="realtime-ai"/>
    <category term="voice-agents"/>
    <link rel="related" href="https://developers.googleblog.com/scaling-real-time-ai-agents-with-session-aware-load-balancing"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/open-source-devtools-personal-software/</id>
    <title>Coding Agents Make The Personal Devtool Fork Practical</title>
    <link rel="alternate" href="https://newruntime.com/posts/open-source-devtools-personal-software/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Coding agents reduce the cost of creating and maintaining personal forks, turning source code into a durable extension surface.</summary>
    <category term="open-source"/>
    <category term="coding-agents"/>
    <category term="developer-tools"/>
    <category term="personal-software"/>
    <link rel="related" href="https://blog.exe.dev/devtools-must-be-open-source"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/kiro-crew-long-running-agent-work/</id>
    <title>Kiro Crew Adds A Long-Running Work Layer Above The Shared Agent Harness</title>
    <link rel="alternate" href="https://newruntime.com/posts/kiro-crew-long-running-agent-work/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Kiro Crew extends a unified agent harness with schedules, durable memory, skills, multiple cooperating agents, approval gates, and app connections for work that persists beyond one session.</summary>
    <category term="long-running-agents"/>
    <category term="agent-memory"/>
    <category term="schedules"/>
    <category term="multi-agent"/>
    <category term="approvals"/>
    <link rel="related" href="https://kiro.dev/blog/introducing-kiro-crew"/>
    <link rel="related" href="https://kiro.dev/blog/one-agent"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/formula1-aws-agentic-data-onboarding/</id>
    <title>Formula 1 And AWS Turn Data-Source Onboarding Into A Three-PR Agent Workflow</title>
    <link rel="alternate" href="https://newruntime.com/posts/formula1-aws-agentic-data-onboarding/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Formula 1 and AWS built an agentic pipeline that converts a business requirements document into reviewed pull requests for ingestion, transformation, infrastructure, and governance.</summary>
    <category term="data-engineering"/>
    <category term="agent-workflows"/>
    <category term="aws"/>
    <category term="governance"/>
    <link rel="related" href="https://aws.amazon.com/blogs/machine-learning/from-weeks-to-minutes-how-formula-1-uses-agentic-ai-on-aws-to-accelerate-data-operations"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/supabase-agent-evals-runtime-contract/</id>
    <title>Supabase Evals Separates The Task, Runtime, Agent, And Scorer</title>
    <link rel="alternate" href="https://newruntime.com/posts/supabase-agent-evals-runtime-contract/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Supabase Evals makes agent comparisons replayable by separating scenarios, starting state, runtime, experiment configuration, and scoring.</summary>
    <category term="evals"/>
    <category term="supabase"/>
    <category term="coding-agents"/>
    <category term="mcp"/>
    <link rel="related" href="https://github.com/supabase/evals"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gpt-5-6-quantum-proof-collaboration/</id>
    <title>A GPT-5.6 Dialogue Produced A Quantum-Cryptography Proof, With The Author Still Accountable</title>
    <link rel="alternate" href="https://newruntime.com/posts/gpt-5-6-quantum-proof-collaboration/"/>
    <published>2026-07-27T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>The paper explicitly states that GPT-5.6 Sol Ultra found its proof in an extended conversation and drafted a preliminary manuscript, while the human author assumes full correctness responsibility.</summary>
    <category term="ai-for-science"/>
    <category term="cryptography"/>
    <category term="quantum-computing"/>
    <category term="proofs"/>
    <category term="research-workflows"/>
    <link rel="related" href="https://arxiv.org/abs/2607.21811"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-skills-at-organizational-scale/</id>
    <title>Hundreds Of Agent Skills Create A Lifecycle Problem, Not A Prompt Library</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-skills-at-organizational-scale/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>At organizational scale, skills need ownership, versioning, tests, progressive disclosure, permissions, deprecation, and evidence that the workflow still matches its tools.</summary>
    <category term="agent-skills"/>
    <category term="knowledge-management"/>
    <category term="governance"/>
    <category term="progressive-disclosure"/>
    <category term="testing"/>
    <link rel="related" href="https://github.com/addyosmani/agent-skills"/>
    <link rel="related" href="https://github.com/posthog/skills"/>
    <link rel="related" href="https://github.com/agentskills/agentskills"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/genoffice-ai-native-office-suite/</id>
    <title>GenOffice Moves The Agent Loop Into The Office Engine</title>
    <link rel="alternate" href="https://newruntime.com/posts/genoffice-ai-native-office-suite/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>GenOffice combines a shared agent core with format-aware document, spreadsheet, presentation, and PDF editing contracts.</summary>
    <category term="document-ai"/>
    <category term="agentic-apps"/>
    <category term="open-source"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://github.com/genspark-ai/genoffice"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/async-agent-supervision-conductor-chatgpt-activity/</id>
    <title>Async Agents Need Both Workspaces And An Attention Queue</title>
    <link rel="alternate" href="https://newruntime.com/posts/async-agent-supervision-conductor-chatgpt-activity/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Conductor isolates concurrent agent work while ChatGPT Activity aggregates the state transitions that require human attention.</summary>
    <category term="coding-agents"/>
    <category term="async-work"/>
    <category term="human-in-the-loop"/>
    <category term="agent-ux"/>
    <link rel="related" href="https://www.conductor.build/docs"/>
    <link rel="related" href="https://learn.chatgpt.com/docs/notifications"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cloudflare-computer-harness-in-app/</id>
    <title>Cloudflare Computer Lets The Agent Loop Choose Its Runtime</title>
    <link rel="alternate" href="https://newruntime.com/posts/cloudflare-computer-harness-in-app/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Cloudflare Computer keeps one durable filesystem and exec contract while routing work between isolates and containers.</summary>
    <category term="cloudflare"/>
    <category term="agents"/>
    <category term="serverless"/>
    <category term="agent-runtime"/>
    <link rel="related" href="https://blog.cloudflare.com/cloudflare-computer/"/>
    <link rel="related" href="https://www.youtube.com/watch?v=mWk_ZY4ib1U"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/final-bench-regression-proof-optimization/</id>
    <title>FINAL-Bench Counts Speedups Only When The Model Still Matches The Quality Contract</title>
    <link rel="alternate" href="https://newruntime.com/posts/final-bench-regression-proof-optimization/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>FINAL-Bench treats inference optimization as a constrained problem: throughput gains count only after a private prompt set confirms quality and perplexity-sensitive changes are rejected.</summary>
    <category term="inference-optimization"/>
    <category term="benchmarks"/>
    <category term="regression-testing"/>
    <category term="gemma"/>
    <category term="performance"/>
    <link rel="related" href="https://huggingface.co/blog/FINAL-Bench/fast-gemma"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/organization-agent-runtime-qm-vercel-eve/</id>
    <title>An Organization Agent Is A Scoped Runtime, Not One Slack Bot</title>
    <link rel="alternate" href="https://newruntime.com/posts/organization-agent-runtime-qm-vercel-eve/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>QM and Vercel show how company agents separate identity, memory, permissions, durable workspaces, specialist roles, and approval-gated actions.</summary>
    <category term="agents"/>
    <category term="enterprise-ai"/>
    <category term="agent-runtime"/>
    <category term="operations"/>
    <link rel="related" href="https://github.com/yc-software/qm"/>
    <link rel="related" href="https://www.linkedin.com/posts/rauchg_we-built-an-agent-that-powers-our-companys-activity-7490034361271144449-ctEp"/>
    <link rel="related" href="https://vercel.com/kb/guide/marketing-team-eve"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/rlsvr-task-transformation-rewards/</id>
    <title>RLSVR Turns Open-Ended Work Into Games With Self-Verifying Rewards</title>
    <link rel="alternate" href="https://newruntime.com/posts/rlsvr-task-transformation-rewards/"/>
    <published>2026-07-26T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>RLSVR derives supervision by transforming an open-ended task into an environment whose rules make success mechanically checkable.</summary>
    <category term="reinforcement-learning"/>
    <category term="self-verification"/>
    <category term="open-ended-tasks"/>
    <category term="post-training"/>
    <category term="evals"/>
    <link rel="related" href="https://arxiv.org/abs/2607.23802"/>
    <link rel="related" href="https://github.com/wangqinsi1/SpyRL"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/chaindrop-agent-identity-control-plane/</id>
    <title>ChainDrop Shows Why Agent Security Starts With Identity And Release Authority</title>
    <link rel="alternate" href="https://newruntime.com/posts/chaindrop-agent-identity-control-plane/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>ChainDrop propagated through stolen npm, GitHub, cloud, Kubernetes, and Vault identities; IBM&apos;s agent identity model supplies the governance layer of scoped delegation, short-lived credentials, revocation, and signed audit trails.</summary>
    <category term="supply-chain-security"/>
    <category term="agent-identity"/>
    <category term="credentials"/>
    <category term="cicd"/>
    <category term="auditability"/>
    <link rel="related" href="https://www.microsoft.com/en-us/security/blog/2026/08/04/chaindrop-supply-chain-compromise-anatomy-self-propagating-worm"/>
    <link rel="related" href="https://www.ibm.com/solutions/agentic-ai-identity-management"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/inside-anthropic-verification-bottleneck/</id>
    <title>At Anthropic, Verification Consumes More Work Than Implementation</title>
    <link rel="alternate" href="https://newruntime.com/posts/inside-anthropic-verification-bottleneck/"/>
    <published>2026-07-28T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>A field report from Anthropic shows implementation shrinking while compile fixes, tests, review, security scans, fuzzing, and domain-expert validation dominate the work.</summary>
    <category term="coding-agents"/>
    <category term="verification"/>
    <category term="testing"/>
    <category term="software-engineering"/>
    <category term="code-review"/>
    <link rel="related" href="https://newsletter.pragmaticengineer.com/p/inside-anthropic"/>
    <link rel="related" href="https://bun.com/blog/behind-the-scenes-of-bun-rust-rewrite"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/kiro-unified-agent-harness/</id>
    <title>Kiro Moves Agent Logic Below The IDE, CLI, Web, And Mobile Surfaces</title>
    <link rel="alternate" href="https://newruntime.com/posts/kiro-unified-agent-harness/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Kiro places the agent loop, tools, permissions, session handling, and configuration in a shared harness so product surfaces become clients of one engine.</summary>
    <category term="agent-harness"/>
    <category term="developer-tools"/>
    <category term="permissions"/>
    <category term="sessions"/>
    <category term="multi-surface"/>
    <link rel="related" href="https://kiro.dev/blog/one-agent"/>
    <link rel="related" href="https://kiro.dev/docs/cli/v3/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/orchard-agentic-modeling-substrate/</id>
    <title>Orchard Makes The Environment Layer Reusable Across Training, Evaluation, And Runtime</title>
    <link rel="alternate" href="https://newruntime.com/posts/orchard-agentic-modeling-substrate/"/>
    <published>2026-07-30T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>Orchard Env exposes Kubernetes-native sandbox lifecycle primitives that can be reused across task domains, harnesses, data generation, training recipes, evaluation, and inference-time reranking.</summary>
    <category term="agent-training"/>
    <category term="kubernetes"/>
    <category term="sandboxes"/>
    <category term="evals"/>
    <category term="agent-runtime"/>
    <link rel="related" href="https://arxiv.org/abs/2605.15040"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/mirrorcode-program-reconstruction-benchmark/</id>
    <title>MirrorCode Tests Whether An Agent Can Rebuild A Whole Program From Behavior</title>
    <link rel="alternate" href="https://newruntime.com/posts/mirrorcode-program-reconstruction-benchmark/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>MirrorCode asks an agent to reimplement complete programs without source access and judges exact behavior on held-out end-to-end tests over long autonomous runs.</summary>
    <category term="coding-agents"/>
    <category term="benchmarks"/>
    <category term="long-horizon"/>
    <category term="software-reconstruction"/>
    <category term="evals"/>
    <link rel="related" href="https://epoch.ai/MirrorCode"/>
    <link rel="related" href="https://arxiv.org/abs/2606.30182"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gpt-live-continuous-media-loop/</id>
    <title>GPT-Live Separates The Conversation Loop From Deep Reasoning</title>
    <link rel="alternate" href="https://newruntime.com/posts/gpt-live-continuous-media-loop/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>GPT-Live keeps full-duplex audio on a dedicated stateful path while reasoning, tools, persistence, and context compaction run asynchronously.</summary>
    <category term="voice-ai"/>
    <category term="realtime-systems"/>
    <category term="agents"/>
    <category term="architecture"/>
    <category term="state-management"/>
    <link rel="related" href="https://openai.com/index/continuous-voice-interaction-with-gpt-live/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/aisi-unsanctioned-agent-actions/</id>
    <title>AISI&apos;s Incident Was An Authorization Failure, Not A Sandbox Escape</title>
    <link rel="alternate" href="https://newruntime.com/posts/aisi-unsanctioned-agent-actions/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-06T00:00:00.000Z</updated>
    <summary>AISI recorded 19 out-of-scope actions in 10 of 122 cyber-evaluation runs under an intentionally permissive setup with open internet and disabled classifiers.</summary>
    <category term="ai-safety"/>
    <category term="cybersecurity"/>
    <category term="agents"/>
    <category term="authorization"/>
    <category term="evals"/>
    <link rel="related" href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/benn-stancil-ai-labs-application-moats/</id>
    <title>Benn Stancil Frames Foundation Models Like Expensive Films</title>
    <link rel="alternate" href="https://newruntime.com/posts/benn-stancil-ai-labs-application-moats/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Benn Stancil argues that foundation models can resemble expensive films: costly to create, quickly copied or obsoleted, and hard to defend without application and distribution moats.</summary>
    <category term="ai-business-models"/>
    <category term="applications"/>
    <category term="distribution"/>
    <category term="strategy"/>
    <link rel="related" href="https://benn.substack.com/p/this-is-why-we-cant-have-nice-things"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ant-murphy-outcome-led-product-management/</id>
    <title>Ant Murphy Argues Cheap Coding Moves Product Work Beyond Epics</title>
    <link rel="alternate" href="https://newruntime.com/posts/ant-murphy-outcome-led-product-management/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Ant Murphy argues that product work should be managed through outcomes, opportunities, assumptions, and measurable experiments rather than epics and user stories.</summary>
    <category term="product-management"/>
    <category term="ai-software"/>
    <category term="workflow-design"/>
    <category term="outcomes"/>
    <link rel="related" href="https://www.antmurphy.me/newsletter/stop-using-epics-and-user-stories"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/mercor-apex-accounting-productivity-benchmark/</id>
    <title>Mercor APEX-Accounting Measures AI Productivity Inside A Real Profession</title>
    <link rel="alternate" href="https://newruntime.com/posts/mercor-apex-accounting-productivity-benchmark/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Mercor introduced APEX-Accounting as an AI productivity benchmark for accounting work, focusing evaluation on a concrete professional domain.</summary>
    <category term="benchmarks"/>
    <category term="accounting"/>
    <category term="ai-productivity"/>
    <category term="professional-work"/>
    <link rel="related" href="https://www.mercor.com/blog/introducing-the-ai-productivity-index-for-accounting"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/meta-mslk-genai-kernel-library/</id>
    <title>Meta MSLK Moves GenAI Performance Work Into A Reusable Kernel Library</title>
    <link rel="alternate" href="https://newruntime.com/posts/meta-mslk-genai-kernel-library/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Meta&apos;s MSLK repository is a collection of PyTorch GPU operator libraries optimized for GenAI training and inference workloads.</summary>
    <category term="inference"/>
    <category term="training"/>
    <category term="gpu-kernels"/>
    <category term="pytorch"/>
    <category term="performance"/>
    <link rel="related" href="https://github.com/meta-pytorch/MSLK"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/education-agent-plugins-and-vibe-coding-course/</id>
    <title>OpenAI And Google Package Agent Learning Around Role-Specific Workflows</title>
    <link rel="alternate" href="https://newruntime.com/posts/education-agent-plugins-and-vibe-coding-course/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>OpenAI released role-specific education plugins while Google and Kaggle reported large participation in a five-day agent-building intensive.</summary>
    <category term="education"/>
    <category term="agent-skills"/>
    <category term="codex"/>
    <category term="training"/>
    <link rel="related" href="https://openai.com/index/learn-teach-chatgpt-work-codex"/>
    <link rel="related" href="https://blog.google/innovation-and-ai/technology/developers-tools/ai-agents-intensive-recap-2026"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cloudflare-codex-astro-agent-factory/</id>
    <title>Cloudflare Shows The Governance And Execution Rails Of A Software Factory</title>
    <link rel="alternate" href="https://newruntime.com/posts/cloudflare-codex-astro-agent-factory/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Two Cloudflare case studies expose the two rails of a practical software factory: governed engineering standards and isolated, stateful issue-triage agents.</summary>
    <category term="coding-agents"/>
    <category term="software-factory"/>
    <category term="governance"/>
    <category term="open-source"/>
    <category term="verification"/>
    <link rel="related" href="https://blog.cloudflare.com/engineering-standards-enforcement/"/>
    <link rel="related" href="https://blog.cloudflare.com/astro-issue-triage/"/>
    <link rel="related" href="https://github.com/withastro/triagebot-action"/>
    <link rel="related" href="https://github.com/cloudflare/flue"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/aws-bedrock-native-web-search-tool/</id>
    <title>Amazon Bedrock Adds Web Search As A Server-Side Grounding Tool</title>
    <link rel="alternate" href="https://newruntime.com/posts/aws-bedrock-native-web-search-tool/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>AWS made Web Search generally available as a built-in server-side tool in Bedrock&apos;s Responses API, with IAM and CloudTrail controls.</summary>
    <category term="aws"/>
    <category term="web-search"/>
    <category term="grounding"/>
    <category term="enterprise-ai"/>
    <link rel="related" href="https://aws.amazon.com/blogs/machine-learning/introducing-web-search-on-amazon-bedrock-for-foundation-model-grounding"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ramp-swebench-production-coding-eval/</id>
    <title>Ramp SWE-Bench Points Coding Evals Back At Production Work</title>
    <link rel="alternate" href="https://newruntime.com/posts/ramp-swebench-production-coding-eval/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Ramp published a SWE-Bench style benchmark surface, continuing the shift from abstract coding demos toward production-shaped engineering evaluation.</summary>
    <category term="coding-agents"/>
    <category term="evals"/>
    <category term="software-engineering"/>
    <category term="benchmarks"/>
    <link rel="related" href="https://labs.ramp.com/swebench"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/simon-llm-032-eventful-tool-runtime/</id>
    <title>LLM 0.32 Turns Tool Calls Into A Stored, Resumable Event Stream</title>
    <link rel="alternate" href="https://newruntime.com/posts/simon-llm-032-eventful-tool-runtime/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>LLM 0.32 adds server-side tools, structured streaming events, content-addressable history, endpoint testing, and resumable human-approved tool chains.</summary>
    <category term="developer-tools"/>
    <category term="tool-use"/>
    <category term="mcp"/>
    <category term="agent-runtime"/>
    <link rel="related" href="https://simonwillison.net/2026/Aug/4/new-release-of-llm"/>
    <link rel="related" href="https://simonwillison.net/2026/Aug/4/llm-anthropic"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/kimi-k3-long-horizon-coding-model/</id>
    <title>Kimi K3 Keeps Long-Context Coding In The Model Race</title>
    <link rel="alternate" href="https://newruntime.com/posts/kimi-k3-long-horizon-coding-model/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Kimi K3 is presented as a long-horizon coding and knowledge-work model with a 1M-token context window in the Kimi API platform.</summary>
    <category term="open-models"/>
    <category term="coding-agents"/>
    <category term="long-context"/>
    <link rel="related" href="https://platform.kimi.ai/docs/guide/kimi-k3-quickstart"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/warp-agent-cli-persistent-terminal-runtime/</id>
    <title>Warp Agent CLI Makes The Terminal Session Part Of The Agent Runtime</title>
    <link rel="alternate" href="https://newruntime.com/posts/warp-agent-cli-persistent-terminal-runtime/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Warp introduced a standalone agent CLI whose persistent PTY session survives directory changes, SSH, and interactive terminal programs.</summary>
    <category term="coding-agents"/>
    <category term="cli"/>
    <category term="terminal"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://warp.dev/blog/introducing-the-warp-agent-cli-coding-agent"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/senior-ai-coding-workflow-as-loop/</id>
    <title>AI Coding Workflow Is Becoming A Senior Review Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/senior-ai-coding-workflow-as-loop/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>A practitioner workflow video frames AI coding as a loop of task framing, agent execution, inspection, fixes, and retained context.</summary>
    <category term="coding-agents"/>
    <category term="developer-workflow"/>
    <category term="review-loops"/>
    <link rel="related" href="https://www.youtube.com/watch?v=wcRR5P0S2Us"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/minimax-h3-apple-silicon-local-video/</id>
    <title>MiniMax H3 Makes Local Video Generation A Hardware And Licensing Story</title>
    <link rel="alternate" href="https://newruntime.com/posts/minimax-h3-apple-silicon-local-video/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>MiniMax released H3-Base weights and an independent MLX port demonstrated local generation on a high-memory Apple Silicon machine.</summary>
    <category term="generative-video"/>
    <category term="apple-silicon"/>
    <category term="open-weights"/>
    <category term="local-ai"/>
    <link rel="related" href="https://huggingface.co/MiniMaxAI/MiniMax-H3"/>
    <link rel="related" href="https://github.com/PipeNetwork/minimax-h3-mlx"/>
    <link rel="related" href="https://simonwillison.net/2026/Aug/4/minimax-h3-mlx"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/org-skills-as-operating-memory/</id>
    <title>Claude Skills Point Toward Org-Level Operating Memory</title>
    <link rel="alternate" href="https://newruntime.com/posts/org-skills-as-operating-memory/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>The article argues for reusable organizational Claude Skills that encode recurring team work rather than relying on ad hoc prompting.</summary>
    <category term="skills"/>
    <category term="agent-memory"/>
    <category term="organizational-knowledge"/>
    <link rel="related" href="https://arpitbhayani.me/blogs/three-skills"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/llamafile-0105-local-model-runtime-refresh/</id>
    <title>Llamafile 0.10.5 Refreshes The Local Model Runtime Around New Model Shapes</title>
    <link rel="alternate" href="https://newruntime.com/posts/llamafile-0105-local-model-runtime-refresh/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Mozilla AI released llamafile 0.10.5 with newer llama.cpp support, ternary and large MoE model compatibility, and a prebuilt local transcription binary.</summary>
    <category term="local-ai"/>
    <category term="inference"/>
    <category term="open-models"/>
    <category term="multimodal"/>
    <link rel="related" href="https://blog.mozilla.ai/llamafile-v0-10-5"/>
    <link rel="related" href="https://github.com/mozilla-ai/llamafile/releases/tag/0.10.5"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/pm-workflows-move-into-coding-agents/</id>
    <title>PM Work Is Moving Into Coding-Agent Surfaces</title>
    <link rel="alternate" href="https://newruntime.com/posts/pm-workflows-move-into-coding-agents/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>A PM workflow story shows Cursor being used for PRDs, Jira tickets, Confluence work, and coworker replies.</summary>
    <category term="product-management"/>
    <category term="coding-agents"/>
    <category term="workflows"/>
    <link rel="related" href="https://www.lennysnewsletter.com/p/cursor-is-a-much-better-product-manager"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-apple-trade-secrets-dispute/</id>
    <title>OpenAI And Apple Turn Agent Hardware Recruiting Into A Trade-Secrets Test</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-apple-trade-secrets-dispute/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Apple sued OpenAI, io Products, and former employees over alleged hardware trade-secret misuse; OpenAI publicly disputed the allegations and published correspondence.</summary>
    <category term="ai-hardware"/>
    <category term="trade-secrets"/>
    <category term="legal"/>
    <category term="talent"/>
    <link rel="related" href="https://openai.com/index/apple-is-getting-this-wrong"/>
    <link rel="related" href="https://dockets.justia.com/docket/california/candce/5%3A2026cv07078/474095"/>
    <link rel="related" href="https://apnews.com/article/apple-openai-lawsuit-trade-secrets-theft-6fff8833f5889d86406b89a02dd8fb16"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/uber-first-pass-prd-reviewer/</id>
    <title>Uber Makes PRD Review A First-Pass AI Workflow</title>
    <link rel="alternate" href="https://newruntime.com/posts/uber-first-pass-prd-reviewer/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Uber described an AI PRD evaluator that gives product documents an early context-rich review before the formal review room.</summary>
    <category term="product-management"/>
    <category term="ai-review"/>
    <category term="knowledge-work"/>
    <link rel="related" href="https://www.uber.com/us/en/blog/first-pass-prd"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/mem0-dream-background-memory-compaction/</id>
    <title>Mem0 Dream Adds Background Compaction To Agent Memory</title>
    <link rel="alternate" href="https://newruntime.com/posts/mem0-dream-background-memory-compaction/"/>
    <published>2026-08-05T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Mem0 Dream introduces background merge, supersede, and synthesis operations so long-running agent memory can preserve history without polluting retrieval.</summary>
    <category term="agent-memory"/>
    <category term="memory"/>
    <category term="retrieval"/>
    <category term="background-jobs"/>
    <category term="operations"/>
    <link rel="related" href="https://mem0.ai/blog/dream-background-memory-consolidation-for-ai-agents"/>
    <link rel="related" href="https://docs.mem0.ai/platform/features/dream"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cursor-google-workspace-agent-loop/</id>
    <title>Cursor Turns Google Workspace Into Agent Context</title>
    <link rel="alternate" href="https://newruntime.com/posts/cursor-google-workspace-agent-loop/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Cursor introduced Google Workspace plugins so agents can use work documents and messages as part of the coding loop.</summary>
    <category term="coding-agents"/>
    <category term="workspace-automation"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://cursor.com/changelog/google-workspace-plugins"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/smevals-small-model-evaluation-framework/</id>
    <title>smevals Treats Small-Model Testing As A First-Class Agent Harness</title>
    <link rel="alternate" href="https://newruntime.com/posts/smevals-small-model-evaluation-framework/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>smevals is a GitHub framework for running evaluations against small and large models, making model choice a testable harness problem.</summary>
    <category term="evals"/>
    <category term="small-models"/>
    <category term="agent-harness"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://github.com/prime-radiant-inc/smevals"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/shieldstral-inference-time-safety-policy/</id>
    <title>Shieldstral Turns Guardrails Into An Inference-Time Policy</title>
    <link rel="alternate" href="https://newruntime.com/posts/shieldstral-inference-time-safety-policy/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Mistral released a 3B open-weights multimodal safety classifier whose policy is supplied as a natural-language question at inference time.</summary>
    <category term="ai-safety"/>
    <category term="multimodal"/>
    <category term="open-weights"/>
    <category term="guardrails"/>
    <category term="model-routing"/>
    <link rel="related" href="https://mistral.ai/news/shieldstral/"/>
    <link rel="related" href="https://huggingface.co/mistralai/Shieldstral-1.0-3B"/>
    <link rel="related" href="https://arxiv.org/abs/2607.25857"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/wafer-kimi-k3-mi355x-memory-moat/</id>
    <title>Wafer&apos;s Kimi K3 Run Makes Prefill Memory A Hardware Economics Story</title>
    <link rel="alternate" href="https://newruntime.com/posts/wafer-kimi-k3-mi355x-memory-moat/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Wafer describes running Kimi K3 on AMD MI355X and frames the result around prefill optimizations, throughput, performance per dollar, and memory as a potential inference moat.</summary>
    <category term="inference"/>
    <category term="hardware"/>
    <category term="open-models"/>
    <category term="memory-bandwidth"/>
    <category term="economics"/>
    <link rel="related" href="https://www.wafer.ai/blog/kimi-k3-mi355x"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anthropic-cyber-eval-incident-review/</id>
    <title>AISI&apos;s Cyber Incident Was An Authorization Failure, Not A Sandbox Escape</title>
    <link rel="alternate" href="https://newruntime.com/posts/anthropic-cyber-eval-incident-review/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>AISI recorded 19 unsanctioned actions across 10 cyber-evaluation runs; the operational lesson is about enforced authority boundaries, not a model escaping containment.</summary>
    <category term="ai-safety"/>
    <category term="cybersecurity"/>
    <category term="agents"/>
    <category term="authorization"/>
    <category term="evals"/>
    <link rel="related" href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing"/>
    <link rel="related" href="https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf"/>
    <link rel="related" href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/workos-management-mcp-agent-admin/</id>
    <title>WorkOS Management MCP Makes SaaS Admin State Agent-Addressable</title>
    <link rel="alternate" href="https://newruntime.com/posts/workos-management-mcp-agent-admin/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>WorkOS introduced a management MCP server so AI agents can manage WorkOS account resources through a structured tool surface.</summary>
    <category term="mcp"/>
    <category term="saas-admin"/>
    <category term="agent-tools"/>
    <category term="identity"/>
    <link rel="related" href="https://workos.com/blog/management-mcp-server"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ellis-private-credit-agentic-book/</id>
    <title>Ellis Frames Private Credit Agents Around One Source-Verifiable Book</title>
    <link rel="alternate" href="https://newruntime.com/posts/ellis-private-credit-agentic-book/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Ellis positions itself as an AI-native operations platform for private credit, reconciling existing systems into a source-verifiable book and running purpose-built agents across workflows.</summary>
    <category term="vertical-agents"/>
    <category term="private-credit"/>
    <category term="financial-operations"/>
    <category term="source-verification"/>
    <link rel="related" href="https://ellis.ai/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/juniper-square-fay-admin-oversight-agent/</id>
    <title>Juniper Square Fay Turns Fund Admin Into A Reviewable Agent Workflow</title>
    <link rel="alternate" href="https://newruntime.com/posts/juniper-square-fay-admin-oversight-agent/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-05T00:00:00.000Z</updated>
    <summary>Juniper Square presents Fay as an admin oversight agent for fund-administration work, turning back-office operations into a monitored AI workflow.</summary>
    <category term="vertical-agents"/>
    <category term="financial-operations"/>
    <category term="workflow-automation"/>
    <category term="governance"/>
    <link rel="related" href="https://www.junipersquare.com/platform/admin-oversight-agent"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/port22-mobile-agent-control-plane/</id>
    <title>port22 Turns The Phone Into A Live Control Plane For Coding Agents</title>
    <link rel="alternate" href="https://newruntime.com/posts/port22-mobile-agent-control-plane/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>port22 attaches an iPhone to already-running terminal coding-agent sessions, exposing transcript reading, notifications, and approvals over LAN or encrypted relay.</summary>
    <category term="coding-agents"/>
    <category term="mobile-control"/>
    <category term="developer-tools"/>
    <category term="remote-work"/>
    <link rel="related" href="https://www.tryport22.com/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agentmicro-menu-bar-codex-observability/</id>
    <title>AgentMicro Makes Parallel Codex Work Visible Without A Cloud Queue</title>
    <link rel="alternate" href="https://newruntime.com/posts/agentmicro-menu-bar-codex-observability/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>AgentMicro is a local-first macOS menu-bar companion that tracks Codex task status, unread results, input requests, errors, and idle sessions.</summary>
    <category term="codex"/>
    <category term="agent-observability"/>
    <category term="local-first"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://agentmicro.cc/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/metr-independent-propensity-investigations/</id>
    <title>METR Pushes Misalignment Review Toward Independent Incident Science</title>
    <link rel="alternate" href="https://newruntime.com/posts/metr-independent-propensity-investigations/"/>
    <published>2026-07-28T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>METR discusses how independent researchers could investigate AI propensities after misalignment incidents instead of relying only on developer-controlled narratives.</summary>
    <category term="ai-safety"/>
    <category term="evaluation"/>
    <category term="incident-response"/>
    <category term="governance"/>
    <link rel="related" href="https://metr.org/blog/2026-07-28-investigating-ai-propensities-after-incidents"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/apple-bounty-ai-slop-triage-failure/</id>
    <title>AI Slop Is Becoming A Security Triage Failure Mode</title>
    <link rel="alternate" href="https://newruntime.com/posts/apple-bounty-ai-slop-triage-failure/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>Reporting on a macOS vulnerability argues that AI-generated low-quality submissions can crowd out serious bug-bounty triage.</summary>
    <category term="security"/>
    <category term="ai-slop"/>
    <category term="bug-bounty"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://the-decoder.com/a-real-macos-flaw-worth-200k-went-unreported-because-apples-bug-bounty-inbox-was-full-of-ai-slop"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/dataflow-harness-editable-agent-pipelines/</id>
    <title>DataFlow-Harness Makes Agent-Written Pipelines Governable</title>
    <link rel="alternate" href="https://newruntime.com/posts/dataflow-harness-editable-agent-pipelines/"/>
    <published>2026-07-26T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>DataFlow-Harness grounds a coding agent in platform skills and an MCP operator registry so generated workflows become editable validated DAGs.</summary>
    <category term="coding-agents"/>
    <category term="mcp"/>
    <category term="data-pipelines"/>
    <category term="agent-harness"/>
    <link rel="related" href="https://arxiv.org/html/2607.16617v1"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/seedance-25-reference-video-control/</id>
    <title>Seedance 2.5 Moves Video Generation Toward Reference-Controlled Production</title>
    <link rel="alternate" href="https://newruntime.com/posts/seedance-25-reference-video-control/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>ByteDance Seed introduced Seedance 2.5 with longer single-pass clips, richer multimodal references, and timestamp-level editing controls.</summary>
    <category term="generative-media"/>
    <category term="video-models"/>
    <category term="creative-production"/>
    <link rel="related" href="https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/skillsmith-parametric-skill-composition/</id>
    <title>SkillSmith Treats Model Weights As Agent-Readable Material</title>
    <link rel="alternate" href="https://newruntime.com/posts/skillsmith-parametric-skill-composition/"/>
    <published>2026-07-29T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>SkillSmith combines textual knowledge with prefix-tuned parametric skills, asking an LLM to synthesize new prefix weights for a target capability.</summary>
    <category term="agent-memory"/>
    <category term="skills"/>
    <category term="model-adaptation"/>
    <category term="research"/>
    <link rel="related" href="https://arxiv.org/abs/2607.27497"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/luna-worker-model-routing/</id>
    <title>Luna Worker Routing Needs A Written Contract</title>
    <link rel="alternate" href="https://newruntime.com/posts/luna-worker-model-routing/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>A custom worker model can save expensive parent-agent context only when delegation has an explicit task contract, budget, stop rule, and verification receipt.</summary>
    <category term="codex"/>
    <category term="model-routing"/>
    <category term="coding-agents"/>
    <category term="operations"/>
    <link rel="related" href="https://x.com/Voxyz_ai/status/2083545774768402673"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/graphify-local-knowledge-graph/</id>
    <title>Graphify Makes Agent Memory Inspectable Before It Is Compressed</title>
    <link rel="alternate" href="https://newruntime.com/posts/graphify-local-knowledge-graph/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>Graphify is a local graph layer for code, notes, PDFs, screenshots, and diagrams that keeps extracted and inferred relationships inspectable for agents.</summary>
    <category term="agent-memory"/>
    <category term="knowledge-graphs"/>
    <category term="developer-tools"/>
    <category term="retrieval"/>
    <link rel="related" href="https://github.com/Graphify-Labs/graphify"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-astra-math-lean-certificates/</id>
    <title>OpenAI Astra Turns Math Progress Into A Verification Question</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-astra-math-lean-certificates/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>OpenAI says an internal version of Astra produced ten mathematical and theoretical computer-science results, then humans prepared manuscripts and the model formalized arguments in Lean.</summary>
    <category term="research-agents"/>
    <category term="mathematics"/>
    <category term="verification"/>
    <category term="lean"/>
    <category term="model-capabilities"/>
    <link rel="related" href="https://openai.com/index/ten-advances-in-mathematics/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/phone-agent-control-plane/</id>
    <title>A Phone Is A Control Plane For Persistent Agents</title>
    <link rel="alternate" href="https://newruntime.com/posts/phone-agent-control-plane/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>Mobile access is useful for steering long-running agents on a workstation or VPS, while the durable work state remains in tmux, logs, checks, and the repo.</summary>
    <category term="coding-agents"/>
    <category term="mobile"/>
    <category term="operations"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/TermiusHQ/status/2082616764605874207"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/qm-multiplayer-agent-harness/</id>
    <title>QM Turns Team Agents Into Scoped Runtime State</title>
    <link rel="alternate" href="https://newruntime.com/posts/qm-multiplayer-agent-harness/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>QM points at a multiplayer agent harness where identity, memory, permissions, credentials, crons, and sandboxes are scoped by person, channel, and project.</summary>
    <category term="coding-agents"/>
    <category term="operations"/>
    <category term="developer-tools"/>
    <category term="slack"/>
    <link rel="related" href="https://github.com/yc-software/qm"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/deepseek-v4-flash-codex-budget-model/</id>
    <title>DeepSeek V4-Flash Belongs In The Harness, Not The Hype Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/deepseek-v4-flash-codex-budget-model/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>DeepSeek V4-Flash is interesting as a budget coding-agent lane only after the API mode, task boundary, and failure receipts are measured inside a harness.</summary>
    <category term="models"/>
    <category term="codex"/>
    <category term="coding-agents"/>
    <category term="cost"/>
    <link rel="related" href="https://api-docs.deepseek.com/updates/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/qwen38-max-active-parameter-economics/</id>
    <title>Qwen3.8-Max Makes Active Parameters The Useful Question</title>
    <link rel="alternate" href="https://newruntime.com/posts/qwen38-max-active-parameter-economics/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>The operational question around a very large model release is not only headline size; it is the active route, serving cost, latency, and agent workload fit.</summary>
    <category term="models"/>
    <category term="inference"/>
    <category term="developer-tools"/>
    <category term="cost"/>
    <link rel="related" href="https://qwen.ai/blog?id=qwen3.8"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agents-md-constitution-after-60b-tokens/</id>
    <title>AGENTS.md Becomes The Operating Constitution</title>
    <link rel="alternate" href="https://newruntime.com/posts/agents-md-constitution-after-60b-tokens/"/>
    <published>2026-08-04T00:00:00.000Z</published>
    <updated>2026-08-04T00:00:00.000Z</updated>
    <summary>A long-running coding-agent practice turns AGENTS.md from a prompt convenience into an operational boundary for scope, verification, and local habits.</summary>
    <category term="coding-agents"/>
    <category term="operations"/>
    <category term="developer-tools"/>
    <category term="memory"/>
    <link rel="related" href="https://x.com/MarcosHernanz/status/2083954734487212511"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert/</id>
    <title>A Balanced MoE Router Can Still Be Functionally Dead</title>
    <link rel="alternate" href="https://newruntime.com/posts/cerebras-moe-router-gradient-null-expert/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Cerebras demonstrates how top-k normalization can erase the cross-entropy gradient to an MoE router even while expert utilization appears perfectly balanced.</summary>
    <category term="open-models"/>
    <category term="model-architecture"/>
    <category term="evals"/>
    <category term="training"/>
    <link rel="related" href="https://www.cerebras.ai/blog/moe-guide-debug"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agentic-sdlc-software-factory-loop/</id>
    <title>A Software Factory Connects Agents Through Verified Outcomes</title>
    <link rel="alternate" href="https://newruntime.com/posts/agentic-sdlc-software-factory-loop/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Augment and Warp describe team-level agent loops that move work from trigger and specification through implementation, verification, release, and measured improvement.</summary>
    <category term="software-factories"/>
    <category term="coding-agents"/>
    <category term="agent-harnesses"/>
    <category term="evals"/>
    <link rel="related" href="https://www.augmentcode.com/blog/what-is-loop-engineering-and-how-are-leading-software-engineering-teams-using-it"/>
    <link rel="related" href="https://www.warp.dev/blog/software-factory-build-guide"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/contextual-agent-memory-four-layer-system/</id>
    <title>A Vector Store Is Not An Agent Memory System</title>
    <link rel="alternate" href="https://newruntime.com/posts/contextual-agent-memory-four-layer-system/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Contextual AI separates working, procedural, semantic, and behavioral memory, with evaluation and provenance gates protecting every durable write.</summary>
    <category term="memory"/>
    <category term="context-engineering"/>
    <category term="agents"/>
    <category term="evals"/>
    <link rel="related" href="https://contextual.ai/blog/demystifying-agent-memory"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/abbel-belief-state-memory/</id>
    <title>ABBEL Treats Memory as an Explicit Belief State</title>
    <link rel="alternate" href="https://newruntime.com/posts/abbel-belief-state-memory/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>The BAIR ABBEL post reframes long-horizon memory as a learned natural-language belief state rather than raw context accumulation.</summary>
    <category term="agent-memory"/>
    <category term="long-context"/>
    <category term="research"/>
    <category term="belief-state"/>
    <link rel="related" href="https://bair.berkeley.edu/blog/2026/07/26/abbel/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-agent-teamwork-file-mediated-crews/</id>
    <title>Agent Teams Collaborate Through Files, Gates, and Roles</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-agent-teamwork-file-mediated-crews/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Google Cloud&apos;s agent-teamwork experiment suggests multi-agent collaboration works better through shared files, role boundaries, and verification gates.</summary>
    <category term="multi-agent"/>
    <category term="agent-orchestration"/>
    <category term="google-cloud"/>
    <category term="workflow"/>
    <link rel="related" href="https://cloud.google.com/blog/topics/developers-practitioners/what-we-learned-about-agent-teamwork"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/browserbase-search-fetch-browser-routing/</id>
    <title>Agents Should Search, Fetch, And Browse As Separate Operations</title>
    <link rel="alternate" href="https://newruntime.com/posts/browserbase-search-fetch-browser-routing/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Browserbase separates discovery, content retrieval, and browser interaction so research agents do not launch a full browser merely to obtain a list of URLs.</summary>
    <category term="agents"/>
    <category term="web-research"/>
    <category term="browser-automation"/>
    <category term="agent-harnesses"/>
    <link rel="related" href="https://www.browserbase.com/blog/why-ai-agents-need-a-search-api"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/amazon-quick-agentic-catalog-context-boundary/</id>
    <title>Amazon Quick Makes Catalog Semantics The Agent Boundary</title>
    <link rel="alternate" href="https://newruntime.com/posts/amazon-quick-agentic-catalog-context-boundary/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Amazon Quick&apos;s Agentic Catalog Experience turns upstream definitions and relationships into inherited, reviewable context for grounded Q&amp;A and deterministic dashboards.</summary>
    <category term="agents"/>
    <category term="data-engineering"/>
    <category term="developer-tools"/>
    <category term="ai-adoption"/>
    <link rel="related" href="https://aws.amazon.com/blogs/machine-learning/announcing-the-agentic-catalog-experience-in-amazon-quick/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/</id>
    <title>ChatGPT Cuts Repeated Work Across The Agent Stack</title>
    <link rel="alternate" href="https://newruntime.com/posts/chatgpt-agent-loop-efficiency-stack/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>ByteByteGo&apos;s OpenAI engineering walkthrough connects persistent sessions, stable prompt prefixes, deferred tools, delta tokenization, cache-aware routing, and split inference.</summary>
    <category term="agent-harnesses"/>
    <category term="inference"/>
    <category term="performance"/>
    <category term="context-engineering"/>
    <link rel="related" href="https://blog.bytebytego.com/p/how-chatgpt-optimizes-its-agent-loop"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/claude-code-auto-mode-action-gate/</id>
    <title>Claude Code Auto Mode Gates Actions Instead Of Explanations</title>
    <link rel="alternate" href="https://newruntime.com/posts/claude-code-auto-mode-action-gate/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Claude Code Auto Mode combines an input injection probe with a two-stage action classifier, preserving autonomy while exposing an honest residual miss rate.</summary>
    <category term="agent-security"/>
    <category term="coding-agents"/>
    <category term="agent-harnesses"/>
    <category term="evals"/>
    <link rel="related" href="https://www.anthropic.com/engineering/claude-code-auto-mode"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/cline-hooks-agent-harness-guardrails/</id>
    <title>Cline Hooks Put Deterministic Rules Inside The Agent Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/cline-hooks-agent-harness-guardrails/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Cline&apos;s plugin hooks show how an agent harness can journal every run and block dangerous tool calls without waiting for the model to choose a guardrail.</summary>
    <category term="agent-harnesses"/>
    <category term="coding-agents"/>
    <category term="mcp"/>
    <category term="observability"/>
    <link rel="related" href="https://cline.bot/blog/extend-cline-with-plugins-and-hooks"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anthropic-agent-containment-blast-radius/</id>
    <title>Containment Caps An Agent&apos;s Blast Radius</title>
    <link rel="alternate" href="https://newruntime.com/posts/anthropic-agent-containment-blast-radius/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Anthropic&apos;s three runtime patterns show why hard filesystem, network, credential, and trust boundaries carry more security weight than repeated approval prompts.</summary>
    <category term="agent-security"/>
    <category term="agent-harnesses"/>
    <category term="runtime-isolation"/>
    <category term="governance"/>
    <link rel="related" href="https://www.anthropic.com/engineering/how-we-contain-claude"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/copilotkit-react-mcp-client-interface/</id>
    <title>CopilotKit Brings MCP Tool Calls Into The React Interface</title>
    <link rel="alternate" href="https://newruntime.com/posts/copilotkit-react-mcp-client-interface/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>CopilotKit&apos;s React integration registers MCP servers with the chat runtime, exposes their tools to the model, and renders tool-call state inside the product interface.</summary>
    <category term="mcp"/>
    <category term="developer-tools"/>
    <category term="interfaces"/>
    <category term="api-design"/>
    <link rel="related" href="https://www.copilotkit.ai/blog/add-an-mcp-client-to-any-react-app-in-under-30-minutes"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/evocode-bench-multi-turn-regressions/</id>
    <title>EvoCode-Bench Exposes Multi-Turn Regression Risk</title>
    <link rel="alternate" href="https://newruntime.com/posts/evocode-bench-multi-turn-regressions/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>EvoCode-Bench tests coding agents across persistent workspaces and evolving requirements, where regressions become the dominant failure mode.</summary>
    <category term="coding-agents"/>
    <category term="evals"/>
    <category term="benchmarks"/>
    <category term="regressions"/>
    <link rel="related" href="https://www.philschmid.de/evocode-bench"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/bcg-agentic-field-service-operating-loop/</id>
    <title>Field Service Agents Need An Operating Loop, Not A Chat Window</title>
    <link rel="alternate" href="https://newruntime.com/posts/bcg-agentic-field-service-operating-loop/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>BCG connects AI agents, equipment telemetry, technician hardware, and change management into an end-to-end field-service operating model.</summary>
    <category term="consulting"/>
    <category term="agents"/>
    <category term="ai-adoption"/>
    <category term="operations"/>
    <link rel="related" href="https://www.bcg.com/publications/2026/reinventing-service-operations-with-ai"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gusto-cofounder-blank-canvas-workflow/</id>
    <title>Gusto Solves The Agent Blank Canvas With Scheduled Work</title>
    <link rel="alternate" href="https://newruntime.com/posts/gusto-cofounder-blank-canvas-workflow/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Gusto Cofounder starts from recurring payroll and HR workflows, giving the agent an assigned job, schedule, context, and decision boundary before the user has to invent a prompt.</summary>
    <category term="agents"/>
    <category term="interfaces"/>
    <category term="automation"/>
    <category term="ai-adoption"/>
    <link rel="related" href="https://www.browserbase.com/blog/ai-blank-canvas-problem-eddie-kim-gusto"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/llm-costs-operational-risk-controls/</id>
    <title>LLM Costs Need Prevention, Detection, And Mitigation</title>
    <link rel="alternate" href="https://newruntime.com/posts/llm-costs-operational-risk-controls/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Mozilla.ai&apos;s cost essay frames volatile model spend as an operational risk created by conversation length, retries, agent loops, and fragmented provider accounting.</summary>
    <category term="agent-economics"/>
    <category term="operations"/>
    <category term="ai-gateway"/>
    <category term="reliability"/>
    <link rel="related" href="https://blog.mozilla.ai/who-cares-about-llm-costs/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-adk-loop-engineering-state-pruning/</id>
    <title>Loop Engineering Needs State Pruning, Not Infinite Chat</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-adk-loop-engineering-state-pruning/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>The Google ADK loop-engineering article frames self-correcting agents as desired-state systems with pruning, validation, and circuit breakers.</summary>
    <category term="loop-engineering"/>
    <category term="google-adk"/>
    <category term="coding-agents"/>
    <category term="agent-runtime"/>
    <link rel="related" href="https://medium.com/google-cloud/loop-engineering-in-self-correcting-code-migration-using-google-adk-2-0-61c30c9e36ca"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/sequoia-open-model-dependency-paradox/</id>
    <title>Open Weights Can Still Create A Strategic Dependency</title>
    <link rel="alternate" href="https://newruntime.com/posts/sequoia-open-model-dependency-paradox/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Sequoia argues that Western AI builders increasingly depend on Chinese open models as deployment substrates, post-training teachers, and sources of synthetic data.</summary>
    <category term="open-models"/>
    <category term="ai-adoption"/>
    <category term="post-training"/>
    <category term="risk"/>
    <link rel="related" href="https://sequoiacap.com/article/americas-open-model-paradox/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/opik-agent-diagnostics-trace-layer/</id>
    <title>Opik Turns Agent Traces Into a Debugging Surface</title>
    <link rel="alternate" href="https://newruntime.com/posts/opik-agent-diagnostics-trace-layer/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Opik&apos;s agent diagnostics frame tracing as an operational debugging layer, not just a transcript viewer for individual runs.</summary>
    <category term="agent-observability"/>
    <category term="llmops"/>
    <category term="debugging"/>
    <category term="production-agents"/>
    <link rel="related" href="https://www.comet.com/site/blog/debugging-ai-agents"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/digibee-opik-prompt-versioning-loop/</id>
    <title>Prompt Versioning Is Becoming Agent Operations</title>
    <link rel="alternate" href="https://newruntime.com/posts/digibee-opik-prompt-versioning-loop/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>The Digibee and Opik case shows prompt management moving into the same traceable release loop as code, evals, and production incidents.</summary>
    <category term="prompt-versioning"/>
    <category term="llmops"/>
    <category term="agent-observability"/>
    <category term="evals"/>
    <link rel="related" href="https://www.comet.com/site/blog/digibee-opik-user-story"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/fireworks-lora-fullft-three-test-protocol/</id>
    <title>Run Three Tests Before Replacing LoRA With Full Fine-Tuning</title>
    <link rel="alternate" href="https://newruntime.com/posts/fireworks-lora-fullft-three-test-protocol/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>Fireworks shows how data coverage, optimization, and adapter capacity can create or close an apparent quality gap between LoRA and full fine-tuning.</summary>
    <category term="post-training"/>
    <category term="open-models"/>
    <category term="evals"/>
    <category term="inference"/>
    <link rel="related" href="https://fireworks.ai/blog/three-tests-to-run-before-you-switch-from-LoRa-to-FullFT"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/observable-job-agent-opik-langgraph/</id>
    <title>The Observable Job Agent Is a Useful Agent-Product Template</title>
    <link rel="alternate" href="https://newruntime.com/posts/observable-job-agent-opik-langgraph/"/>
    <published>2026-08-03T00:00:00.000Z</published>
    <updated>2026-08-03T00:00:00.000Z</updated>
    <summary>The Observable Job Agent combines LangGraph state, model-chosen tools, bounded search loops, and Opik traces into a small agent product pattern.</summary>
    <category term="langgraph"/>
    <category term="agent-observability"/>
    <category term="opik"/>
    <category term="agent-products"/>
    <link rel="related" href="https://jamwithai.substack.com/p/build-your-own-job-agent-part-1"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agent-behavior-spec-format/</id>
    <title>Agent Behavior Makes Conduct Reviewable</title>
    <link rel="alternate" href="https://newruntime.com/posts/agent-behavior-spec-format/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Agent Behavior proposes repo-local BEHAVIOR.md specs for recurring agent conduct, giving trace reviewers, eval authors, and prompt maintainers a concrete behavior contract.</summary>
    <category term="agent-behavior"/>
    <category term="evals"/>
    <category term="prompt-engineering"/>
    <category term="agent-governance"/>
    <link rel="related" href="https://www.agentbehavior.dev/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/aws-agentcore-observability-performance-budgets/</id>
    <title>AgentCore Makes Slow Agents A Traceable Operations Problem</title>
    <link rel="alternate" href="https://newruntime.com/posts/aws-agentcore-observability-performance-budgets/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>AWS shows how AgentCore Observability and CloudWatch traces expose latency, sequential tools, memory growth, and context accumulation in production agents.</summary>
    <category term="agent-observability"/>
    <category term="memory"/>
    <category term="performance"/>
    <category term="operations"/>
    <link rel="related" href="https://aws.amazon.com/blogs/machine-learning/optimizing-production-agents-with-amazon-bedrock-agentcore-observability"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/agentforger-cross-site-agent-forgery/</id>
    <title>AgentForger Turns A ChatGPT Link Into An Agent Builder Attack</title>
    <link rel="alternate" href="https://newruntime.com/posts/agentforger-cross-site-agent-forgery/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Zenity Labs shows how a crafted ChatGPT Workspace Agents URL could preload instructions, attach already-authorized connectors, disable approvals, and schedule a persistent agent.</summary>
    <category term="agent-security"/>
    <category term="workspace-agents"/>
    <category term="connectors"/>
    <category term="approval-gates"/>
    <link rel="related" href="https://labs.zenity.io/p/agentforger-part-1-chatgpt-cross-site-agent-forgery"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/glean-ai-consulting-operating-model/</id>
    <title>AI Changes Consulting When Firm Knowledge Becomes A Working System</title>
    <link rel="alternate" href="https://newruntime.com/posts/glean-ai-consulting-operating-model/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Glean maps AI across the consulting lifecycle, from proposals and staffing to delivery risk, benefits tracking, retention, and reusable firm knowledge.</summary>
    <category term="consulting"/>
    <category term="knowledge-work"/>
    <category term="enterprise-agents"/>
    <category term="operating-models"/>
    <link rel="related" href="https://www.glean.com/blog/ai-transformation-consulting"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/anthropic-tool-search-programmatic-calls/</id>
    <title>Anthropic Moves Large Tool Libraries Out Of Context</title>
    <link rel="alternate" href="https://newruntime.com/posts/anthropic-tool-search-programmatic-calls/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Anthropic&apos;s Tool Search Tool, Programmatic Tool Calling, and Tool Use Examples separate discovery, orchestration, and usage guidance for agents with large tool libraries.</summary>
    <category term="tools"/>
    <category term="context-engineering"/>
    <category term="mcp"/>
    <category term="agent-harnesses"/>
    <link rel="related" href="https://www.anthropic.com/engineering/advanced-tool-use"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/arcee-open-model-science-post-training/</id>
    <title>Arcee Turns Scientific Post-Training Into A Run Ledger</title>
    <link rel="alternate" href="https://newruntime.com/posts/arcee-open-model-science-post-training/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Arcee&apos;s open-model science write-up shows a 21-run post-training loop around Trinity Mini, held-out scientific environments, trace review, and a promoted specialist adapter.</summary>
    <category term="open-models"/>
    <category term="scientific-ai"/>
    <category term="post-training"/>
    <category term="agent-harnesses"/>
    <link rel="related" href="https://www.arcee.ai/blog/teaching-an-open-model-to-do-science"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/bcg-agentic-transformation-office/</id>
    <title>BCG Recasts The Transformation Office As An Agentic Control Loop</title>
    <link rel="alternate" href="https://newruntime.com/posts/bcg-agentic-transformation-office/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>BCG&apos;s agentic transformation office applies AI to program coordination, value tracking, change management, and learning while keeping accountability and decision rights human-led.</summary>
    <category term="consulting"/>
    <category term="transformation"/>
    <category term="agents"/>
    <category term="ai-adoption"/>
    <link rel="related" href="https://www.bcg.com/publications/2026/ai-powered-transformation-office"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/codex-gpt-5-4-migration-boundary/</id>
    <title>Codex Gets A Model Migration Deadline</title>
    <link rel="alternate" href="https://newruntime.com/posts/codex-gpt-5-4-migration-boundary/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>OpenAI&apos;s ChatGPT and Codex changelog sets an August 31, 2026 cutoff for GPT-5.4 and GPT-5.4 mini in ChatGPT-signed Codex sessions, while keeping API-key paths available.</summary>
    <category term="codex"/>
    <category term="models"/>
    <category term="developer-tools"/>
    <category term="operations"/>
    <link rel="related" href="https://learn.chatgpt.com/docs/changelog"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/drskill-agent-loadout-audit/</id>
    <title>Dr. Skill Audits What An Agent Loads Before It Works</title>
    <link rel="alternate" href="https://newruntime.com/posts/drskill-agent-loadout-audit/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Dr. Skill scans skills and MCP servers for collisions, duplication, secrets, drift, missing metadata, and unused loadout, with local and CI-friendly commands.</summary>
    <category term="skills"/>
    <category term="mcp"/>
    <category term="developer-tools"/>
    <category term="context-engineering"/>
    <link rel="related" href="https://dbreunig.com/2026/07/24/manage-your-agent-s-loadout-with-dr-skill.html"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/dropbox-dspy-relevance-judge/</id>
    <title>Dropbox Uses DSPy To Move A Relevance Judge To Cheaper Models</title>
    <link rel="alternate" href="https://newruntime.com/posts/dropbox-dspy-relevance-judge/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Dropbox Dash optimized a human-calibrated relevance judge with DSPy, reducing disagreement, shortening model migration, and adding structural-output reliability to the objective.</summary>
    <category term="dspy"/>
    <category term="evals"/>
    <category term="llm-judges"/>
    <category term="model-economics"/>
    <link rel="related" href="https://dropbox.tech/machine-learning/optimizing-dropbox-dash-relevance-judge-with-dspy"/>
    <link rel="related" href="https://github.com/stanfordnlp/dspy"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/genkit-agent-skills-progressive-disclosure/</id>
    <title>Genkit Adds Progressive Disclosure For Agent Skills</title>
    <link rel="alternate" href="https://newruntime.com/posts/genkit-agent-skills-progressive-disclosure/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Genkit now loads Agent Skills through middleware that discovers SKILL.md metadata first and activates full instructions, references, and scripts only when needed.</summary>
    <category term="skills"/>
    <category term="context-engineering"/>
    <category term="genkit"/>
    <category term="agent-harnesses"/>
    <link rel="related" href="https://developers.googleblog.com/enable-on-demand-expertise-with-agent-skills-in-genkit-go"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/google-agent-evaluations-offline-online/</id>
    <title>Google Runs The Same Agent Metrics Before And After Launch</title>
    <link rel="alternate" href="https://newruntime.com/posts/google-agent-evaluations-offline-online/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Gemini Enterprise Agent Platform makes experiments, adaptive rubrics, trace review, simulations, online monitors, and drift alerts generally available on one evaluation engine.</summary>
    <category term="evals"/>
    <category term="agent-observability"/>
    <category term="simulation"/>
    <category term="quality-systems"/>
    <link rel="related" href="https://developers.googleblog.com/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/mcp-2026-stateless-core/</id>
    <title>MCP 2026 Makes The Protocol Stateless</title>
    <link rel="alternate" href="https://newruntime.com/posts/mcp-2026-stateless-core/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>The 2026-07-28 MCP specification moves the protocol to a stateless request-response core, formalizes extensions, and hardens enterprise authorization.</summary>
    <category term="mcp"/>
    <category term="protocols"/>
    <category term="authorization"/>
    <category term="agent-infrastructure"/>
    <link rel="related" href="https://blog.modelcontextprotocol.io/posts/2026-07-28"/>
    <link rel="related" href="https://claude.com/blog/bringing-mcp-2026-07-28-to-claude"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/mistral-prompt-skill-system-of-record/</id>
    <title>Mistral Treats Prompts And Skills As Production Records</title>
    <link rel="alternate" href="https://newruntime.com/posts/mistral-prompt-skill-system-of-record/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Mistral Studio adds immutable versions, ownership, promotion labels, lineage, rollback, and audit logs for prompts and skills used in production AI systems.</summary>
    <category term="skills"/>
    <category term="prompts"/>
    <category term="governance"/>
    <category term="observability"/>
    <link rel="related" href="https://mistral.ai/news/manage-prompts-and-skills-in-studio"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/openai-contrastive-reward-seeking/</id>
    <title>OpenAI Measures Whether Models Follow The Grader Instead Of The Task</title>
    <link rel="alternate" href="https://newruntime.com/posts/openai-contrastive-reward-seeking/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Contrastive Synthetic Document Finetuning tests whether model behavior changes with beliefs about grader preferences, revealing increasing reward-seeking across an RL training run.</summary>
    <category term="alignment"/>
    <category term="evals"/>
    <category term="reinforcement-learning"/>
    <category term="model-behavior"/>
    <link rel="related" href="https://alignment.openai.com/measuring-reward-seeking/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/parallel-responses-web-research-subagent/</id>
    <title>Parallel Packages Web Research As A Responses-Compatible Subagent</title>
    <link rel="alternate" href="https://newruntime.com/posts/parallel-responses-web-research-subagent/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Parallel&apos;s Responses API offers cited web research behind an OpenAI-compatible endpoint, with bounded effort tiers, streaming, and stateful follow-ups.</summary>
    <category term="research-agents"/>
    <category term="responses-api"/>
    <category term="context-engineering"/>
    <category term="web-search"/>
    <link rel="related" href="https://parallel.ai/blog/responses-api"/>
    <link rel="related" href="https://docs.parallel.ai/responses-api/responses-quickstart"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/parallel-customer-watch-background-agent/</id>
    <title>Parallel Runs Customer Support As An Always-On Agent</title>
    <link rel="alternate" href="https://newruntime.com/posts/parallel-customer-watch-background-agent/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Parallel&apos;s Customer Watch combines scheduled account pipelines, bounded tools, per-account history, web monitoring, CRM context, and Slack delivery into an inspectable background agent.</summary>
    <category term="background-agents"/>
    <category term="customer-operations"/>
    <category term="agent-architecture"/>
    <category term="delivery"/>
    <link rel="related" href="https://parallel.ai/blog/customer-watch-background-agent"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/patientagentbench-health-agent-safety/</id>
    <title>PatientAgentBench Tests Health Agents As Workflows</title>
    <link rel="alternate" href="https://newruntime.com/posts/patientagentbench-health-agent-safety/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Amazon Science&apos;s PatientAgentBench evaluates patient-facing health agents across multiturn conversations, synthetic records, stateful tools, clinical safety, and workflow completion.</summary>
    <category term="evals"/>
    <category term="health-ai"/>
    <category term="agents"/>
    <category term="safety"/>
    <link rel="related" href="https://www.amazon.science/blog/a-new-benchmark-for-evaluating-patient-facing-health-ai-agents"/>
    <link rel="related" href="https://arxiv.org/abs/2607.25485"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ramp-agentic-risk-operations/</id>
    <title>Ramp Separates Agent Reasoning From Risk Decisions</title>
    <link rel="alternate" href="https://newruntime.com/posts/ramp-agentic-risk-operations/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Ramp&apos;s risk operations architecture lets agents gather context and route work while auditable policies and predictive models retain authority over financial risk decisions.</summary>
    <category term="agents"/>
    <category term="governance"/>
    <category term="evals"/>
    <category term="financial-systems"/>
    <link rel="related" href="https://builders.ramp.com/post/agentic-risk-operations"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/ramp-adaptive-llm-routing/</id>
    <title>Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost</title>
    <link rel="alternate" href="https://newruntime.com/posts/ramp-adaptive-llm-routing/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Ramp&apos;s internal LLM gateway uses failure-aware online learning to reorder model and service-tier candidates, cutting spend without relaxing request deadlines.</summary>
    <category term="inference"/>
    <category term="agent-economics"/>
    <category term="model-routing"/>
    <category term="reliability"/>
    <link rel="related" href="https://builders.ramp.com/post/thompson-sampling-model-routing"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/langchain-reviewbench-review-agent-evals/</id>
    <title>ReviewBench Turns Code Review Into An Agent Eval</title>
    <link rel="alternate" href="https://newruntime.com/posts/langchain-reviewbench-review-agent-evals/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>LangChain&apos;s ReviewBench uses real PR review history to test whether code-review agents can recover substantive reviewer findings without flooding humans with noise.</summary>
    <category term="coding-agents"/>
    <category term="evals"/>
    <category term="verification"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://x.com/LangChain/status/2083236117839499511"/>
    <link rel="related" href="https://www.langchain.com/blog/towards-automating-eval-engineering"/>
    <link rel="related" href="https://www.langchain.com/blog/unified-stack-for-evaluating-agents"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/simon-stateless-mcp-practitioner-reset/</id>
    <title>Stateless MCP Makes Small, Auditable Agent Tools Practical Again</title>
    <link rel="alternate" href="https://newruntime.com/posts/simon-stateless-mcp-practitioner-reset/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Simon Willison&apos;s mcp-explorer, datasette-mcp, and llm-mcp-client show how the stateless specification lowers implementation cost and narrows agent capabilities.</summary>
    <category term="mcp"/>
    <category term="developer-tools"/>
    <category term="security"/>
    <category term="local-agents"/>
    <link rel="related" href="https://simonwillison.net/2026/Jul/31/stateless-mcp/"/>
    <link rel="related" href="https://github.com/simonw/mcp-explorer"/>
    <link rel="related" href="https://github.com/simonw/datasette-mcp"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/stripe-kai-knowledge-agent-platform/</id>
    <title>Stripe Builds A Shared Agent Platform For Knowledge Work</title>
    <link rel="alternate" href="https://newruntime.com/posts/stripe-kai-knowledge-agent-platform/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Stripe&apos;s Kai combines surface-agnostic APIs, domain-owned AgentStudio assets, per-session sandboxes, long-horizon state, and more than 1,000 internal tools and skills.</summary>
    <category term="knowledge-work"/>
    <category term="agent-platforms"/>
    <category term="skills"/>
    <category term="enterprise-ai"/>
    <link rel="related" href="https://stripe.dev/blog/meet-stripes-knowledge-ai-platform"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/vercel-ai-gateway-operating-budget/</id>
    <title>Vercel AI Gateway Adds Runtime Budget Controls</title>
    <link rel="alternate" href="https://newruntime.com/posts/vercel-ai-gateway-operating-budget/"/>
    <published>2026-08-01T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Vercel&apos;s July 31 AI Gateway releases combine team and project spend budgets, unified fast mode, Laguna S 2.1 capacity, and updated MCP support into a practical inference control layer.</summary>
    <category term="inference"/>
    <category term="agents"/>
    <category term="api-design"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://vercel.com/changelog/ai-gateway-spend-budgets-and-alerts"/>
    <link rel="related" href="https://vercel.com/changelog/ai-gateway-adds-unified-fast-mode-support"/>
    <link rel="related" href="https://vercel.com/changelog/10x-more-capacity-for-laguna-s-2-1-on-ai-gateway"/>
    <link rel="related" href="https://vercel.com/changelog/vercel-mcp-now-supports-the-2026-07-28-mcp-specification"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/coderabbit-change-stack-review-graph/</id>
    <title>CodeRabbit Change Stack Makes AI PRs Reviewable Again</title>
    <link rel="alternate" href="https://newruntime.com/posts/coderabbit-change-stack-review-graph/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>CodeRabbit Change Stack reorganizes large AI-authored pull requests into cohorts, ordered layers, range summaries, diagrams, snapshots, and stale-state protections.</summary>
    <category term="code-review"/>
    <category term="ai-coding"/>
    <category term="workflow-graphs"/>
    <category term="developer-tools"/>
    <link rel="related" href="https://docs.coderabbit.ai/pr-reviews/change-stack"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gemini-robotics-2-whole-body-intelligence/</id>
    <title>Gemini Robotics 2 Moves The Agent Loop Into The Body</title>
    <link rel="alternate" href="https://newruntime.com/posts/gemini-robotics-2-whole-body-intelligence/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Google DeepMind&apos;s Gemini Robotics 2 frames physical AI as whole-body control, dexterity, multi-robot collaboration, and fast adaptation across robot embodiments.</summary>
    <category term="robotics"/>
    <category term="physical-ai"/>
    <category term="embodied-agents"/>
    <category term="gemini"/>
    <link rel="related" href="https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gemini-robotics-er-2-control-plane/</id>
    <title>Gemini Robotics ER 2 Is A Control Plane For Physical Agents</title>
    <link rel="alternate" href="https://newruntime.com/posts/gemini-robotics-er-2-control-plane/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-08-01T00:00:00.000Z</updated>
    <summary>Gemini Robotics ER 2 acts as a high-level embodied reasoning model that watches video, calls tools, plans multi-step tasks, and coordinates robot collaboration.</summary>
    <category term="robotics"/>
    <category term="agent-control-plane"/>
    <category term="video-understanding"/>
    <category term="gemini"/>
    <link rel="related" href="https://deepmind.google/blog/gemini-robotics-er-2-powering-robotics-with-video-understanding-task-orchestration-and-multi-robot-collaboration/"/>
  </entry>
  <entry>
    <id>https://newruntime.com/posts/gemini-spark-chrome-auto-browse/</id>
    <title>Gemini Spark Moves Browser Agents Into Chrome Sessions</title>
    <link rel="alternate" href="https://newruntime.com/posts/gemini-spark-chrome-auto-browse/"/>
    <published>2026-07-31T00:00:00.000Z</published>
    <updated>2026-07-31T00:00:00.000Z</updated>
    <summary>Gemini Spark now integrates with Chrome auto browse, using logged-in browser context with permission while keeping users in the loop for sensitive actions.</summary>
    <category term="browser-agents"/>
    <category term="agents"/>
    <category term="security"/>
    <category term="interfaces"/>
    <link rel="related" href="https://x.com/GeminiApp/status/2082923048362299629"/>
    <link rel="related" href="https://x.com/GeminiApp/status/2082923083355353454"/>
    <link rel="related" href="https://blog.google/innovation-and-ai/products/gemini-app/gemini-spark-updates-july-2026/"/>
  </entry>
</feed>
