Topic hub
Models
A New Runtime topic hub collecting signals, patterns, field notes, and public sources about models.
Field notes
What should readers understand next?
Codex Gets A Model Migration Deadline
OpenAI's ChatGPT and Codex changelog sets an August 31, 2026 cutoff for GPT-5.4 and GPT-5.4 mini in ChatGPT-signed Codex sessions, while keeping API-key paths available.
OpenAI Shows Efficiency Is a Full-Stack Agent Problem
OpenAI's GPT-5.6 efficiency write-up connects model training, inference optimization, and the Codex/ChatGPT Work harness into one compounding cost-performance loop.
OpenAI's ARC-AGI-3 Jump Was a Harness Result
OpenAI's ARC-AGI-3 write-up shows why agent benchmarks measure the model plus the runtime harness: retained reasoning and compaction changed both score and token use.
DeepsecBench Makes Security Agents a Cost/Recall Tradeoff
Vercel's DeepsecBench reframes security-agent evaluation around recall, precision, cost, total scan time, and a hidden benchmark that resists training leakage.
OpenAI Splits Transcription Into File and Live Workflows
OpenAI's new transcription guide makes recorded audio and live audio separate product paths, with gpt-transcribe and gpt-live-transcribe as the recommended starting models.
Raw signals
What changed recently?
Anthropic: In a review of our cybersecurity evaluations, we found three incidents in which a Claude model...
A public X post from Anthropic as a public source in its own right flags In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation...
Google DeepMind: One brain.
A public X post from Google DeepMind as a public source in its own right flags One brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwo....
Google DeepMind: This is how Gemini Robotics 2 helps @Apptronik’s Apollo 2 use whole body intelligence to pack...
A public X post from Google DeepMind as a public source in its own right flags This is how Gemini Robotics 2 helps @Apptronik’s Apollo 2 use whole body intelligence to pack for a sports game ↓
LangChain: ✅ Avoid cost overruns ✅ Limit runaway traffic ✅ Improve agent reliability ✅ Reduce sensitive d...
A public X post from LangChain with a linked primary source flags ✅ Avoid cost overruns ✅ Limit runaway traffic ✅ Improve agent reliability ✅ Reduce sensitive data exposure One gateway across your models and providers.
LangChain: LangSmith LLM Gateway is now available in public beta.
A public X post from LangChain as a public source in its own right flags LangSmith LLM Gateway is now available in public beta. Set spend and rate limits, determine fallback policies, and redact sensitive data before it reaches a model provider, all fr...
mem0: Experiment 2: Does a preference survive a hard context wipe?
A public X post from mem0 with a linked primary source flags Experiment 2: Does a preference survive a hard context wipe? We told Claude(mid-conversation) to write functions with standard for loops instead of list comprehensions. Then ran /...
OpenAI Developers: ImageGen in Codex just got a new lightbox and canvas.
A public X post from OpenAI Developers with a linked primary source flags ImageGen in Codex just got a new lightbox and canvas. Now, it’s even easier to explore and refine visuals in your workflow.
OpenAI: We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
A public X post from OpenAI with a linked primary source flags We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna. Combined with Luna’s new price, we expect Auto-review to cost about 10x less, makin...
Vercel Developers: Grok Voice Think Fast 2.0 is now on AI Gateway.
A public X post from Vercel Developers with a linked primary source flags Grok Voice Think Fast 2.0 is now on AI Gateway. A speech-to-speech model that reasons while it speaks, so it thinks through a query without adding latency. 𝚖𝚘𝚍𝚎𝚕: '𝚡𝚊𝚒/𝚐𝚛𝚘𝚔-𝚟𝚘𝚒𝚌𝚎-...
Augment Code: The best value for every token: GPT-5.6 Sol is now our default in Cosmos Eight models in eight...
A public X post from Augment Code as a public source in its own right flags The best value for every token: GPT-5.6 Sol is now our default in Cosmos Eight models in eight weeks! After testing them all on real, long-horizon software development tasks, we’r...
Cline: Cline is open source, so you can fork it and run this with your favorite model as well.
A public X post from Cline with a linked primary source flags Cline is open source, so you can fork it and run this with your favorite model as well. Read more about how we did this here:
Google AI Studio: we're looking for feedback on the ai studio model playground how would you rate your current e...
A public X post from Google AI Studio as a public source in its own right flags we're looking for feedback on the ai studio model playground how would you rate your current experience, and what specific features or improvements would you like to see added? le...
Hugging Face: Pro tip: add in a "Sign-In with Hugging Face" OAuth button to your website/app so that communi...
A public X post from Hugging Face with a linked primary source flags Pro tip: add in a "Sign-In with Hugging Face" OAuth button to your website/app so that community members can easily share their email, create model/dataset repos, store data in Bu...
OpenAI Developers: We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance.
A public X post from OpenAI Developers as a public source in its own right flags We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance. These improvements compound across inference and the agent loop, producing more useful work from...
OpenAI: A benchmark score reflects the model as well as the harness and settings used to run it.
A public X post from OpenAI with a linked primary source flags A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build o...
OpenAI: We quietly released the open-source Codex Security CLI, but Hacker News found it before we had...
A public X post from OpenAI with a linked primary source flags We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings...
Anthropic: Full technical details of both attacks are provided in our new papers: On HAWK: On AES: And th...
A public X post from Anthropic with a linked primary source flags Full technical details of both attacks are provided in our new papers: On HAWK: On AES: And the associated model chain-of-thought for AES:
Anthropic: We support this petition, signed by our CEO, several co-founders, and senior staff.
A public X post from Anthropic as a public source in its own right flags We support this petition, signed by our CEO, several co-founders, and senior staff. Our own research on recursive self-improvement, published last month, points to the need for to...
Augment Code: Our users have been really excited to try Kimi K3 and we're excited to bring it to Augment's p...
A public X post from Augment Code as a public source in its own right flags Our users have been really excited to try Kimi K3 and we're excited to bring it to Augment's product family. It is the most capable open-source model that we have tested to date....
Factory: We're firm believers in the open security ecosystem, contributing across our platform: • Publi...
A public X post from Factory with a linked primary source flags We're firm believers in the open security ecosystem, contributing across our platform: • Publishing two open-weight models behind Droid Shield 2.0 • Using Autonomous Security Revi...
Grok: Grok 4.5 is now live in GitHub Copilot.
A public X post from Grok with a linked primary source flags Grok 4.5 is now live in GitHub Copilot. Switch from the model picker for frontier intelligence at top speeds.
Hugging Face: Training Agents 3: Learn how to do train a local/ open weight agent with reinforcement learning
A public X post from Hugging Face as a public source in its own right flags Training Agents 3: Learn how to do train a local/ open weight agent with reinforcement learning
LangChain: Our data agent now handles roughly 40x the request volume our 3 person data team could manage...
A public X post from LangChain with a linked primary source flags Our data agent now handles roughly 40x the request volume our 3 person data team could manage directly. Now, our data team can focus on the models, context, and guardrails that ma...
OpenAI Developers: We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for lo...
A public X post from OpenAI Developers with a linked primary source flags We're introducing two new transcription models in the API: • GPT-Live-Transcribe: built for low-latency live transcription. • GPT-Transcribe: optimized for asynchronous transcript...
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | alignment.anthropic.comsource | primary receipt | source_urls |
| 2 | anthropic.comsource | supporting receipt | source_urls |
| 3 | arxiv.orgpaper | supporting receipt | source_urls |
| 4 | blogs.nvidia.comarticle | supporting receipt | source_urls |
| 5 | blogs.nvidia.comarticle | supporting receipt | source_urls |
| 6 | claude.comsource | supporting receipt | source_urls |
| 7 | cline.botsource | supporting receipt | source_urls |
| 8 | cline.botsource | supporting receipt | source_urls |
| 9 | cursor.comsource | supporting receipt | source_urls |
| 10 | cursor.comsource | supporting receipt | source_urls |
Showing 10 of 150; the complete set is exposed in the JSON route.