Topic hub
Coding agents
A living map of coding agents as they move from autocomplete into supervised software work, review loops, environments, and operating cost.
Pattern memory
What patterns are emerging?
- high
Verification bandwidth is the scarce engineering resource
The primary bottleneck in agentic software delivery is moving from code production to the human and machine capacity required to verify it.
Field notes
What should readers understand next?
A Software Factory Connects Agents Through Verified Outcomes
Augment and Warp describe team-level agent loops that move work from trigger and specification through implementation, verification, release, and measured improvement.
Claude Code Auto Mode Gates Actions Instead Of Explanations
Claude Code Auto Mode combines an input injection probe with a two-stage action classifier, preserving autonomy while exposing an honest residual miss rate.
Cline Hooks Put Deterministic Rules Inside The Agent Loop
Cline's plugin hooks show how an agent harness can journal every run and block dangerous tool calls without waiting for the model to choose a guardrail.
EvoCode-Bench Exposes Multi-Turn Regression Risk
EvoCode-Bench tests coding agents across persistent workspaces and evolving requirements, where regressions become the dominant failure mode.
Loop Engineering Needs State Pruning, Not Infinite Chat
The Google ADK loop-engineering article frames self-correcting agents as desired-state systems with pruning, validation, and circuit breakers.
GPT-5.6 Turns Efficiency Work Into API Economics
OpenAI turned GPT-5.6 serving and kernel efficiency gains into lower Luna and Terra prices, plus a faster Sol mode for latency-sensitive API workloads.
ReviewBench Turns Code Review Into An Agent Eval
LangChain's ReviewBench uses real PR review history to test whether code-review agents can recover substantive reviewer findings without flooding humans with noise.
GitHub Stacked PRs Turn Large Agent Changes Into Reviewable Chains
GitHub's stacked pull request preview gives large dependent code changes a native review path, which matters as agents produce broader diffs.
Cline Turns Recursive Self-Improvement Into Harness Work
Cline's Terminal-Bench run is not a singularity story; it is a concrete loop where an agent reads traces, patches the harness, reruns evals, and hands a PR to humans.
Cursor Treats the Cloud Agent Environment as the Product
Cursor's cloud-agent environment write-up shows why agent performance depends on dependencies, commands, security boundaries, end-to-end tests, and self-healing diagnostics.
Cursor's SQLite Swarm Makes Coordination the Expensive Part
Cursor's SQLite experiment shows why agent-swarm economics depend on task trees, shared memory, conflict handling, review lenses, and selective use of expensive planners.
Factory and Comarch Show the Night-Shift Shape of Agent Work
Factory's Comarch case study is less about one coding assistant and more about governed agent missions that continue execution while humans set direction and review.
miniSQLite Makes the Coding-Swarm Claim Inspectable
The released Rust database turns Cursor's swarm experiment into an artifact that can be read, built, tested, and challenged instead of accepted as a benchmark chart.
A Codex Skill Turns Multi-Agent Work Into a Reusable Control Surface
A shared Codex orchestration skill points to a future where agent workflows are packaged as portable operational routines.
Agentic Engineering Looks Like Workflow Design, Not Hands-Free Coding
A full-stack agentic engineering walkthrough shows the operating pattern around planning, validation, browser checks, and human review.
Claude Code Certifications Turn Agent Use Into a Credential Layer
Anthropic certification activity around Claude Code is a signal that agentic development is becoming a managed enterprise capability.
Claude.md Shrinkage Says the Harness Is Learning What Not to Say
A reported 80 percent reduction in Claude Code prompt material points to a maturing harness discipline around concise system context.
Code Review Becomes the Agent Bottleneck
Coles shows that reviewing generated code is becoming the scarce engineering work, not merely a cleanup pass after agent output.
Codex Hooks Close the Type Error Loop
Codex hooks make validation output an active part of the agent loop, so type errors and checks can be returned before the task drifts.
Codex Security CLI Turns Security Review Into a Scannable Workbench
OpenAI's Codex Security CLI packages repository, diff, working-tree, export, validation, and patch flows into a security-review workbench rather than a single scanner command.
Prompt Caching Turns Agent Context Into Infrastructure
Earendil frames prompt caching as an agent systems primitive, where stable context becomes a cost, latency, and architecture concern.
RAPTOR Loop Hunt Packages Security Hunting as an Agent Skill
RAPTOR Loop Hunt shows security research moving toward looped agent skills with altitude changes, evidence collection, and review checkpoints.
GitHub Copilot App Turns Agent Work Into a Desktop Control Surface
GitHub's Copilot app GA is a catch-up signal: agentic coding is being packaged as sessions, worktrees, canvases, validation, automations, MCP servers, and skills.
Kimi K3 Is Moving From Model Launch to Routed Coding Component
Vercel and Factory show the next phase of Kimi K3 adoption: one open model becomes a routable, priced, regional, fast-or-standard component inside coding-agent platforms.
The Coding Harness Is Becoming Independent From the Model
Practitioners are routing different models through coding-agent workflows, while production systems increasingly choose model and effort per role instead of per product.
Uber's Agentic Adoption Claims Need Two Kinds of Evidence
An Uber employee reports 99% AI-tool adoption and agent attribution for over 70% of pull requests, while official engineering evidence confirms broad AI review at production scale.
Devin Outposts Splits the Agent Brain from the Execution Plane
Devin Outposts moves command execution, repository access, and sandbox lifecycle into customer-controlled infrastructure while the agent loop remains in Cognition's cloud.
Gemini 3.6 Flash Moves the Agent Race Toward Cost per Task
Google's Gemini 3.6 Flash release frames the model race around token efficiency, built-in computer use, and specialized cyber agents rather than raw chat intelligence alone.
AI Coding Workflow: From Idea to Verifiable Work
AI coding works better when checkable artifacts stand between the idea and the code: specs, tickets, TDD, fresh-context review, and manual QA.
Coding Agent Cost Is Cut in Environment Config, Not Prompts
Claude Code and Codex costs are reduced by environment design, not by asking the agent to read less: output filtering, repo maps, model routing, and stable prompt caching.
Coding Agent Sandboxes Break in Places Teams Do Not Expect
Pillar shows that agent sandboxes must be assessed not only around the agent process, but around files, configs, allowlisted commands, and local daemons the host later trusts.
Kimi K3: Frontend Is Becoming the Model Race Arena
Kimi K3 moves open model competition into visual software engineering: frontend benchmarks require not only code, but layout, screenshots, accessibility, and human preference.
Raw signals
What changed recently?
Factory: At Comarch, hundreds of engineers now use the Factory platform to run agents end-to-end across...
A public X post from Factory with a linked primary source flags At Comarch, hundreds of engineers now use the Factory platform to run agents end-to-end across the software development lifecycle. • 40% greater engineering efficiency • >30% high...
GitHub: Stacked PRs now on GitHub 🥞
A public X post from GitHub with a linked primary source flags Stacked PRs now on GitHub 🥞
Google Gemini: Create, edit, and summarize with Gemini on macOS using your voice.
A public X post from Google Gemini with a linked primary source flags Create, edit, and summarize with Gemini on macOS using your voice. Stay in your flow by speaking into any active window: Dictate clean text or ask Gemini to transform highlighted...
mem0: Experiment 2: Does a preference survive a hard context wipe?
A public X post from mem0 with a linked primary source flags Experiment 2: Does a preference survive a hard context wipe? We told Claude(mid-conversation) to write functions with standard for loops instead of list comprehensions. Then ran /...
OpenAI Developers: ImageGen in Codex just got a new lightbox and canvas.
A public X post from OpenAI Developers with a linked primary source flags ImageGen in Codex just got a new lightbox and canvas. Now, it’s even easier to explore and refine visuals in your workflow.
OpenAI: We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
A public X post from OpenAI with a linked primary source flags We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna. Combined with Luna’s new price, we expect Auto-review to cost about 10x less, makin...
A Codex Skill Turns Multi-Agent Work Into a Reusable Control Surface
A shared Codex orchestration skill points to a future where agent workflows are packaged as portable operational routines.
Agentic Engineering Looks Like Workflow Design, Not Hands-Free Coding
A full-stack agentic engineering walkthrough shows the operating pattern around planning, validation, browser checks, and human review.
Augment Code: The best value for every token: GPT-5.6 Sol is now our default in Cosmos Eight models in eight...
A public X post from Augment Code as a public source in its own right flags The best value for every token: GPT-5.6 Sol is now our default in Cosmos Eight models in eight weeks! After testing them all on real, long-horizon software development tasks, we’r...
Claude Code Certifications Turn Agent Use Into a Credential Layer
Anthropic certification activity around Claude Code is a signal that agentic development is becoming a managed enterprise capability.
Claude.md Shrinkage Says the Harness Is Learning What Not to Say
A reported 80 percent reduction in Claude Code prompt material points to a maturing harness discipline around concise system context.
Code Review Becomes the Agent Bottleneck
Coles shows that reviewing generated code is becoming the scarce engineering work, not merely a cleanup pass after agent output.
Codex Hooks Close the Type Error Loop
Codex hooks make validation output an active part of the agent loop, so type errors and checks can be returned before the task drifts.
Cursor: Cursor is now on iPad.
A public X post from Cursor with a linked primary source flags Cursor is now on iPad. All the power of Cursor on iPhone, with more room to work with agents.
OpenAI Developers: We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance.
A public X post from OpenAI Developers as a public source in its own right flags We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance. These improvements compound across inference and the agent loop, producing more useful work from...
OpenAI: We quietly released the open-source Codex Security CLI, but Hacker News found it before we had...
A public X post from OpenAI with a linked primary source flags We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings...
Prompt Caching Turns Agent Context Into Infrastructure
Earendil frames prompt caching as an agent systems primitive, where stable context becomes a cost, latency, and architecture concern.
RAPTOR Loop Hunt Packages Security Hunting as an Agent Skill
RAPTOR Loop Hunt shows security research moving toward looped agent skills with altitude changes, evidence collection, and review checkpoints.
Augment Code: Our users have been really excited to try Kimi K3 and we're excited to bring it to Augment's p...
A public X post from Augment Code as a public source in its own right flags Our users have been really excited to try Kimi K3 and we're excited to bring it to Augment's product family. It is the most capable open-source model that we have tested to date....
Cursor: Today we're launching Cursor Start, a new ₹649/month plan for developers in India.
A public X post from Cursor with a linked primary source flags Today we're launching Cursor Start, a new ₹649/month plan for developers in India. Start includes generous access to Grok 4.5 and Composer, so you can plan, build, test, and ship...
Factory: We're firm believers in the open security ecosystem, contributing across our platform: • Publi...
A public X post from Factory with a linked primary source flags We're firm believers in the open security ecosystem, contributing across our platform: • Publishing two open-weight models behind Droid Shield 2.0 • Using Autonomous Security Revi...
GitHub: 5.
A public X post from GitHub with a linked primary source flags 5. Turn on code scanning Code scanning uses CodeQL to identify patterns that lead to vulnerabilities, including injection flaws, unsafe deserialization, and insecure GitHub Action...
GitHub: You're not behind.
A public X post from GitHub as a public source in its own right flags You're not behind. There's no secret everyone else has. There's just the harness, and it's mostly all you need. @burkeholland gives you a simple, repeatable workflow for GitHub Co...
Zed: Congrats to @poolsideai on launching Poolside Desktop Assistant.
A public X post from Zed as a public source in its own right flags Congrats to @poolsideai on launching Poolside Desktop Assistant. Their Laguna models power an ACP agent that runs in any client, Zed included, and now their desktop app is an ACP...
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | addyo.substack.comsource | primary receipt | source_urls |
| 2 | addyosmani.comsource | supporting receipt | source_urls |
| 3 | addyosmani.comsource | supporting receipt | source_urls |
| 4 | agent-cookbook.comsource | supporting receipt | source_urls |
| 5 | agentos-sdk.devsource | supporting receipt | source_urls |
| 6 | ai.google.devsource | supporting receipt | source_urls |
| 7 | aitmpl.comsource | supporting receipt | source_urls |
| 8 | alignment.anthropic.comsource | supporting receipt | source_urls |
| 9 | alilleybrinker.comsource | supporting receipt | source_urls |
| 10 | americanbanker.comsource | supporting receipt | source_urls |
Showing 10 of 309; the complete set is exposed in the JSON route.