Topic hub
Model routing
How teams route work across model families, effort levels, cheap executors, judges, and local/runtime constraints.
Pattern memory
What patterns are emerging?
- high
Harness architecture outlives model choice
For production agents, the harness is becoming a more durable product boundary than the identity of the model running inside it.
- high
Agent economics moves to completed work
The economically meaningful unit for agent systems is becoming cost per verified completed task rather than cost per token or model call.
Field notes
What should readers understand next?
Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost
Ramp's internal LLM gateway uses failure-aware online learning to reorder model and service-tier candidates, cutting spend without relaxing request deadlines.
Kimi K3 Is Moving From Model Launch to Routed Coding Component
Vercel and Factory show the next phase of Kimi K3 adoption: one open model becomes a routable, priced, regional, fast-or-standard component inside coding-agent platforms.
The Coding Harness Is Becoming Independent From the Model
Practitioners are routing different models through coding-agent workflows, while production systems increasingly choose model and effort per role instead of per product.
Raw signals
What changed recently?
OpenCode makes the model a swappable dependency
OpenCode keeps rules, skills, permissions, sessions, and tools stable while the underlying model changes.
A pricier model can be cheaper per completed task
Cognition reports that Fable 5 completed coding work with fewer steps and output tokens than its previous lead model.
Claude Code separates model choice from effort
Anthropic exposes model selection and effort level as different controls for capability, token use, latency, and persistence.
Improve routes architecture and execution to different models
The Improve skill audits a repository with a stronger planning model and delegates bounded fixes to cheaper executors.
An advisor model can guide a cheaper executor
The advisor-tool pattern lets a fast executor request bounded analysis from a stronger model while keeping control of the task loop.
An inference gateway makes model switching operational
Inference.net exposes model routing through a stable gateway so clients can change providers without rewriting every integration.
Coinbase Reports Lower AI Spend as Token Use Rises
A production account suggests routing, caching, model choice, and workflow design can reduce unit cost even while aggregate agent usage expands.
GLM-5.2 Changes the Economics of Claude Code Workflows
A capable open model routed through compatible infrastructure can replace premium inference for parts of a coding-agent workload.
Sakana Fugu Presents Multiple Agents as One Model Endpoint
Fugu hides a multi-agent debate and synthesis system behind a model-like interface, making orchestration an implementation detail of inference.
Fusion API Uses Multiple Models and a Judge for One Answer
A single endpoint can route a question to several models and synthesize their outputs, moving model selection and comparison behind an orchestration layer.
GitHub / yichuan-w/LEANN: Routable Model Components
The archive captures GitHub / yichuan-w/LEANN as a dated public record from GitHub / yichuan-w/LEANN. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.
Vercel Treats Production AI as Multi-Model Infrastructure
Vercel's gateway view assumes model providers are routable dependencies with fallback, observability, and policy around them.
GitHub / chenglou/pretext: Executable Design Context
The archive captures GitHub / chenglou/pretext as a dated public record from GitHub / chenglou/pretext. It documents design systems becoming machine-readable context, constraints, and review loops and is retained as supporting evidence for the executable design context trend.
Devl: Executable Design Context
The archive captures Devl as a dated public record from Devl. It documents design systems becoming machine-readable context, constraints, and review loops and is retained as supporting evidence for the executable design context trend.
OpenAI: Routable Model Components
The archive captures OpenAI as a dated public record from OpenAI. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.
Claude: Agent Security Runtime Boundaries
The archive captures YouTube source as a dated public record from YouTube source. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.
Google / Developers Tools: Routable Model Components
The archive captures Google / Developers Tools as a dated public record from Google / Developers Tools. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.
Google Developers / Bring State Of Art Agentic Skills To: Ai-Native Operating Models
The archive captures Google Developers / Bring State Of Art Agentic Skills To as a dated public record from Google Developers / Bring State Of Art Agentic Skills To. It documents AI adoption shifting jobs, coordination, review, and organizational capacity and is retained as supporting evidence for the AI-native operating models trend.
Openinterpreter: Task-Specific Ai Interfaces
The archive captures Openinterpreter as a dated public record from Openinterpreter. It documents model output moving from plain answers into generated task interfaces and actions and is retained as supporting evidence for the task-specific AI interfaces trend.
Techcrunch / Meta Ai Security Researcher Said Openclaw Agent: Task-Specific Ai Interfaces
The archive captures Techcrunch / Meta Ai Security Researcher Said Openclaw Agent as a dated public record from Techcrunch / Meta Ai Security Researcher Said Openclaw Agent. It documents model output moving from plain answers into generated task interfaces and actions and is retained as pressure-testing evidence for the task-specific AI interfaces trend.
Anthropic / Claude Sonnet: Routable Model Components
The archive captures Anthropic / Claude Sonnet as a dated public record from Anthropic / Claude Sonnet. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.
GitHub / memodb-io/Acontext: Agent-Ready Software
The archive captures GitHub / memodb-io/Acontext as a dated public record from GitHub / memodb-io/Acontext. It documents software exposing explicit capabilities, permissions, and machine-readable actions and is retained as pressure-testing evidence for the agent-ready software trend.
GitHub / pydantic/monty: Agent-Ready Software
The archive captures GitHub / pydantic/monty as a dated public record from GitHub / pydantic/monty. It documents software exposing explicit capabilities, permissions, and machine-readable actions and is retained as supporting evidence for the agent-ready software trend.
X source / Akshay Pachaar: Verification Bandwidth
The archive captures X source / Akshay Pachaar as a dated public record from X source / Akshay Pachaar. It documents evaluation, review, and observability becoming the bottleneck after generation accelerates and is retained as pressure-testing evidence for the verification bandwidth trend.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | ai.google.devsource | primary receipt | source_urls |
| 2 | aiechoes.substack.comsource | supporting receipt | source_urls |
| 3 | anthropic.comsource | supporting receipt | source_urls |
| 4 | app.klingai.comsource | supporting receipt | source_urls |
| 5 | arxiv.orgpaper | supporting receipt | source_urls |
| 6 | arxiv.orgpaper | supporting receipt | source_urls |
| 7 | arxiv.orgpaper | supporting receipt | source_urls |
| 8 | blog.googlearticle | supporting receipt | source_urls |
| 9 | blog.googlearticle | supporting receipt | source_urls |
| 10 | blog.nilenso.comarticle | supporting receipt | source_urls |
Showing 10 of 114; the complete set is exposed in the JSON route.