Topic hub
Model routing
How teams route work across model families, effort levels, cheap executors, judges, and local/runtime constraints.
Short answer
- Model routing treats models as swappable execution resources rather than a single product choice.
- The pattern is to match task risk, latency, cost, and verification needs to the right model or model sequence.
- This makes the harness more durable than any one frontier model release.
Pattern memory
Current hypotheses
- high
Harness architecture outlives model choice
For production agents, the harness is becoming a more durable product boundary than the identity of the model running inside it.
- high
Agent economics moves to completed work
The economically meaningful unit for agent systems is becoming cost per verified completed task rather than cost per token or model call.
Raw signals
Recent observations
OpenCode makes the model a swappable dependency
OpenCode keeps rules, skills, permissions, sessions, and tools stable while the underlying model changes.
A pricier model can be cheaper per completed task
Cognition reports that Fable 5 completed coding work with fewer steps and output tokens than its previous lead model.
Claude Code separates model choice from effort
Anthropic exposes model selection and effort level as different controls for capability, token use, latency, and persistence.
Improve routes architecture and execution to different models
The Improve skill audits a repository with a stronger planning model and delegates bounded fixes to cheaper executors.
An advisor model can guide a cheaper executor
The advisor-tool pattern lets a fast executor request bounded analysis from a stronger model while keeping control of the task loop.
An inference gateway makes model switching operational
Inference.net exposes model routing through a stable gateway so clients can change providers without rewriting every integration.
Coinbase Reports Lower AI Spend as Token Use Rises
A production account suggests routing, caching, model choice, and workflow design can reduce unit cost even while aggregate agent usage expands.
GLM-5.2 Changes the Economics of Claude Code Workflows
A capable open model routed through compatible infrastructure can replace premium inference for parts of a coding-agent workload.
Sakana Fugu Presents Multiple Agents as One Model Endpoint
Fugu hides a multi-agent debate and synthesis system behind a model-like interface, making orchestration an implementation detail of inference.
Fusion API Uses Multiple Models and a Judge for One Answer
A single endpoint can route a question to several models and synthesize their outputs, moving model selection and comparison behind an orchestration layer.
Vercel Treats Production AI as Multi-Model Infrastructure
Vercel's gateway view assumes model providers are routable dependencies with fallback, observability, and policy around them.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | claude.comsource | primary receipt | source_urls |
| 2 | code.claude.comsource | supporting receipt | source_urls |
| 3 | cognition.comsource | supporting receipt | source_urls |
| 4 | databricks.comsource | supporting receipt | source_urls |
| 5 | docs.inference.netdocs | supporting receipt | source_urls |
| 6 | engineering.ramp.comsource | supporting receipt | source_urls |
| 7 | github.comrepo | supporting receipt | source_urls |
| 8 | github.comrepo | supporting receipt | source_urls |
| 9 | github.comrepo | supporting receipt | source_urls |
| 10 | github.comrepo | supporting receipt | source_urls |
Showing 10 of 25; the complete set is exposed in the JSON route.