Topic hub

Model routing

How teams route work across model families, effort levels, cheap executors, judges, and local/runtime constraints.

Retrieval answer

How teams route work across model families, effort levels, cheap executors, judges, and local/runtime constraints. Model routing treats models as swappable execution resources rather than a single product choice. The pattern is to match task risk, latency, cost, and verification needs to the right model or model sequence. This makes the harness more durable than any one frontier model release.

Pattern memory

What patterns are emerging?

2 patterns
  1. high

    Harness architecture outlives model choice

    For production agents, the harness is becoming a more durable product boundary than the identity of the model running inside it.

  2. high

    Agent economics moves to completed work

    The economically meaningful unit for agent systems is becoming cost per verified completed task rather than cost per token or model call.

Field notes

What should readers understand next?

3 notes
  1. Ramp Teaches Its Gateway To Route By Failure, Latency, And Cost

    Ramp's internal LLM gateway uses failure-aware online learning to reorder model and service-tier candidates, cutting spend without relaxing request deadlines.

  2. Kimi K3 Is Moving From Model Launch to Routed Coding Component

    Vercel and Factory show the next phase of Kimi K3 adoption: one open model becomes a routable, priced, regional, fast-or-standard component inside coding-agent platforms.

  3. The Coding Harness Is Becoming Independent From the Model

    Practitioners are routing different models through coding-agent workflows, while production systems increasingly choose model and effort per role instead of per product.

Raw signals

What changed recently?

53 signals
  1. OpenCode makes the model a swappable dependency

    OpenCode keeps rules, skills, permissions, sessions, and tools stable while the underlying model changes.

  2. A pricier model can be cheaper per completed task

    Cognition reports that Fable 5 completed coding work with fewer steps and output tokens than its previous lead model.

  3. Claude Code separates model choice from effort

    Anthropic exposes model selection and effort level as different controls for capability, token use, latency, and persistence.

  4. Improve routes architecture and execution to different models

    The Improve skill audits a repository with a stronger planning model and delegates bounded fixes to cheaper executors.

  5. An advisor model can guide a cheaper executor

    The advisor-tool pattern lets a fast executor request bounded analysis from a stronger model while keeping control of the task loop.

  6. An inference gateway makes model switching operational

    Inference.net exposes model routing through a stable gateway so clients can change providers without rewriting every integration.

  7. Coinbase Reports Lower AI Spend as Token Use Rises

    A production account suggests routing, caching, model choice, and workflow design can reduce unit cost even while aggregate agent usage expands.

  8. GLM-5.2 Changes the Economics of Claude Code Workflows

    A capable open model routed through compatible infrastructure can replace premium inference for parts of a coding-agent workload.

  9. Sakana Fugu Presents Multiple Agents as One Model Endpoint

    Fugu hides a multi-agent debate and synthesis system behind a model-like interface, making orchestration an implementation detail of inference.

  10. Fusion API Uses Multiple Models and a Judge for One Answer

    A single endpoint can route a question to several models and synthesize their outputs, moving model selection and comparison behind an orchestration layer.

  11. GitHub / yichuan-w/LEANN: Routable Model Components

    The archive captures GitHub / yichuan-w/LEANN as a dated public record from GitHub / yichuan-w/LEANN. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.

  12. Vercel Treats Production AI as Multi-Model Infrastructure

    Vercel's gateway view assumes model providers are routable dependencies with fallback, observability, and policy around them.

  13. GitHub / chenglou/pretext: Executable Design Context

    The archive captures GitHub / chenglou/pretext as a dated public record from GitHub / chenglou/pretext. It documents design systems becoming machine-readable context, constraints, and review loops and is retained as supporting evidence for the executable design context trend.

  14. Devl: Executable Design Context

    The archive captures Devl as a dated public record from Devl. It documents design systems becoming machine-readable context, constraints, and review loops and is retained as supporting evidence for the executable design context trend.

  15. OpenAI: Routable Model Components

    The archive captures OpenAI as a dated public record from OpenAI. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.

  16. Claude: Agent Security Runtime Boundaries

    The archive captures YouTube source as a dated public record from YouTube source. It documents agent security expanding from prompt policy into memory, tools, sandboxes, and execution boundaries and is retained as branch-opening evidence for the agent security runtime boundaries trend.

  17. Google / Developers Tools: Routable Model Components

    The archive captures Google / Developers Tools as a dated public record from Google / Developers Tools. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.

  18. Google Developers / Bring State Of Art Agentic Skills To: Ai-Native Operating Models

    The archive captures Google Developers / Bring State Of Art Agentic Skills To as a dated public record from Google Developers / Bring State Of Art Agentic Skills To. It documents AI adoption shifting jobs, coordination, review, and organizational capacity and is retained as supporting evidence for the AI-native operating models trend.

  19. Openinterpreter: Task-Specific Ai Interfaces

    The archive captures Openinterpreter as a dated public record from Openinterpreter. It documents model output moving from plain answers into generated task interfaces and actions and is retained as supporting evidence for the task-specific AI interfaces trend.

  20. Techcrunch / Meta Ai Security Researcher Said Openclaw Agent: Task-Specific Ai Interfaces

    The archive captures Techcrunch / Meta Ai Security Researcher Said Openclaw Agent as a dated public record from Techcrunch / Meta Ai Security Researcher Said Openclaw Agent. It documents model output moving from plain answers into generated task interfaces and actions and is retained as pressure-testing evidence for the task-specific AI interfaces trend.

  21. Anthropic / Claude Sonnet: Routable Model Components

    The archive captures Anthropic / Claude Sonnet as a dated public record from Anthropic / Claude Sonnet. It documents models becoming replaceable or specialized components inside a more durable runtime and is retained as supporting evidence for the routable model components trend.

  22. GitHub / memodb-io/Acontext: Agent-Ready Software

    The archive captures GitHub / memodb-io/Acontext as a dated public record from GitHub / memodb-io/Acontext. It documents software exposing explicit capabilities, permissions, and machine-readable actions and is retained as pressure-testing evidence for the agent-ready software trend.

  23. GitHub / pydantic/monty: Agent-Ready Software

    The archive captures GitHub / pydantic/monty as a dated public record from GitHub / pydantic/monty. It documents software exposing explicit capabilities, permissions, and machine-readable actions and is retained as supporting evidence for the agent-ready software trend.

  24. X source / Akshay Pachaar: Verification Bandwidth

    The archive captures X source / Akshay Pachaar as a dated public record from X source / Akshay Pachaar. It documents evaluation, review, and observability becoming the bottleneck after generation accelerates and is retained as pressure-testing evidence for the verification bandwidth trend.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01related materialHarness architecture outlives model choiceContinue through the Model routing topic.
  2. 02related materialAgent economics moves to completed workContinue through the Model routing topic.
  3. 03related materialRamp Teaches Its Gateway To Route By Failure, Latency, And CostContinue through the Model routing topic.
  4. 04related materialKimi K3 Is Moving From Model Launch to Routed Coding ComponentContinue through the Model routing topic.
  5. 05related materialThe Coding Harness Is Becoming Independent From the ModelContinue through the Model routing topic.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract