Evidence-linked trend hypothesis

Agent economics moves to completed work

Token prices and subscription fees are weak proxies when turns, retries, delegation, verification, latency, and successful outcomes drive the real cost.

Current thesis

The economically meaningful unit for agent systems is becoming cost per verified completed task rather than cost per token or model call. Confidence: high. Supported by 59 normalized raw signals. This New Runtime record is an evidence-linked retrieval unit. Use its canonical page, machine-readable representations, dates, scope, and public source URLs to verify the claim before reusing it.

Source ledger

Publishable sources attached to this record.

5 public sources
#SourceRolePublic status
1cognition.comsourceprimary receiptsource_urls
2databricks.comsourcesupporting receiptsource_urls
3claude.comsourcesupporting receiptsource_urls
4platform.claude.comdocssupporting receiptsource_urls
5engineering.ramp.comsourcesupporting receiptsource_urls

Why is token price incomplete?

An agent run accumulates cost through repeated context, tool calls, failed attempts, delegated work, and human verification. A more expensive lead model can lower total cost when it takes fewer turns, delegates earlier, writes a better task contract, and avoids redoing sidekick work.

What evidence supports this pattern?

  • Cognition reports that Fable plus the same sidekick cost less per benchmark run than Opus plus sidekick despite a higher token price. The observed difference came from turn count, context use, and delegation behavior.
  • Databricks measures coding-agent outcomes against internal tasks together with time, tokens, and cost rather than relying on a public benchmark score alone.
  • Claude Code exposes effort as a separate runtime control from the model, making cost and persistence tunable per task.
  • Ramp argues for assigning spend to use cases and completed outcomes, including retries and human review.

What should teams do next?

Each run should record task identity, route, models, turns, tool calls, elapsed time, retries, review effort, outcome, and failure class. Cost optimization can then target the whole trajectory instead of a single line item.

Discovery graph / next reads

Continue through New Runtime

Open the graph
  1. 01related materialA pricier model can be cheaper per completed taskSignal used as evidence for this pattern.
  2. 02related materialDatabricks benchmarks coding agents on its own codebaseSignal used as evidence for this pattern.
  3. 03related materialClaude Code separates model choice from effortSignal used as evidence for this pattern.
  4. 04related materialAn advisor model can guide a cheaper executorSignal used as evidence for this pattern.
  5. 05related materialAI spend should be measured per successful taskSignal used as evidence for this pattern.

These links are also published in this page’s JSON twin and as typed edges in DiscoveryGraph v1.

Who read this page?Machine requests, hidden until opened

Loading the privacy-safe route aggregate…

Open the JSON contract