Field note
Four signals point to the same operating problem: more capable coding agents can consume far more tokens, organizations are introducing budgets, product teams are adding spend controls, and executives still struggle to connect aggregate AI expenditure to outcomes. Raw token counts cannot resolve that tension.
The correct denominator is a completed, accepted task. Its cost includes the first run, retries, tool calls, review, rejected changes, test infrastructure, and any later rework. A model that costs twice as much per call may be cheaper if it completes a difficult task once; a cheap model becomes expensive when it creates repeated failures.
A practical control plane sets per-run and per-project budgets, segments workloads, routes bounded tasks to cheaper models, records stop reasons, and requires quality gates before marking completion. The personal GPT-5.6 versus GPT-5.5 token comparison is useful as a counterexample, not a universal benchmark; workload-level evidence must decide the policy.
