Field note
OpenAI and Cerebras added an ultrafast GPT-5.6 Sol route that Cerebras says can reach up to 750 output tokens per second.
New Runtime reading: Throughput matters when an agent repeatedly plans, calls tools, inspects results, and replans. The relevant test is end-to-end task latency and cost under a real harness, not the provider's peak token rate in isolation.
Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.