Field note
DeepSeek's update page positions V4-Flash as a public-beta model adapted for coding-agent use through API workflows. That makes it worth tracking, but not as a vague cheaper-model story.
The useful question is whether it can hold a bounded lane inside a real harness. A budget model is valuable when it reliably handles small edits, refactors, tests, summaries, or retrieval passes while emitting enough failure evidence for the parent agent to route around mistakes.
Price claims should therefore be treated as inputs to an eval, not as the conclusion. The model needs a task boundary, a retry policy, a fallback lane, and receipts for the cases it should not own.
For New Runtime, this is a routing candidate: measure it inside Codex-style tasks, compare failure receipts against current lanes, and only then decide whether it belongs in a daily agent loop.
