Field note
OpenAI published ten claimed advances across mathematics and theoretical computer science, produced while evaluating an internal version of Astra, described as its next major model.
The operational signal is not just that a model produced proofs. The structure matters more: the search cost is presented as roughly $2,000 at Sol API rates, humans prepared the arguments into manuscripts, and each argument was formalized in a Lean certificate. That turns model capability into an evidence pipeline: generate candidate arguments, make them legible to humans, formalize them, and invite the mathematical community to test the results.
For New Runtime, this belongs near the front of the queue because it changes the unit of AI research work. The interesting object is no longer a benchmark score or a demo answer; it is a package of claims, manuscripts, certificates, and accountability. If this pattern holds, the next frontier is less about whether a model can be brilliant in isolation and more about whether the surrounding verification loop can absorb model-generated research at speed.
