The Fable and Jacobi-hypothesis counterexample story matters less as another leaderboard entry and more as a different kind of evaluation. The model produces an artifact that must survive outside the benchmark harness.
In that scenario, the main question is not the model’s score. It is whether a specialist can independently verify the result, find errors, reproduce the reasoning, and see where trust should stop.
That is closer to scientific artifact evaluation than a familiar model benchmark. If the artifact survives external review, it changes the capability map more than an abstract percentage.
New Runtime read: the next serious eval for frontier systems is not a knowledge quiz. It is a verifiable contribution in a domain where an expert can say: this is new, this is wrong, or this needs more work.