Field note
Practical evaluation increasingly measures a model together with its harness, retained state, deterministic tests, review gates, and shipped artifacts.
New Runtime reading: A headline score without runner policy and task evidence explains less. The stronger unit is a reproducible diagnostic that can trace a failure or accepted result through the runtime.
Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.
