Field note
Four narrow releases split evaluation into testable units: transferable agent memory, perceptual object-removal quality, a concrete tutoring moment, and large-scale multilingual refactoring.
New Runtime reading: The set is useful precisely because the metrics cannot be collapsed into one score. It shows evaluation moving from a headline leaderboard toward task-specific diagnostics.
Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.