Field note
A mathematical taxonomy, an individual writing preference, and a small-model catalog show why asking which model is best is meaningless without task class, budget, and an accepted result.
New Runtime reading: The sources do not form a benchmark. Their value is a qualified guide for designing one: define the task, constraints, evaluation artifact, and deployment boundary first.
Evidence boundary: this item uses the listed public sources and keeps vendor, author, or reporter claims attributed. The queued page is an editorial synthesis, not an independent validation of every reported metric.