Field note
Mercor's APEX-Accounting is useful because it moves AI productivity evaluation into a recognizable professional domain.
Accounting is a good stress test for applied AI because the work is structured but not trivial: source documents, rules, reconciliation, judgment, and review all matter. A benchmark in this space can expose whether AI improves real task throughput or only performs well on disconnected examples.
The New Runtime angle is that productivity benchmarks are becoming domain assets. Every profession will need its own task suites, reviewer expectations, and failure taxonomy before anyone can make a serious claim about AI productivity gains.
