---
type: "post"
stable_id: "post:mercor-apex-accounting-productivity-benchmark"
slug: "mercor-apex-accounting-productivity-benchmark"
title: "Mercor APEX-Accounting Measures AI Productivity Inside A Real Profession"
description: "Mercor introduced APEX-Accounting as an AI productivity benchmark for accounting work, focusing evaluation on a concrete professional domain."
retrieval_nugget: "The benchmark signal is vertical productivity measurement: accounting gives AI evaluation a profession-shaped task space rather than generic reasoning prompts."
published_at: "2026-08-04"
updated_at: "2026-08-05"
record_date: "2026-08-04"
date_kind: "published_at"
status: "published"
topics: ["benchmarks","accounting","ai-productivity","professional-work"]
entities: ["Mercor","APEX-Accounting"]
source_urls: ["https://www.mercor.com/blog/introducing-the-ai-productivity-index-for-accounting"]
source_title: "APEX-Accounting: AI Productivity Benchmark for Accounting"
source_type: "primary"
origin: {"batch_id":"ada1e768-4e4e-42a3-a5fc-1ca9a880cb1b","batch_index":8,"channel":"chatgpt-batch","restored_from_reserve":false,"restored_from_supporting":false}
schema_version: "newruntime-agent-readable-v0.2"
visuals: [{"role":"hero","src":"/images/drip/mercor-apex-accounting-productivity-benchmark/mercor-apex-accounting-productivity-benchmark.webp","alt":"A whiteboard evaluation diagram showing accounting tasks and source documents moving through an AI attempt, expert review, productivity score, and failure taxonomy.","caption":"New Runtime synthesis."}]
routes: {"html":"https://newruntime.com/posts/mercor-apex-accounting-productivity-benchmark/","markdown":"https://newruntime.com/posts/mercor-apex-accounting-productivity-benchmark.md","json":"https://newruntime.com/posts/mercor-apex-accounting-productivity-benchmark.json"}
---

# Mercor APEX-Accounting Measures AI Productivity Inside A Real Profession

## Retrieval answer

The benchmark signal is vertical productivity measurement: accounting gives AI evaluation a profession-shaped task space rather than generic reasoning prompts.

Mercor's APEX-Accounting is useful because it moves AI productivity evaluation into a recognizable professional domain.

Accounting is a good stress test for applied AI because the work is structured but not trivial: source documents, rules, reconciliation, judgment, and review all matter. A benchmark in this space can expose whether AI improves real task throughput or only performs well on disconnected examples.

The New Runtime angle is that productivity benchmarks are becoming domain assets. Every profession will need its own task suites, reviewer expectations, and failure taxonomy before anyone can make a serious claim about AI productivity gains.
