Topic hub
Agent evals
A record of evals moving from model benchmarks toward behavior, workflow completion, skill reliability, and review burden.
Short answer
- Agent evals need to measure behavior in a workflow, not only answer quality in a prompt.
- The relevant unit is completed work under constraints: tools, files, permissions, time, cost, and review outcome.
- A good eval makes regressions visible before the agent is trusted with larger work.
Pattern memory
Current hypotheses
- medium
Skills become a portable capability layer
Agent skills are emerging as a portable capability layer, but their value depends on progressive disclosure, provenance, security review, and behavioral evaluation.
- high
Verification bandwidth is the scarce engineering resource
The primary bottleneck in agentic software delivery is moving from code production to the human and machine capacity required to verify it.
Raw signals
Recent observations
Agent skills need behavioral evals, not prose review
Benchmarks show that expert-authored skills can help while self-generated skills can underperform a no-skill baseline.
Verification capacity becomes the coding-agent bottleneck
Coding agents can parallelize implementation faster than engineering teams can expand review judgment.
AI product feedback needs behavioral traces
For nondeterministic products, edits, retries, overrides, abandonment, and sampled outputs reveal more than isolated ratings.
Databricks benchmarks coding agents on its own codebase
Databricks evaluates agents on fresh internal pull-request tasks and measures success alongside runtime, tokens, and cost.
Newer agent models can reward shorter prompts
OpenAI's migration guidance emphasizes concise instructions, explicit contracts, and less compensating prose for newer models.
Multi-agent systems need continuous eval pipelines
A Google workflow evaluates not only the final answer but also routing, delegation, tool calls, and the trajectory between agents.
Agent memory is being evaluated as system behavior
MemoryData evaluates what an agent stores, retrieves, updates, and forgets across a sequence instead of grading one final answer.
Harvey trained a legal agent with an applied compute loop
Harvey describes domain experts, task environments, evaluation, and iterative training as one system for legal agent performance.
Verifiability is an AI product feature
Hamel Husain's eval-smell framework treats missing traces, weak rubrics, and unverifiable outputs as product defects.
Prompt debt turns natural language into legacy code
System prompts accumulate patches for edge cases until model changes and local fixes make the behavior fragile and opaque.
AI compresses coding, not the whole software lifecycle
Implementation time can collapse while requirements, architecture, validation, review, and maintenance stay constrained by human judgment.
Source ledger
Publishable sources attached to this record.
| # | Source | Role | Public status |
|---|---|---|---|
| 1 | addyo.substack.comsource | primary receipt | source_urls |
| 2 | addyosmani.comsource | supporting receipt | source_urls |
| 3 | anthropic.comsource | supporting receipt | source_urls |
| 4 | arxiv.orgpaper | supporting receipt | source_urls |
| 5 | arxiv.orgpaper | supporting receipt | source_urls |
| 6 | arxiv.orgpaper | supporting receipt | source_urls |
| 7 | databricks.comsource | supporting receipt | source_urls |
| 8 | dbreunig.comsource | supporting receipt | source_urls |
| 9 | developers.openai.comdocs | supporting receipt | source_urls |
| 10 | github.comrepo | supporting receipt | source_urls |
Showing 10 of 22; the complete set is exposed in the JSON route.