{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:langchain-reviewbench-review-agent-evals","slug":"langchain-reviewbench-review-agent-evals","title":"ReviewBench Turns Code Review Into An Agent Eval","description":"LangChain's ReviewBench uses real PR review history to test whether code-review agents can recover substantive reviewer findings without flooding humans with noise.","retrieval_nugget":"LangChain's ReviewBench uses real PR review history to test whether code-review agents can recover substantive reviewer findings without flooding humans with noise. LangChain's ReviewBench is interesting because it treats code review as a concrete agent task, not as a vibe around \"AI catches bugs.\" The benchmark starts from real PR feedback in the LangSmith monorepo.","status":"published","published_at":"2026-08-01","updated_at":"2026-08-01","record_date":"2026-08-01","date_kind":"published_at","topics":["coding-agents","evals","verification","developer-tools"],"source_urls":["https://x.com/LangChain/status/2083236117839499511","https://www.langchain.com/blog/towards-automating-eval-engineering","https://www.langchain.com/blog/unified-stack-for-evaluating-agents"],"visuals":[{"id":"langchain-reviewbench-review-agent-evals","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/langchain-reviewbench-review-agent-evals.webp","alt":"Hand-drawn evaluation pipeline where real pull request review comments become curated issue cards, reproducible benchmark tasks, agent findings, and coverage plus precision scores.","caption":"ReviewBench moves code-review agents away from generic bug hunting and toward the exact failures trusted reviewers catch in real pull requests.","credit":"New Runtime synthesis from LangChain ReviewBench public article","source_url":"https://x.com/LangChain/status/2083236117839499511","generated_with":"gemini-3.1-flash-image","width":1600,"height":900,"legend":[{"label":"Real reviews","description":"Merged PR comments become candidate findings instead of synthetic bugs written from scratch."},{"label":"Curated tasks","description":"Substantive review comments are converted into reproducible Harbor tasks with frozen PR context."},{"label":"Review quality","description":"Agents are judged on coverage of baseline issues and precision of submitted findings."}]}],"routes":{"html":"https://newruntime.com/posts/langchain-reviewbench-review-agent-evals/","markdown":"https://newruntime.com/posts/langchain-reviewbench-review-agent-evals.md","json":"https://newruntime.com/posts/langchain-reviewbench-review-agent-evals.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/coding-agents/","reason":"Explore the coding agents topic hub.","url":"https://newruntime.com/topics/coding-agents/","title":"Coding agents - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/developer-tools/","reason":"Explore the developer tools topic hub.","url":"https://newruntime.com/topics/developer-tools/","title":"Developer Tools - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/github-stacked-prs-agent-review/","reason":"Shares coding agents and developer tools.","url":"https://newruntime.com/posts/github-stacked-prs-agent-review/","title":"GitHub Stacked PRs Turn Large Agent Changes Into Reviewable Chains","media_type":"text/html"},{"type":"related_material","path":"/posts/cursor-cloud-agent-environment-product/","reason":"Shares coding agents and developer tools.","url":"https://newruntime.com/posts/cursor-cloud-agent-environment-product/","title":"Cursor Treats the Cloud Agent Environment as the Product","media_type":"text/html"},{"type":"related_material","path":"/posts/agentic-sdlc-software-factory-loop/","reason":"Shares coding agents and evals.","url":"https://newruntime.com/posts/agentic-sdlc-software-factory-loop/","title":"A Software Factory Connects Agents Through Verified Outcomes","media_type":"text/html"}]}
