{"schema_version":"newruntime-agent-readable-v0.2","type":"post","stable_id":"post:deepsecbench-security-agent-economics","slug":"deepsecbench-security-agent-economics","title":"DeepsecBench Makes Security Agents a Cost/Recall Tradeoff","description":"Vercel's DeepsecBench reframes security-agent evaluation around recall, precision, cost, total scan time, and a hidden benchmark that resists training leakage.","retrieval_nugget":"Vercel's DeepsecBench reframes security-agent evaluation around recall, precision, cost, total scan time, and a hidden benchmark that resists training leakage. Vercel's DeepsecBench is useful because it refuses to make security-agent evaluation a simple leaderboard story. The benchmark measures models on vulnerability discovery in application code, but the report includes recall, precision, cost, and total time.","status":"published","published_at":"2026-07-29","updated_at":"2026-07-29","record_date":"2026-07-29","date_kind":"published_at","topics":["security","evals","models","agent-economics"],"source_urls":["https://x.com/vercel/status/2081846100173177313","https://vercel.com/blog/deepsecbench-evaluating-model-performance-in-finding-cybersecurity-6O29ShaGQlHGINBsMEJSzR/21706b5aba"],"visuals":[{"id":"deepsecbench-security-agent-economics-nano-banana","kind":"editorial-diagram","role":"hero","src":"https://newruntime.com/images/posts/deepsecbench-security-agent-economics-nano-banana.webp","alt":"Hand-drawn benchmark funnel showing a hidden repository and 50 entry files going through model scans, judge review, F2 score, cost and time, then model policy.","caption":"DeepsecBench turns defensive scanning into a portfolio choice across recall, precision, cost, and elapsed time.","credit":"New Runtime synthesis from public source inspection","source_url":"https://vercel.com/blog/deepsecbench-evaluating-model-performance-in-finding-cybersecurity-6O29ShaGQlHGINBsMEJSzR/21706b5aba","generated_with":"nano-banana-style-imagegen","width":1600,"height":900,"legend":[{"label":"Golden set","description":"The benchmark uses 231 human-judged findings from a hidden pre-fix codebase."},{"label":"F2 score","description":"Recall is weighted twice as much as precision because missed vulnerabilities stay unfixed."},{"label":"Policy","description":"Teams can choose model mixes and scan frequency from score, cost, and time."}]}],"routes":{"html":"https://newruntime.com/posts/deepsecbench-security-agent-economics/","markdown":"https://newruntime.com/posts/deepsecbench-security-agent-economics.md","json":"https://newruntime.com/posts/deepsecbench-security-agent-economics.json"},"source_format":"markdown","next_reads":[{"type":"topic","path":"/topics/agent-economics/","reason":"Explore the agent economics topic hub.","url":"https://newruntime.com/topics/agent-economics/","title":"Agent economics - New Runtime","media_type":"text/html"},{"type":"topic","path":"/topics/evals/","reason":"Explore the evals topic hub.","url":"https://newruntime.com/topics/evals/","title":"Agent evals - New Runtime","media_type":"text/html"},{"type":"related_material","path":"/posts/anthropic-claude-cryptographic-weaknesses/","reason":"Shares evals and security.","url":"https://newruntime.com/posts/anthropic-claude-cryptographic-weaknesses/","title":"Claude Mythos Moves Cryptanalysis Into the Verification Bottleneck","media_type":"text/html"},{"type":"related_material","path":"/posts/cline-recursive-self-improvement-coding-agent/","reason":"Shares agent economics and evals.","url":"https://newruntime.com/posts/cline-recursive-self-improvement-coding-agent/","title":"Cline Turns Recursive Self-Improvement Into Harness Work","media_type":"text/html"},{"type":"related_material","path":"/posts/openai-arc-agi-settings-harness/","reason":"Shares evals and models.","url":"https://newruntime.com/posts/openai-arc-agi-settings-harness/","title":"OpenAI's ARC-AGI-3 Jump Was a Harness Result","media_type":"text/html"}]}
