Field note
Ramp SWE-Bench is another sign that coding-agent evaluation is moving closer to real production work.
Generic benchmarks are useful, but they do not answer the question a company actually has: can this agent operate inside our codebase, our issue shapes, our review norms, and our failure modes? A company-specific SWE-Bench surface turns that question into a repeatable test.
The New Runtime angle is that internal evals are becoming infrastructure. As coding agents become cheaper to run and easier to plug into repositories, the scarce asset is a high-signal task set that reflects the work the organization really needs done.
