1,865 human-verified, contamination-resistant enterprise-grade GitHub issues across 41 actively-maintained repos (public/held-out/commercial splits); patches average 107 lines across 4.1 files, versus 11.6 lines/1 file on SWE-bench Verified; an agent must produce a patch passing human-reviewed fail2pass/pass2pass test suites.
discriminates
[build-verify-reflect][debate][ensemble]
saturation
open
patternLabFit
Explicitly built to replace SWE-bench Verified once it saturates: with the same build-verify-reflect loop but a much harder ceiling (best model 43.6% Pass@1 on the public set, under 20% on the private commercial set), it gives roundsToFirstPass and error-decorrelation far more headroom to discriminate multi-agent patterns than the near-saturated original.
Distinct from SWE-bench Verified (id: swe-bench-verified): larger/harder tasks (multi-file, 100+ LOC changes common), GPL/held-out/commercial repos chosen specifically to resist training-data contamination, and a much lower current ceiling (sub-45% vs 80%+), so it is not a trivial rename or subset of the tracked seed.