1,400+ real Upwork freelance software-engineering tasks (bug fixes to full feature builds) worth $1M total in real-world payouts, plus separate managerial tasks where a model must choose between competing technical implementation proposals.
discriminates
[build-verify-reflect][judge-calibration]
saturation
open
patternLabFit
Independent tasks give a real-money-weighted build-verify-reflect signal via end-to-end tests triple-verified by engineers. The managerial subset, where a model picks between proposals and is scored against the real hired manager's actual choice, is a direct analog to judge-calibration for a decision-making role rather than a generation role.
Distinct from swe-bench-pro (enterprise repos, patch-only): adds real dollar-value weighting per task and a managerial decision-making task type (choosing between proposals) absent from any tracked or other proposed coding entry. Verification marked mixed because independent tasks are det-test (triple-verified end-to-end tests) while managerial tasks are scored against a human ground-truth decision.