Competition mathematics (algebra, geometry, number theory, combinatorics) with symbolically-checkable gold answers; the 500-item subset of the Hendrycks MATH benchmark.
discriminates
[debate][build-verify-reflect][ensemble]
saturation
saturated
patternLabFit
The cleanest standard benchmark for debate: one of the few where multi-agent debate reliably beats self-consistency (+4 to 7pp). Gold answers feed build-verify-reflect directly and make roundsToFirstPass measurable.