Evaluates LLM judges on knowledge, reasoning, math and coding tasks where correctness is objectively verifiable; finds even strong judges are near chance on hard items.
discriminates
[judge-calibration]
saturation
open
patternLabFit
The benchmark to run against any judge used in Pattern Lab's judge-score or creative-tension patterns before trusting its verdicts. Guards the LLM-as-judge verification path.