500 challenging open-ended prompts curated automatically (BenchBuilder pipeline from Chatbot Arena/WildChat data) and scored via LLM-as-a-Judge; validated against human preference rankings and model-separation ability rather than gold correctness.
discriminates
[judge-calibration]
saturation
mid
patternLabFit
A methodology reference for validating an LLM-as-judge pipeline's alignment with human preference and its ability to separate close model performances (relevant wherever Pattern Lab uses judge-score to rank near-tied outputs).
Distinct from JudgeBench: JudgeBench checks whether a judge picks the objectively-correct answer on verifiable tasks; Arena-Hard-Auto instead validates a judge pipeline's correlation with human preference and its power to separate models on open-ended prompts.