- axis
- methodology
- verification
- det-gold
- whatItTests
- Meta-evaluation of LLM-as-judge methods for mobile GUI agents: 931 human-annotated agent trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps, against which 6 judge methods are scored across multiple LLM backends.
- discriminates
[judge-calibration]
- saturation
- open
- patternLabFit
- Extends the JudgeBench-style judge-calibration check from single-response pairs to full agent trajectories, exactly the shape of artifact a Pattern Lab judge scores; its findings that a simple screenshot-sampling baseline judge often beats purpose-built judge methods, and that benchmark quality metrics correlate with real-world judge utility, are directly actionable for choosing and trusting a trajectory judge.
- notes
- Submitted 2026-08-11, within the discovery window. Distinct from JudgeBench (judges on static response pairs with objectively verifiable answers), Arena-Hard-Auto (judge alignment with human preference on open-ended prompts), and RewardBench 2 (reward models): MobileJudgeBench scores judges against human-annotated success labels on multi-step mobile agent trajectories, i.e. trajectory-level rather than response-level judge calibration. Verification marked det-gold since judges are scored against human-annotated ground-truth trajectory labels.