- axis
- agentic
- verification
- mixed
- whatItTests
- An unsaturated, holistic, text-only benchmark that chains two to thirteen single-domain subproblems (visual reasoning, coding, math, information extraction/web search, problem-solving, general knowledge, data analysis) into composite challenges answered within a single prompt, with added complexity from prompt encoding and deliberate context bloat; models may freely use code execution and web search.
- saturation
- open
- patternLabFit
- Its composite, cross-domain chaining of subproblems (2-13 per item) is a candidate stress test for whether a Pattern Lab pipeline's per-domain specialist roles actually compose correctly end to end, rather than measuring any single skill in isolation; the paper evaluates single models/agents only, so no multi-agent collaboration mechanism is evidenced.
- notes
- Submitted 2026-07-20 (within the discovery window). Distinct from every tracked entry: unlike GAIA/BrowseComp (single-answer lookup) or PolyWorkBench (multilingual long-horizon workplace tasks), Relay-Bench's defining mechanism is forced composition of many single-domain subproblems (up to 13) inside one prompt, with explicit context-bloat and prompt-encoding stressors. Verification marked 'mixed' rather than a specific kind: the abstract does not state the grading methodology (exact-match per subproblem vs an aggregate judge), and no fetched section confirmed this, so it is left unresolved per the no-fabrication policy rather than asserted as det-gold.