- axis
- agentic
- verification
- det-gold
- whatItTests
- 327 retail-domain tasks across a large ecosystem of 1,665 tools, requiring an agent to iteratively discover relevant tools, invoke them to surface intermediate evidence, and reach a final answer; includes an optional blocking mode injecting missing, failing, or distracting tools to test recovery.
- discriminates
[build-verify-reflect]
- saturation
- open
- patternLabFit
- Datatype-checked final answers plus explicit failure-injection (missing/failing/distracting tools) make this a strong build-verify-reflect and error-recovery testbed; no multi-agent collaboration mechanism is described in the paper.
- notes
- Published 2026-06-22, within the discovery window; distinct from τ-bench by testing tool-discovery-at-scale and induced tool failure rather than user-simulated dialogue.