- axis
- agentic
- verification
- det-test
- whatItTests
- 65 agentic tasks modeled on employees following company handbooks: an agent operates in a self-contained company environment (file workspace plus mock email, chat, calendar, issue-tracking, and commerce services over MCP) and must carry out routine professional work governed by an expert-written 20-124 page standard operating procedure, spanning five domains (finance, medical billing, insurance, logistics, HR) and 10 fictional companies whose specific rules and thresholds are altered per task to resist memorization; graded by a rubric of 824 programmatic criteria checking both required and prohibited actions.
- saturation
- open
- patternLabFit
- Tests whether a standing policy document actually constrains agent behavior over an extended tool-use horizon rather than just measuring task completion, directly relevant to any Pattern Lab pattern that hands an agent a persistent rulebook/persona brief; its named failure mode (agent performs a required check, then acts against its own result) is a sharp discriminator for whether build-verify-reflect loops are genuinely load-bearing or decorative. The paper evaluates single agents only, so no multi-agent collaboration mechanism is evidenced.
- notes
- Submitted 2026-07-28 (within the discovery window). Distinct from tau-bench (user-simulated dialogue against a goal database state), PlanBench-XL (tool-discovery-at-scale with injected tool failures), and PolyWorkBench (multilingual long-horizon workplace tasks): HANDBOOK.md's defining feature is a long, binding, per-task-mutated policy document (20-124 pages) that the agent must obey over an extended MCP-tool-use horizon, graded against an 824-criterion deterministic rubric rather than a single goal-state or exact-match answer.