- axis
- agentic
- verification
- mixed
- whatItTests
- Web-agent task execution unified with instructional guide-writing over rendered screenshots (not DOM/accessibility trees), with two grounding schemes (Set-of-Mark element selection and raw pixel coordinates); scored on joint action-success and guide-quality metrics via LLM-assisted annotation, human verification, and live-environment evaluation.
- saturation
- open
- patternLabFit
- A computer-use-adjacent benchmark testing whether an agent can both act correctly AND externalize the action sequence as a human-readable guide; relevant wherever a Pattern Lab producer must document its process for a critic/reviewer, though the paper evaluates single agents, not a producer-critic multi-agent loop, so no discriminator is asserted.
- notes
- Distinct from OSWorld/BrowseComp (execution-only) by requiring the agent to also produce a step-by-step natural-language guide alongside acting, scored jointly; grounded in rendered screenshots rather than DOM/accessibility-tree text; scoped to Digital-Adoption-Platform-style web tasks rather than full desktop/OS control. Even the strongest model completes fewer than 40% of tasks per the abstract.