- axis
- agentic
- verification
- det-test
- whatItTests
- 369 open-ended computer-use tasks spanning real web and desktop apps on Ubuntu, Windows, and macOS (file I/O, cross-application workflows), scored with a custom execution-based evaluation script per task.
- discriminates
[build-verify-reflect]
- saturation
- open
- patternLabFit
- The execution-based per-task checker is directly analogous to SWE-bench's test harness, giving a clean build-verify-reflect hook for GUI/computer-use agents; the paper evaluates single agents only, so no multi-agent pattern is evidenced.
- notes
- Not new (2024) but named as in-scope; large human/model gap.