- axis
- coding
- verification
- det-test
- whatItTests
- 10 frozen research repositories spanning 10 training-algorithm families, where an agent must improve the repository's training algorithm; the agent's code is rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent against the repository's original algorithm under the same procedure.
- saturation
- open
- patternLabFit
- The rerun-from-scratch, hidden-evaluator protocol is a strict outcome oracle for research-code improvement, and the finding that most agent systems never change how the model learns at all (while the minority that do score 0.226 versus 0.126) is a sharp probe of whether an agent pipeline produces substantive changes or cosmetic ones, relevant to judging depth-of-contribution in Pattern Lab producer roles. Single-agent-system evaluation only, so no collaboration discriminator is asserted.
- notes
- Submitted 2026-08-20, within the discovery window. Distinct from KernelBench (single GPU-kernel speedup oracle) and the SWE-bench family (issue patches): AI4AI-Bench targets end-to-end training-algorithm redesign inside frozen research repos scored by long-horizon reruns, with an explicit recursive-self-improvement framing. Abstract states the task suite, evaluators, and every scored submission are released; no repo URL is given on the arXiv page. Axis placed as coding (research-code modification) rather than agentic; flagged for the assembler's judgment.