- axis
- coding
- verification
- det-test
- whatItTests
- 105 Python code-editing challenges (210 total problems, since each pairs a descriptive and a lazy instruction), classified into three change kinds (corrective, perfective, adaptive), scored pass@1 against a hidden test suite.
- saturation
- open
- patternLabFit
- The descriptive-vs-lazy instruction pairing per task is a ready-made probe for whether a clarifying or role-specialized agent improves handling of under-specified instructions, but the paper itself evaluates single models on single-shot edits, not a multi-agent or iterative-repair mechanism, so no discriminator is asserted.
- notes
- Distinct from BigCodeBench/LiveCodeBench (generation from scratch) and the SWE-bench family (repo-scale patches): scoped narrowly to instruction-conditioned edits of a given snippet, with an explicit descriptive-vs-lazy instruction-phrasing axis not present elsewhere in this catalogue.