- axis
- coding
- verification
- det-test
- whatItTests
- 1,140 practical Python tasks requiring an agent to compose multiple function calls across 139 libraries and 7 domains, following complex natural-language instructions, scored against hidden test suites averaging 99% branch coverage.
- discriminates
[build-verify-reflect]
- saturation
- open
- patternLabFit
- Per-task hidden test suites with near-complete branch coverage make this a strong build-verify-reflect candidate for tool-use-heavy coding (compositional library calls rather than single-function synthesis). The paper evaluates 60 LLMs individually; it does not test multi-agent collaboration.
- notes
- Distinct from swe-bench-verified/pro (issue-resolution patches) and livecodebench (algorithmic/competitive problems): BigCodeBench targets practical library and tool composition (data analysis, web development) rather than bug repair or puzzle-solving.