- axis
- multi-agent
- verification
- mixed
- whatItTests
- Fine-grained evaluation across seven sub-stages of three difficulty levels, testing single-agent navigation, paired-agent task execution, and multi-agent collaboration AND competition capability, evaluated on four closed-source and seven open-source models.
- saturation
- open
- patternLabFit
- The staged single-to-paired-to-multi-agent difficulty progression is a candidate structure for graduated Pattern Lab collaboration tests that isolate where cooperation breaks down as agent count and adversarial pressure increase.
- notes
- Distinct from MultiAgentBench: BattleAgentBench uses a fixed seven-sub-stage, three-difficulty-level progression (single agent -> paired agent -> multi-agent) with explicit competitive scenarios, rather than milestone KPIs across coordination topologies. discriminates left empty: abstract does not detail a specific collaboration-pattern mechanism (e.g. debate, reputation-elo) beyond general cooperation/competition capability.