- axis
- multi-agent
- verification
- det-test
- whatItTests
- JAX-based, open-ended multi-agent coordination benchmark built on Craftax-like long-horizon survival-world dynamics (exploration, crafting, trading, combat), embedding procedurally generated coordination tasks, soft specialisation, communication, and a controllable coordination-difficulty setting; scored via normalised return.
- discriminates
[role-specialization]
- saturation
- open
- patternLabFit
- Alem's controllable coordination-difficulty knob and soft-specialisation roles give a long-horizon, open-ended testbed for Pattern Lab's dynamic-role and role-specialization patterns beyond short structured tasks; its finding that strong individual-task competence does not imply coordination competence is a direct caution against relying on passRate alone as a Pattern Lab collaboration-quality proxy.
- notes
- Distinct from all three tracked multi-agent entries: Alem is a long-horizon, procedurally generated survival-world benchmark rather than a short structured scenario (Sotopia-style) or staged difficulty ladder (BattleAgentBench), and evaluates 13 current frontier models (Gemini-3.1-Pro-High, GPT-5.4-High) zero-shot against trained MARL reference agents, finding LLM teams average only ~6% normalised return. Published June 2026, within the discovery window's recency band.