- axis
- methodology
- verification
- mixed
- whatItTests
- An evaluation suite for testing, reliability, and observability of LLM-based multi-agent systems: standardizes MAS configuration/execution behind a unified interface, integrates native and third-party MAS via adapters, and exports framework-agnostic execution traces plus system signals (latency, cost, failures); instantiated across 12 representative MAS.
- discriminates
[producer-agreement]
- saturation
- n-a
- patternLabFit
- MAESTRO's framework-agnostic execution traces plus system signals (latency, cost, run-to-run variance) is a template for instrumenting Pattern Lab experiments so that a coordination-topology comparison isn't confounded by architecture-driven variance, which the paper shows dominates model-level differences in cost-latency-accuracy trade-offs.
- notes
- Distinct from MAST (human-annotated failure taxonomy) and the tracked judge/reward methodology entries (JudgeBench, Arena-Hard-Auto, RewardBench 2, Chatbot Arena, all of which validate judges or preference rankings): MAESTRO is an execution/observability harness across 12 real multi-agent frameworks measuring reproducibility and architecture-driven resource/variance trade-offs, not an accuracy or preference benchmark. Published January 2026, open-source at github.com/sands-lab/maestro.