- axis
- methodology
- verification
- mixed
- whatItTests
- A statistical methodology that measures a 'noise floor' via configuration-equivalent paired trials on the same model, then checks whether a reported multi-agent coordination gain exceeds that floor; proposes 'coordination-active pass^k' (restricted to trials where the coordination mechanism actually engages) as a minimum reporting standard.
- discriminates
[producer-agreement]
- saturation
- n-a
- patternLabFit
- Directly answers the question Pattern Lab's own metrics need before trusting a reported lift: is a coordination-pattern improvement bigger than the paired noise floor, or within normal run-to-run variance. Its coordination-active pass^k metric is a candidate minimum-reporting bar for any Pattern Lab experiment claiming a multi-agent-pattern gain, and pairs naturally with roundsToFirstPass and producerAgreement.
- notes
- Distinct from JudgeBench/Arena-Hard-Auto (which validate LLM judges, not coordination-gain claims) and from MAST (a human-annotated failure taxonomy): this is a statistical-rigor methodology paper testing whether reported multi-agent coordination gains exceed measurement noise across ten recent architectures. Very recent (June 2026); could not retrieve the full verbatim abstract due to a fetch-tool length limit, but the quoted finding sentence and the 'coordination-active pass^k' terminology were independently confirmed.