- axis
- multi-agent
- verification
- mixed
- whatItTests
- Competition-based benchmark using two social-deduction games (Undercover, Chameleon) and three game-theory scenarios (cost-sharing, multi-player Prisoner's Dilemma, Public Good) to quantify judgment, reasoning, deception, self-awareness, cooperation, coordination, and rationality in multi-agent LLM settings.
- discriminates
[role-specialization]
- saturation
- open
- patternLabFit
- The Undercover/Chameleon social-deduction setup isolates deception and self-awareness as separate scored dimensions from cooperation and coordination, a decomposition Pattern Lab could borrow to separate 'trust' failures from 'capability' failures in its collaboration score; the reported PGM enhancement (+37% average across models) is a concrete lift figure for a structured-reasoning intervention.
- notes
- Distinct from MAST (a post-hoc failure taxonomy) and MultiAgentBench (milestone-KPI topology comparison): MAgIC uses social-deduction and game-theory scenarios plus a probabilistic-graphic-modeling (PGM) enhancement layer to score deception/self-awareness/rationality as separate dimensions. Exact scoring mechanism per dimension (deterministic game outcome vs LLM judge) is not fully specified on the project page, hence 'mixed' and moderate confidence.