- axis
- multi-agent
- verification
- mixed
- whatItTests
- Open-ended social-simulation benchmark where LLM agents role-play 40 characters with private goals, secrets, and relationships across scenarios spanning negotiation, collaboration, and competition; scored on a 7-dimension holistic rubric (Sotopia-Eval).
- discriminates
[role-specialization]
- saturation
- open
- patternLabFit
- Sotopia-Eval's seven-dimension holistic scoring (goal completion, relationship, social rules, believability) is a template for judging Pattern Lab's role-play-heavy games beyond simple task success; its published GPT-4-vs-human Pearson correlations per dimension (0.71 on goal completion down to 0.22 on secret-keeping) are a ready calibration check for any Pattern Lab judge scoring character-driven collaboration.
- notes
- Distinct from MultiAgentBench (milestone-KPI topology comparison) and BattleAgentBench (staged single-to-multi-agent difficulty ladder): Sotopia is role-play-based social simulation scored on a 7-dimension rubric with reported human-correlation figures per dimension. ICLR 2024, widely cited, missing from the tracked 16.