Simulated multi-turn conversations between an LLM-played user and a tool-using agent in domain-specific settings (e.g. retail), scored by comparing the final database state against an annotated goal state.
discriminates
[build-verify-reflect]
saturation
open
patternLabFit
The goal-state comparison is a deterministic checker an agent could consult mid-conversation, giving a natural build-verify-reflect hook; the paper does not test multi-agent collaboration patterns directly.