Loading...
Mutual referencing among debating agents is modelled as an explicit, mutable graph of Debate Relationships. A Selection RL-Agent dynamically rewires which peers each agent attends to each round based on measured group-level consensus/divergence evidence, and a Behavior RL-Agent adapts each agent's generation strategy in response, the two optimised jointly as a sequential multi-agent RL problem.
Source: https://arxiv.org/abs/2608.03648