When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL: Effective communication in...
Multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I pr...
来源简报
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I pr...