In many common-payoff games, achieving good performance requires players to develop protocols for communicating their private information implicitly -- i.e., using actions that have non-communicative effects on the environment. Multi-agent reinforcement learning practitioners typically approach this problem using independent learning methods in the hope that agents will learn implicit communication as a byproduct of expected return maximization. Unfortunately, independent learning methods are incapable of doing this in many settings. In this work, we isolate the implicit communication problem by identifying a class of partially observable common-payoff games, which we call implicit referential games, whose difficulty can be attributed to implicit communication. Next, we introduce a principled method based on minimum entropy coupling that leverages the structure of implicit referential games, yielding a new perspective on implicit communication. Lastly, we show that this method can discover performant implicit communication protocols in settings with very large spaces of messages.

本文介绍了一种被称为Markov编码游戏(MCG)的方法来处理在分散控制环境下进行通信的问题，并且介绍了一种新的理论算法MEME来进行最大熵强化学习和最小熵耦合的平衡。同时进行的实验也表明了这种算法在解决小型和大型MCG问题时具有良好性能。

通过马尔可夫决策过程进行通信