Multi-agent reinforcement learning (MARL) algorithms often struggle to find strategies close to Pareto optimal Nash Equilibrium, owing largely to the lack of efficient exploration. The problem is exacerbated in sparse-reward settings, caused by the larger variance exhibited in policy learning. This paper introduces MESA, a novel meta-exploration method for cooperative multi-agent learning. It learns to explore by first identifying the agents' high-rewarding joint state-action subspace from training tasks and then learning a set of diverse exploration policies to "cover" the subspace. These trained exploration policies can be integrated with any off-policy MARL algorithm for test-time tasks. We first showcase MESA's advantage in a multi-step matrix game. Furthermore, experiments show that with learned exploration policies, MESA achieves significantly better performance in sparse-reward tasks in several multi-agent particle environments and multi-agent MuJoCo environments, and exhibits the ability to generalize to more challenging tasks at test time.

MESA 是一种新颖的元探索方法，通过从训练任务中识别代理的高奖励联合状态-动作子空间，然后学习一组多样性的探索策略来解决多智能体协同学习中有效探索的问题。实验证明，通过学习到的探索策略，MESA 在稀疏奖励环境和挑战性任务中均能显著提高性能，并具备在测试时泛化到更复杂任务的能力。

MESA：基于状态动作空间结构的多智能体学习中的合作元探索