In this paper, we propose a novel benchmark called the StarCraft Multi-Agent Challenges+, where agents learn to perform multi-stage tasks and to use environmental factors without precise reward functions. The previous challenges (SMAC) recognized as a standard benchmark of Multi-Agent Reinforcement Learning are mainly concerned with ensuring that all agents cooperatively eliminate approaching adversaries only through fine manipulation with obvious reward functions. This challenge, on the other hand, is interested in the exploration capability of MARL algorithms to efficiently learn implicit multi-stage tasks and environmental factors as well as micro-control. This study covers both offensive and defensive scenarios. In the offensive scenarios, agents must learn to first find opponents and then eliminate them. The defensive scenarios require agents to use topographic features. For example, agents need to position themselves behind protective structures to make it harder for enemies to attack. We investigate MARL algorithms under SMAC+ and observe that recent approaches work well in similar settings to the previous challenges, but misbehave in offensive scenarios. Additionally, we observe that an enhanced exploration approach has a positive effect on performance but is not able to completely solve all scenarios. This study proposes new directions for future research.

本文提出了一个叫做SMAC + 的新型基准，该基准旨在探索MARL算法在StarCraft遊戲中学习隐含的多阶段任务、环境因素和微控制的能力。在攻击和防御场景中，该基准要求智能体进行多方面探索，进一步提高算法的探索能力。研究结果表明，近年来的一些算法在该基准中表现良好，但在攻击场景方面表现不佳，为未来的研究提供了新的方向。

StarCraft多智能体挑战+: 在没有精确奖励函数的情况下学习多阶段任务和环境因素