Many applications, e.g., in shared mobility, require coordinating a large number of agents. Mean-field reinforcement learning addresses the resulting scalability challenge by optimizing the policy of a representative agent. In this paper, we address an important generalization where there exist global constraints on the distribution of agents (e.g., requiring capacity constraints or minimum coverage requirements to be met). We propose Safe-$\text{M}^3$-UCRL, the first model-based algorithm that attains safe policies even in the case of unknown transition dynamics. As a key ingredient, it uses epistemic uncertainty in the transition model within a log-barrier approach to ensure pessimistic constraints satisfaction with high probability. We showcase Safe-$\text{M}^3$-UCRL on the vehicle repositioning problem faced by many shared mobility operators and evaluate its performance through simulations built on Shenzhen taxi trajectory data. Our algorithm effectively meets the demand in critical areas while ensuring service accessibility in regions with low demand.

本研究提出了 Safe-M3-UCRL 算法，使用平均场强化学习来为大量智能体寻找优化方法，并且可以在面临未知转换动态时实现建模优化问题，保证悲观约束条件的满足。在这个基础上，我们以共享代步交通问题为例进行了模拟评估，结果表明，该算法在保证服务可用性的同时，能够有效地维持区域内的供需平衡。

安全的基于模型的多智能体均场强化学习