Beyond specific settings, many multi-agent learning algorithms fail to converge to an equilibrium solution, and instead display complex, non-stationary behaviours such as recurrent or chaotic orbits. In fact, recent literature suggests that such complex behaviours are likely to occur when the number of agents increases. In this paper, we study Q-learning dynamics in network polymatrix games where the network structure is drawn from classical random graph models. In particular, we focus on the Erdos-Renyi model, a well-studied model for social networks, and the Stochastic Block model, which generalizes the above by accounting for community structures within the network. In each setting, we establish sufficient conditions under which the agents' joint strategies converge to a unique equilibrium. We investigate how this condition depends on the exploration rates, payoff matrices and, crucially, the sparsity of the network. Finally, we validate our theoretical findings through numerical simulations and demonstrate that convergence can be reliably achieved in many-agent systems, provided network sparsity is controlled.

本研究解决了多智能体学习算法在多元环境中可能无法收敛到均衡解的问题，尤其是在代理数量增加时表现出的复杂非平稳行为。通过研究随机图模型下的Q学习动态，本文提出了一种新的条件，阐明了探索率、收益矩阵及网络稀疏性对智能体策略收敛的影响。在控制网络稀疏性的情况下，研究表明在多智能体系统中能够实现可靠的收敛。

随机网络中的多智能体Q学习动态：由于探索和稀疏性而导致的收敛