Applying reinforcement learning (RL) to combinatorial optimization problems is attractive as it removes the need for expert knowledge or pre-solved instances. However, it is unrealistic to expect an agent to solve these (often NP-)hard problems in a single shot at inference due to their inherent complexity. Thus, leading approaches often implement additional search strategies, from stochastic sampling and beam-search to explicit fine-tuning. In this paper, we argue for the benefits of learning a population of complementary policies, which can be simultaneously rolled out at inference. To this end, we introduce Poppy, a simple theoretically grounded training procedure for populations. Instead of relying on a predefined or hand-crafted notion of diversity, Poppy induces an unsupervised specialization targeted solely at maximizing the performance of the population. We show that Poppy produces a set of complementary policies, and obtains state-of-the-art RL results on three popular NP-hard problems: the traveling salesman (TSP), the capacitated vehicle routing (CVRP), and 0-1 knapsack (KP) problems. On TSP specifically, Poppy outperforms the previous state-of-the-art, dividing the optimality gap by 5 while reducing the inference time by more than an order of magnitude.

通过引入基于Population的强化学习思想，由于其在最大化性能时尚未预定义特定的多样性，证明了该方法产生一组互补的策略，并在三个著名的NP-hard问题上获得最新的强化学习结果：旅行推销员问题(TSP)，分配式车辆路径规划问题(CVRP)和01背包问题(KP)。在特定的TSP问题上，其超过先前的最先进技术，将最优性差距分为5个，同时缩短了推理时间超过一个数量级。

基于人群的组合优化强化学习