We consider a non stationary multi-armed bandit in which the population preferences are positively and negatively reinforced by the observed rewards. The objective of the algorithm is to shape the population preferences to maximize the fraction of the population favouring a predetermined arm. For the case of binary opinions, two types of opinion dynamics are considered -- decreasing elasticity (modeled as a Polya urn with increasing number of balls) and constant elasticity (using the voter model). For the first case, we describe an Explore-then-commit policy and a Thompson sampling policy and analyse the regret for each of these policies. We then show that these algorithms and their analyses carry over to the constant elasticity case. We also describe a Thompson sampling based algorithm for the case when more than two types of opinions are present. Finally, we discuss the case where presence of multiple recommendation systems gives rise to a trade-off between their popularity and opinion shaping objectives.

该研究论文探讨了非平稳的多臂赌博机中，通过观察到的奖励来积极和消极地加强人群偏好，算法的目标是塑造人群偏好，从而最大化人群中支持特定臂的比例，提出了不同意见动态模型，包括两种二元意见动态（弹性递减和常数弹性），探讨了不同策略及其遗憾值的分析，针对多于两种意见的情况提出了基于Thompson采样的算法，同时讨论了多个推荐系统存在时受欢迎度和意见塑造目标之间的权衡问题。

影响性强盗：偏好塑造的臂选择