Reinforcement learning (RL) -- finding the optimal behaviour (also referred
to as policy) maximizing the collected long-term cumulative reward -- is among
the most influential approaches in machine learning with a large number of
successful applications. In several decision problems, however, one faces the
possibility of policy switching -- changing from the current policy to a new
one -- which incurs a non-negligible cost (examples include the shifting of the
currently applied educational technology, modernization of a computing cluster,
and the introduction of a new webpage design), and in the decision one is
limited to using historical data without the availability for further online
interaction. Despite the inevitable importance of this offline learning
scenario, to our best knowledge, very little effort has been made to tackle the
key problem of balancing between the gain and the cost of switching in a
flexible and principled way. Leveraging ideas from the area of optimal
transport, we initialize the systematic study of policy switching in offline
RL. We establish fundamental properties and design a Net Actor-Critic algorithm
for the proposed novel switching formulation. Numerical experiments demonstrate
the efficiency of our approach on multiple benchmarks of the Gymnasium.

采用最优输运的思想，我们对离线强化学习中的政策切换问题进行了系统研究，并设计了一种新颖的切换公式的 Net Actor-Critic 算法，数值实验证实了我们方法在多个 Gymnasium 基准测试上的效率。

离线强化学习中的均衡策略切换：切换还是不切换？

To Switch or Not to Switch? Balanced Policy Switching in Offline  Reinforcement Learning

We study the problem of reinforcement learning (RL) with low (policy)
switching cost - a problem well-motivated by real-life RL applications in which
deployments of new policies are costly and the number of policy updates must be
low. In this paper, we propose a new algorithm based on stage-wise exploration
and adaptive policy elimination that achieves a regret of
$\widetilde{O}(\sqrt{H^4S^2AT})$ while requiring a switching cost of $O(HSA
\log\log T)$. This is an exponential improvement over the best-known switching
cost $O(H^2SA\log T)$ among existing methods with
$\widetilde{O}(\mathrm{poly}(H,S,A)\sqrt{T})$ regret. In the above, $S,A$
denotes the number of states and actions in an $H$-horizon episodic Markov
Decision Process model with unknown transitions, and $T$ is the number of
steps. As a byproduct of our new techniques, we also derive a reward-free
exploration algorithm with a switching cost of $O(HSA)$. Furthermore, we prove
a pair of information-theoretical lower bounds which say that (1) Any no-regret
algorithm must have a switching cost of $\Omega(HSA)$; (2) Any
$\widetilde{O}(\sqrt{T})$ regret algorithm must incur a switching cost of
$\Omega(HSA\log\log T)$. Both our algorithms are thus optimal in their
switching costs.

本文针对实际强化学习应用中新策略部署的高成本和策略更新次数必须较少的问题，提出了一种基于分阶段探索和自适应策略消除算法，实现了在低换乘成本下的回报 并且在已知的换乘成本中实现了指数级的改善。