Offline reinforcement learning (RL) is a challenging setting where existing
off-policy actor-critic methods perform poorly due to the overestimation of
out-of-distribution state-action pairs. Thus, various additional augmentations
are proposed to keep the learned policy close to the offline dataset (or the
behavior policy). In this work, starting from the analysis of offline monotonic
policy improvement, we get a surprising finding that some online on-policy
algorithms are naturally able to solve offline RL. Specifically, the inherent
conservatism of these on-policy algorithms is exactly what the offline RL
method needs to overcome the overestimation. Based on this, we propose Behavior
Proximal Policy Optimization (BPPO), which solves offline RL without any extra
constraint or regularization introduced compared to PPO. Extensive experiments
on the D4RL benchmark indicate this extremely succinct method outperforms
state-of-the-art offline RL algorithms. Our implementation is available at
this https URL

本文通过对线下单调策略改进的分析得出有趣结论，即一些在线策略算法天生就能解决离线 RL 问题，而 Behavior Proximal Policy Optimization (BPPO) 正是基于这个结论提出的，无需额外约束或正则化就能在 D4RL 基准测试中超越最先进的线下 RL 算法。