Model-based planners and controllers are commonly used to solve complex manipulation problems as they can efficiently optimize diverse objectives and generalize to long horizon tasks. However, they are limited by the fidelity of their model which oftentimes leads to failures during deployment. To enable a robot to recover from such failures, we propose to use hierarchical reinforcement learning to learn a separate recovery policy. The recovery policy is triggered when a failure is detected based on sensory observations and seeks to take the robot to a state from which it can complete the task using the nominal model-based controllers. Our approach, called RecoveryChaining, uses a hybrid action space, where the model-based controllers are provided as additional \emph{nominal} options which allows the recovery policy to decide how to recover, when to switch to a nominal controller and which controller to switch to even with \emph{sparse rewards}. We evaluate our approach in three multi-step manipulation tasks with sparse rewards, where it learns significantly more robust recovery policies than those learned by baselines. Finally, we successfully transfer recovery policies learned in simulation to a physical robot to demonstrate the feasibility of sim-to-real transfer with our method.

本研究针对现有模型在复杂操作中的失效问题，提出了一种新颖的层次强化学习方法，旨在学习独立的恢复策略。当检测到失败时，该策略可以使机器人恢复至可完成任务的状态。研究结果表明，相比于基线方法，该方法在学习恢复策略方面显著提高了鲁棒性，并成功实现了从仿真到物理机器人的转移。

RecoveryChaining：学习本地恢复策略以实现稳健操作