Although Deep Reinforcement Learning (DRL) has been popular in many disciplines including robotics, state-of-the-art DRL algorithms still struggle to learn long-horizon, multi-step and sparse reward tasks, such as stacking several blocks given only a task-completion reward signal. To improve learning efficiency for such tasks, this paper proposes a DRL exploration technique, termed A^2, which integrates two components inspired by human experiences: Abstract demonstrations and Adaptive exploration. A^2 starts by decomposing a complex task into subtasks, and then provides the correct orders of subtasks to learn. During training, the agent explores the environment adaptively, acting more deterministically for well-mastered subtasks and more stochastically for ill-learnt subtasks. Ablation and comparative experiments are conducted on several grid-world tasks and three robotic manipulation tasks. We demonstrate that A^2 can aid popular DRL algorithms (DQN, DDPG, and SAC) to learn more efficiently and stably in these environments.

本文提出了一种DRL探索技术A^2，通过将复杂任务分解成子任务、提供正确的子任务顺序以及自适应探索环境的方式，改善了学习效率，实验表明在多个任务中，A^2有助于DQN、DDPG和SAC等普通DRL算法在这些环境中更高效、更稳定地学习。

高效稳定的多步稀疏奖励强化学习的抽象演示和自适应探索