The predictive information, the mutual information between the past and future, has been shown to be a useful representation learning auxiliary loss for training reinforcement learning agents, as the ability to model what will happen next is critical to success on many control tasks. While existing studies are largely restricted to training specialist agents on single-task settings in simulation, in this work, we study modeling the predictive information for robotic agents and its importance for general-purpose agents that are trained to master a large repertoire of diverse skills from large amounts of data. Specifically, we introduce Predictive Information QT-Opt (PI-QT-Opt), a QT-Opt agent augmented with an auxiliary loss that learns representations of the predictive information to solve up to 297 vision-based robot manipulation tasks in simulation and the real world with a single set of parameters. We demonstrate that modeling the predictive information significantly improves success rates on the training tasks and leads to better zero-shot transfer to unseen novel tasks. Finally, we evaluate PI-QT-Opt on real robots, achieving substantial and consistent improvement over QT-Opt in multiple experimental settings of varying environments, skills, and multi-task configurations.

通过引入准确的表示学习机制——Predictive Information QT-Opt（PI-QT-Opt），我们研究了预测信息对机器人智能代理的建模以及其在从大量数据中培养具备各种技能的通用代理方面的重要性。实验结果表明，这种机制的应用能有效地提高任务求解的速度，并实现对无尝试性新任务的更好的转移学习。

PI-QT-Opt: 预测信息提升多任务机器人增强学习规模