Reinforcement Learning is divided in two main paradigms: model-free and model-based. Each of these two paradigms has strengths and limitations, and has been successfully applied to real world domains that are appropriate to its corresponding strengths. In this paper, we present a new approach aimed at bridging the gap between these two paradigms. We aim to take the best of the two paradigms and combine them in an approach that is at the same time data-efficient and cost-savvy. We do so by learning a probabilistic dynamics model and leveraging it as a prior for the intertwined model-free optimization. As a result, our approach can exploit the generality and structure of the dynamics model, but is also capable of ignoring its inevitable inaccuracies, by directly incorporating the evidence provided by the direct observation of the cost. As a proof-of-concept, we demonstrate on simulated tasks that our approach outperforms purely model-based and model-free approaches, as well as the approach of simply switching from a model-based to a model-free setting.

本文提出了一种新的方法，旨在将模型自由和模型相关两种范式结合起来，通过学习概率动力学模型和利用它作为模型自由优化的先验概率来实现数据有效和成本节约，并证明这种方法优于单纯的模型相关和模型自由方法，以及从模型相关模式切换到模型自由模式的方法。

MBMF:基于模型的先验知识用于无模型强化学习