Applying probabilistic models to reinforcement learning (RL) has become an exciting direction of research owing to powerful optimisation tools such as variational inference becoming applicable to RL. However, due to their formulation, existing inference frameworks and their algorithms pose significant challenges for learning optimal policies, for example, the absence of mode capturing behaviour in pseudo-likelihood methods and difficulties in optimisation of learning objective in maximum entropy RL based approaches. We propose VIREL, a novel, theoretically grounded probabilistic inference framework for RL that utilises the action-value function in a parametrised form to capture future dynamics of the underlying Markov decision process. Owing to it's generality, our framework lends itself to current advances in variational inference. Applying the variational expectation-maximisation algorithm to our framework, we show that actor-critic algorithm can be reduced to expectation-maximization. We derive a family of methods from our framework, including state-of-the-art methods based on soft value functions. We evaluate two actor-critic algorithms derived from this family, which perform on par with soft actor critic, demonstrating that our framework offers a promising perspective on RL as inference.

提出一种新的基于概率模型的强化学习方法VIREL，通过应用参数化的动作值函数来总结底层MDP系统的未来动态，使VIREL具有KL散度的寻找峰值形式、自然地从推断中学习确定性最佳策略的能力和分别优化价值函数和策略的能力。通过对VIREL应用变分期望最大化方法，我们表明可以将Actor-critic算法简化为期望最大化，其中策略改进对应E步骤，策略评估对应M步骤，最后，我们展示了来自这个家族的Actor-critic算法在几个领域优于基于软值函数的最新方法。

VIREL：一种变分推断框架的强化学习