Real-world decision-making problems are usually accompanied by delayed rewards, which affects the sample efficiency of Reinforcement Learning, especially in the extremely delayed case where the only feedback is the episodic reward obtained at the end of an episode. Episodic return decomposition is a promising way to deal with the episodic-reward setting. Several corresponding algorithms have shown remarkable effectiveness of the learned step-wise proxy rewards from return decomposition. However, these existing methods lack either attribution or representation capacity, leading to inefficient decomposition in the case of long-term episodes. In this paper, we propose a novel episodic return decomposition method called Diaster (Difference of implicitly assigned sub-trajectory reward). Diaster decomposes any episodic reward into credits of two divided sub-trajectories at any cut point, and the step-wise proxy rewards come from differences in expectation. We theoretically and empirically verify that the decomposed proxy reward function can guide the policy to be nearly optimal. Experimental results show that our method outperforms previous state-of-the-art methods in terms of both sample efficiency and performance.

我们提出了一种名为Diaster（隐式分配子轨道奖励差异）的新的分解方法，将任何情节奖励分解为两个分割点处的两个子轨迹的学分，并且步骤性代理奖励来自期望的差异。我们在理论和实证上验证了分解后的代理奖励函数可以使策略趋近于最优。实验结果表明，我们的方法在样本效率和性能方面优于先前的最新方法。

通过隐含分配子轨迹奖励差异进行情节回归分解