Deep reinforcement learning is a promising approach to training a dialog
manager, but current methods struggle with the large state and action spaces of
multi-domain dialog systems. Building upon Deep Q-learning from Demonstrations
(DQfD), an algorithm that scores highly in difficult Atari games, we leverage
dialog data to guide the agent to successfully respond to a user's requests. We
make progressively fewer assumptions about the data needed, using labeled,
reduced-labeled, and even unlabeled data to train expert demonstrators. We
introduce Reinforced Fine-tune Learning, an extension to DQfD, enabling us to
overcome the domain gap between the datasets and the environment. Experiments
in a challenging multi-domain dialog system framework validate our approaches,
and get high success rates even when trained on out-of-domain data.

本研究提出一种基于 Deep Q-learning from Demonstrations 的 Reinforced Fine-tune Learning 方法，利用 labeled、reduced-labeled 和 unlabeled data 训练 expert demonstrators，以解决多领域对话系统中 state 和 action 空间较大的问题，并在实验中取得了较高的成功率。