Learning to navigate in a visual environment following natural language instructions is a challenging task because natural language instructions are highly variable, ambiguous, and under-specified. In this paper, we present a novel training paradigm, Learn from EveryOne (LEO), which leverages multiple instructions (as different views) for the same trajectory to resolve language ambiguity and improve generalization. By sharing parameters across instructions, our approach learns more effectively from limited training data and generalizes better in unseen environments. On the recent Room-to-Room (R2R) benchmark dataset, LEO achieves 16% improvement (absolute) over a greedy agent as the base agent (25.3% $\rightarrow$ 41.4%) in Success Rate weighted by Path Length (SPL). Further, LEO is complementary to most existing models for vision-and-language navigation, allowing for easy integration with the existing techniques, leading to LEO+, which creates the new state of the art, pushing the R2R benchmark to 62% (9% absolute improvement).

通过利用多条不同视角的指令，共享参数，解决语义歧义和提高广义性，学习在视觉环境中遵循自然语言指令导航的新训练范式 LEarn from EveryOne (LEO) 在 R2R 基准测试数据集上比贪婪代理 (25.3%->41.4%) 提高 16% 的成功率重量化路径长度 (SPL)，并且与大多数现有的视觉和语言导航模型互补，易于与现有技术集成，推动 R2R 基线提升至 62%（绝对提升 9%）的最新技术 LEO+被创造出来。

视觉语言导航的多视图学习