First-person video naturally brings the use of a physical environment to the forefront, since it shows the camera wearer interacting fluidly in a space based on his intentions. However, current methods largely separate the observed actions from the persistent space itself. We introduce a model for environment affordances that is learned directly from egocentric video. The main idea is to gain a human-centric model of a physical space (such as a kitchen) that captures (1) the primary spatial zones of interaction and (2) the likely activities they support. Our approach decomposes a space into a topological map derived from first-person activity, organizing an ego-video into a series of visits to the different zones. Further, we show how to link zones across multiple related environments (e.g., from videos of multiple kitchens) to obtain a consolidated representation of environment functionality. On EPIC-Kitchens and EGTEA+, we demonstrate our approach for learning scene affordances and anticipating future actions in long-form video.

通过学习人类源动作，我们引入了一种从第一人称视频中直接学习物理空间环境能力的模型，该模型将空间分解为基于活动的拓扑图，并展示了如何跨多个相关环境链接区域以获得其功能性的整合表示。我们在实验中展示了学习场景能力预测未来操作的方法。

EGO-TOPO: 从自我中心视频中提取环境能力