We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method, WildRefer, for this task by fully utilizing the appearance features in images, the location and geometry features in point clouds, and the dynamic features in consecutive input frames to match the semantic features in language. In particular, we propose two novel datasets, STRefer and LifeRefer, which focus on large-scale human-centric daily-life scenarios with abundant 3D object and natural language annotations. Our datasets are significant for the research of 3D visual grounding in the wild and has huge potential to boost the development of autonomous driving and service robots. Extensive comparisons and ablation studies illustrate that our method achieves state-of-the-art performance on two proposed datasets. Code and dataset will be released when the paper is published.

本研究提出了一种基于自然语言描述和多模式视觉数据的大规模动态场景的3D视觉定位任务的方法，并且通过利用图像的外观特征、点云中的位置和几何特征以及连续输入帧中的动态特征，匹配语言中的语义特征。我们提出了两个新的数据集，STRefer和LifeRefer，这些数据集对于野外3D视觉定位的研究具有重要意义，并且有着提升自动驾驶和服务机器人发展的巨大潜力。广泛的比较和消融研究证明，我们的方法在两个提出的数据集上实现了最先进的性能。

WildRefer: 基于多模态视觉数据和自然语言的大规模动态场景中的3D物体定位