We introduce the task of 3D visual grounding in large-scale dynamic scenes
based on natural linguistic descriptions and online captured multi-modal visual
data, including 2D images and 3D LiDAR point clouds. We present a novel method,
WildRefer, for this task by fully utilizing the appearance features in images,
the location and geometry features in point clouds, and the dynamic features in
consecutive input frames to match the semantic features in language. In
particular, we propose two novel datasets, STRefer and LifeRefer, which focus
on large-scale human-centric daily-life scenarios with abundant 3D object and
natural language annotations. Our datasets are significant for the research of
3D visual grounding in the wild and has huge potential to boost the development
of autonomous driving and service robots. Extensive comparisons and ablation
studies illustrate that our method achieves state-of-the-art performance on two
proposed datasets. Code and dataset will be released when the paper is
published.

本研究提出了一种基于自然语言描述和多模式视觉数据的大规模动态场景的 3D 视觉定位任务的方法，并且通过利用图像的外观特征、点云中的位置和几何特征以及连续输入帧中的动态特征，匹配语言中的语义特征。我们提出了两个新的数据集，STRefer 和 LifeRefer，这些数据集对于野外 3D 视觉定位的研究具有重要意义，并且有着提升自动驾驶和服务机器人发展的巨大潜力。广泛的比较和消融研究证明，我们的方法在两个提出的数据集上实现了最先进的性能。