Image-to-video adaptation seeks to efficiently adapt image models for use in the video domain. Instead of finetuning the entire image backbone, many image-to-video adaptation paradigms use lightweight adapters for temporal modeling on top of the spatial module. However, these attempts are subject to limitations in efficiency and interpretability. In this paper, we propose a novel and efficient image-to-video adaptation strategy from the object-centric perspective. Inspired by human perception, which identifies objects as key components for video understanding, we integrate a proxy task of object discovery into image-to-video transfer learning. Specifically, we adopt slot attention with learnable queries to distill each frame into a compact set of object tokens. These object-centric tokens are then processed through object-time interaction layers to model object state changes across time. Integrated with two novel object-level losses, we demonstrate the feasibility of performing efficient temporal reasoning solely on the compressed object-centric representations for video downstream tasks. Our method achieves state-of-the-art performance with fewer tunable parameters, only 5\% of fully finetuned models and 50\% of efficient tuning methods, on action recognition benchmarks. In addition, our model performs favorably in zero-shot video object segmentation without further retraining or object annotations, proving the effectiveness of object-centric video understanding.

通过采用对象为中心的视角，本文提出了一种新颖高效的图像到视频适应策略。结合可学习查询的槽注意力，将每帧图像压缩为一组紧凑的对象令牌，并通过对象时间交互层建模对象在时间上的状态变化。通过两种新颖的对象级损失，我们的方法在行动识别基准测试上以较少的可调参数（仅为完全微调模型的5％和高效微调方法的50％）达到了最先进的性能。此外，我们的模型在零样本视频对象分割中表现良好，无需进一步的重新训练或对象注释，证明了对象为中心的视频理解的有效性。

重新思考图像到视频的适应：一个以物体为中心的视角