Our objective in this work is video-text retrieval - in particular a joint
embedding that enables efficient text-to-video retrieval. The challenges in
this area include the design of the visual architecture and the nature of the
training data, in that the available large scale video-text training datasets,
such as HowTo100M, are noisy and hence competitive performance is achieved only
at scale through large amounts of compute. We address both these challenges in
this paper. We propose an end-to-end trainable model that is designed to take
advantage of both large-scale image and video captioning datasets. Our model is
an adaptation and extension of the recent ViT and Timesformer architectures,
and consists of attention in both space and time. The model is flexible and can
be trained on both image and video text datasets, either independently or in
conjunction. It is trained with a curriculum learning schedule that begins by
treating images as 'frozen' snapshots of video, and then gradually learns to
attend to increasing temporal context when trained on video datasets. We also
provide a new video-text pretraining dataset WebVid-2M, comprised of over two
million videos with weak captions scraped from the internet. Despite training
on datasets that are an order of magnitude smaller, we show that this approach
yields state-of-the-art results on standard downstream video-retrieval
benchmarks including MSR-VTT, MSVD, DiDeMo and LSMDC.

本研究目标是视频文本检索 - 特别是一种联合嵌入，可以实现高效的文本到视频检索。作者们提出了一种端到端可训练的模型，旨在利用大规模的图像和视频字幕数据集。该模型是近期 ViT 和 Timesformer 框架的改进扩展，包括时间和空间方面的注意力机制。通过训练 WebVid-2M 数据集，作者们表明这种方法在标准下游的视频检索基准测试中取得了最先进的结果。