Learning high-quality video representation has shown significant applications
in computer vision and remains challenging. Previous work based on mask
autoencoders such as ImageMAE and VideoMAE has proven the effectiveness of
learning representations in images and videos through reconstruction strategy
in the visual modality. However, these models exhibit inherent limitations,
particularly in scenarios where extracting features solely from the visual
modality proves challenging, such as when dealing with low-resolution and
blurry original videos. Based on this, we propose AV-MaskEnhancer for learning
high-quality video representation by combining visual and audio information.
Our approach addresses the challenge by demonstrating the complementary nature
of audio and video features in cross-modality content. Moreover, our result of
the video classification task on the UCF101 dataset outperforms the existing
work and reaches the state-of-the-art, with a top-1 accuracy of 98.8% and a
top-5 accuracy of 99.9%.

通过结合视听信息，我们提出了 AV-MaskEnhancer 方法来学习高质量的视频表示，解决了从低分辨率和模糊的原始视频中提取特征的挑战，并在 UCF101 数据集上的视频分类任务中取得了 98.8% 的 top-1 准确率和 99.9% 的 top-5 准确率，超越了现有工作并达到了最先进水平。

AV-MaskEnhancer：通过音频 - 视觉蒙版自编码器增强视频表达

AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual  Masked Autoencoder

This study explores the application of self-supervised learning (SSL) to the
task of motion forecasting, an area that has not yet been extensively
investigated despite the widespread success of SSL in computer vision and
natural language processing. To address this gap, we introduce Forecast-MAE, an
extension of the mask autoencoders framework that is specifically designed for
self-supervised learning of the motion forecasting task. Our approach includes
a novel masking strategy that leverages the strong interconnections between
agents' trajectories and road networks, involving complementary masking of
agents' future or history trajectories and random masking of lane segments. Our
experiments on the challenging Argoverse 2 motion forecasting benchmark show
that Forecast-MAE, which utilizes standard Transformer blocks with minimal
inductive bias, achieves competitive performance compared to state-of-the-art
methods that rely on supervised learning and sophisticated designs. Moreover,
it outperforms the previous self-supervised learning method by a significant
margin. Code is available at this https URL

通过引入 Forecast-MAE，一种专为自我监督学习运动预测任务设计的掩模自编码器框架的扩展，利用标准 Transformer 块以及最小的内在偏差，我们在具有挑战性的 Argoverse 2 运动预测基准测试上进行的实验表明，Forecast-MAE 取得了与依赖于监督学习和复杂设计的最先进方法竞争性的性能，并且明显优于以前的自我监督学习方法。