Unsupervised domain adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. The most recent UDA methods always resort to adversarial training to yield state-of-the-art results and a dominant number of existing UDA methods employ convolutional neural networks (CNNs) as feature extractors to learn domain invariant features. Vision transformer (ViT) has attracted tremendous attention since its emergence and has been widely used in various computer vision tasks, such as image classification, object detection, and semantic segmentation, yet its potential in adversarial domain adaptation has never been investigated. In this paper, we fill this gap by employing the ViT as the feature extractor in adversarial domain adaptation. Moreover, we empirically demonstrate that ViT can be a plug-and-play component in adversarial domain adaptation, which means directly replacing the CNN-based feature extractor in existing UDA methods with the ViT-based feature extractor can easily obtain performance improvement. The code is available at https://github.com/LluckyYH/VT-ADA.

无监督领域自适应方法已取得了最新的优秀成果，并大多数使用了卷积神经网络作为特征提取器，然而视觉转换器虽已被广泛应用于计算机视觉任务，却未在对抗性领域自适应中研究过，本文填补了这一空白，并证明视觉转换器可直接应用于现有的对抗性领域自适应方法，以提升性能。

基于视觉转换器的对抗领域自适应