Transformers are successfully applied to computer vision due to their powerful modeling capacity with self-attention. However, the excellent performance of transformers heavily depends on enormous training images. Thus, a data-efficient transformer solution is urgently needed. In this work, we propose an early knowledge distillation framework, which is termed as DearKD, to improve the data efficiency required by transformers. Our DearKD is a two-stage framework that first distills the inductive biases from the early intermediate layers of a CNN and then gives the transformer full play by training without distillation. Further, our DearKD can be readily applied to the extreme data-free case where no real images are available. In this case, we propose a boundary-preserving intra-divergence loss based on DeepInversion to further close the performance gap against the full-data counterpart. Extensive experiments on ImageNet, partial ImageNet, data-free setting and other downstream tasks prove the superiority of DearKD over its baselines and state-of-the-art methods.

本文提出了一种早期知识蒸馏框架(DearKD)，通过从卷积神经网络的早期中间层中提取归纳偏差然后通过无蒸馏进行训练，以提高变压器所需的数据效率。我们还针对极端的零数据情况提出了一种基于DeepInversion的边界保留内部分歧损失，从而进一步缩小与完整数据对照组之间的性能差距。针对ImageNet、partial ImageNet、无数据设置和其他下游任务的大量实验证明DearKD优于其基准和最先进的方法。

DearKD：用于Vision Transformers的数据高效早期知识蒸馏