Data augmentation has been pivotal in successfully training deep learning models on classification tasks over the past decade. An important subclass of data augmentation techniques - which includes both label smoothing and Mixup - involves modifying not only the input data but also the input label during model training. In this work, we analyze the role played by the label augmentation aspect of such methods. We prove that linear models on linearly separable data trained with label augmentation learn only the minimum variance features in the data, while standard training (which includes weight decay) can learn higher variance features. An important consequence of our results is negative: label smoothing and Mixup can be less robust to adversarial perturbations of the training data when compared to standard training. We verify that our theory reflects practice via a range of experiments on synthetic data and image classification benchmarks.

我们分析了标签增强方法在模型训练中的作用，证明了采用标签增强的线性模型仅仅学习数据中的最小方差特征，而标准训练则能够学习到更高方差的特征。我们的结果表明，与标准训练相比，标签平滑和Mixup在对抗性扰动下对训练数据的鲁棒性较差。通过对合成数据和图像分类基准的一系列实验，我们验证了我们的理论与实践的一致性。

学习最小方差特征通过标签增强