Recent studies demonstrate that deep networks, even robustified by the state-of-the-art adversarial training (AT), still suffer from large robust generalization gaps, in addition to the much more expensive training costs than standard training. In this paper, we investigate this intriguing problem from a new perspective, i.e., injecting appropriate forms of sparsity during adversarial training. We introduce two alternatives for sparse adversarial training: (i) static sparsity, by leveraging recent results from the lottery ticket hypothesis to identify critical sparse subnetworks arising from the early training; (ii) dynamic sparsity, by allowing the sparse subnetwork to adaptively adjust its connectivity pattern (while sticking to the same sparsity ratio) throughout training. We find both static and dynamic sparse methods to yield win-win: substantially shrinking the robust generalization gap and alleviating the robust overfitting, meanwhile significantly saving training and inference FLOPs. Extensive experiments validate our proposals with multiple network architectures on diverse datasets, including CIFAR-10/100 and Tiny-ImageNet. For example, our methods reduce robust generalization gap and overfitting by 34.44% and 4.02%, with comparable robust/standard accuracy boosts and 87.83%/87.82% training/inference FLOPs savings on CIFAR-100 with ResNet-18. Besides, our approaches can be organically combined with existing regularizers, establishing new state-of-the-art results in AT. Codes are available in https://github.com/VITA-Group/Sparsity-Win-Robust-Generalization.

本文提出两种新颖的在对抗训练期间注入适当稀疏形式的方法，即：通过利用最近的彩票假设的结果识别早期训练中出现的关键稀疏子网络来实现静态稀疏，以及通过在训练期间使稀疏子网络自适应调整其连接模式（同时保持相同的稀疏比率）来实现动态稀疏，并发现这两种新方法都可以显著缩减稳健泛化差距和减轻过度拟合，同时大大减少训练和推理的FLOPs，实验证明此方法在各种数据集上有着显著作用，包括CIFAR-10/100和Tiny-ImageNet。

稀疏性双赢：更高效的训练带来更好的鲁棒泛化