Deep neural networks are known to exhibit a `double descent' behavior as the number of parameters increases. Recently, it has also been shown that an `epochwise double descent' effect exists in which the generalization error initially drops, then rises, and finally drops again with increasing training time. This presents a practical problem in that the amount of time required for training is long, and early stopping based on validation performance may result in suboptimal generalization. In this work we develop an analytically tractable model of epochwise double descent that allows us to characterise theoretically when this effect is likely to occur. This model is based on the hypothesis that the training data contains features that are slow to learn but informative. We then show experimentally that deep neural networks behave similarly to our theoretical model. Our findings indicate that epochwise double descent requires a critical amount of noise to occur, but above a second critical noise level early stopping remains effective. Using insights from theory, we give two methods by which epochwise double descent can be removed: one that removes slow to learn features from the input and reduces generalization performance, and another that instead modifies the training dynamics and matches or exceeds the generalization performance of standard training. Taken together, our results suggest a new picture of how epochwise double descent emerges from the interplay between the dynamics of training and noise in the training data.

本文研究表明，随着参数数量的增加，深度神经网络会呈现出“双下降”的特性，同时，随着训练时间的增长，也存在着“按时间下降的双重下降”效应，这在实践中导致训练时间过长，基于验证表现的早停可能导致非最优泛化。作者提出了一种可以从理论上解释“按时间下降的双重下降”的模型，并提供了两种方法来消除这种效应。通过理论分析和实验验证表明，消除缓慢学习特征或修改训练方式可以消除“按时间下降的双重下降”，并且改善模型泛化性能。

epochwise双重下降发生的时间和方式