Label noise is inherent in many deep learning tasks when the training set becomes large. A typical approach to tackle noisy labels is using robust loss functions. Categorical cross entropy (CCE) is a successful loss function in many applications. However, CCE is also notorious for fitting samples with corrupted labels easily. In contrast, mean absolute error (MAE) is noise-tolerant theoretically, but it generally works much worse than CCE in practice. In this work, we have three main points. First, to explain why MAE generally performs much worse than CCE, we introduce a new understanding of them fundamentally by exposing their intrinsic sample weighting schemes from the perspective of every sample's gradient magnitude with respect to logit vector. Consequently, we find that MAE's differentiation degree over training examples is too small so that informative ones cannot contribute enough against the non-informative during training. Therefore, MAE generally underfits training data when noise rate is high. Second, based on our finding, we propose an improved MAE (IMAE), which inherits MAE's good noise-robustness. Moreover, the differentiation degree over training data points is controllable so that IMAE addresses the underfitting problem of MAE. Third, the effectiveness of IMAE against CCE and MAE is evaluated empirically with extensive experiments, which focus on image classification under synthetic corrupted labels and video retrieval under real noisy labels.

本文探究了基于经验损失函数中内置的例子加权对抗不正常训练数据的鲁棒性深度学习，重点研究了与对数相关的梯度幅度以及未进行彻底研究的角度。研究发现，均方误差并没有平等地处理例子，梯度幅度的方差很重要，提出了一种称为改进均方误差（IMAE）的解决方案，证明了其在图像分类方面具有出色的效果。

噪声鲁棒性学习的IMAE: 平均绝对误差不能平等地对待样本，梯度的大小方差很重要