Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may be responsible for such success. In particular, several authors have noted that \emph{over-parametrization} appears to act as a remedy against non-convexity. In this paper, we address this phenomenon by studying key topological properties of the loss, such as the presence or absence of "spurious valleys", defined as connected components of sub-level sets that do not include a global minimum. Focusing on a class of two-layer neural networks defined by smooth (but generally non-linear) activation functions, our main contribution is to prove that as soon as the hidden layer size matches the \emph{intrinsic} dimension of the reproducing space, defined as the linear functional space generated by the activations, no spurious valleys exist, thus allowing the existence of descent directions. Our setup includes smooth activations such as polynomials, both in the empirical and population risk, and generic activations in the empirical risk case.

本文主要研究神经网络中存在的局部极小值问题。针对两层神经网络，定义了其固有维度，并证明了有限的固有维度保证了超参数化的模型不存在局部极小值，而无限的固有维度意味着在某些数据分布下必然存在局部极小值。此外，尽管在一般情况下可能存在局部极小值，但其出现在低风险水平，并高概率地避免在超参数化的模型上。

双层神经网络优化景观中的虚假峰谷