Neural networks require careful weight initialization to prevent signals from exploding or vanishing. Existing initialization schemes solve this problem in specific cases by assuming that the network has a certain activation function or topology. It is difficult to derive such weight initialization strategies, and modern architectures therefore often use these same initialization schemes even though their assumptions do not hold. This paper introduces AutoInit, a weight initialization algorithm that automatically adapts to different neural network architectures. By analytically tracking the mean and variance of signals as they propagate through the network, AutoInit is able to appropriately scale the weights at each layer to avoid exploding or vanishing signals. Experiments demonstrate that AutoInit improves performance of various convolutional and residual networks across a range of activation function, dropout, weight decay, learning rate, and normalizer settings. Further, in neural architecture search and activation function meta-learning, AutoInit automatically calculates specialized weight initialization strategies for thousands of unique architectures and hundreds of unique activation functions, and improves performance in vision, language, tabular, multi-task, and transfer learning scenarios. AutoInit thus serves as an automatic configuration tool that makes design of new neural network architectures more robust. The AutoInit package provides a wrapper around existing TensorFlow models and is available at https://github.com/cognizant-ai-labs/autoinit.

本文介绍了一种自适应不同神经网络结构的权重初始化算法AutoInit，该算法通过跟踪信号传播时的均值和方差，适当地调整每层的权重，从而避免信号爆炸或消失。实验证明，AutoInit在各种激活函数、正则化、学习率和归一化设置下，都能提高卷积、残差和Transformer网络的性能，并比依赖数据的初始化方法更可靠。该算法的灵活性使其能够为各种规模的任务初始化模型，是神经架构搜索和激活函数发现等领域一种自动化配置工具，使新神经网络结构的设计更加鲁棒。AutoInit package提供了一个TensorFlow的封装，可在此 URL中获得。

AutoInit: 神经网络分析信号保持的权重初始化