Recent work in network quantization has substantially reduced the time and space complexity of neural network inference, enabling their deployment on embedded and mobile devices with limited computational and memory resources. However, existing quantization methods often represent all weights and activations with the same precision (bit-width). In this paper, we explore a new dimension of the design space: quantizing different layers with different bit-widths. We formulate this problem as a neural architecture search problem and propose a novel differentiable neural architecture search (DNAS) framework to efficiently explore its exponential search space with gradient-based optimization. Experiments show we surpass the state-of-the-art compression of ResNet on CIFAR-10 and ImageNet. Our quantized models with 21.1x smaller model size or 103.9x lower computational cost can still outperform baseline quantized or even full precision models.

该研究探索了一种新的神经网络压缩方法，通过不同比特宽度的量化不同层并使用可微分神经架构搜索框架进行优化，成功地实现了比现有方法更高的压缩率，模型尺寸缩小21.1倍或计算量降低103.9倍

可微神经架构搜索进行卷积网络的混合精度量化