We explore techniques to significantly improve the compute efficiency and performance of Deep Convolution Networks without impacting their accuracy. To improve the compute efficiency, we focus on achieving high accuracy with extremely low-precision (2-bit) weight networks, and to accelerate the execution time, we aggressively skip operations on zero-values. We achieve the highest reported accuracy of 76.6% Top-1/93% Top-5 on the Imagenet object classification challenge with low-precision network\footnote{github release of the source code coming soon} while reducing the compute requirement by ~3x compared to a full-precision network that achieves similar accuracy. Furthermore, to fully exploit the benefits of our low-precision networks, we build a deep learning accelerator core, dLAC, that can achieve up to 1 TFLOP/mm^2 equivalent for single-precision floating-point operations (~2 TFLOP/mm^2 for half-precision).

本研究旨在通过采用极低精度（2位）权重网络，并在零值上进行操作跳过以提高计算效率和性能，以在低精度网络下获得更高精度。实验结果表明，与全精度网络相比，在并非影响相似准确度的情况下，计算需求降低了约3倍，且在Imagenet物体分类挑战上取得了最高报道准确度。为了充分利用低精度网络优势，研究小组开发了一种深度学习加速器核心dLAC，可实现每平方毫米单精度浮点运算的TFLOP当量，半精度时可达到每平方毫米的2个TFLOP。

低精度和稀疏性加速深度卷积网络