We show how to improve the inference efficiency of an LLM by expanding it into a mixture of sparse experts, where each expert is a copy of the original weights, one-shot pruned for a specific cluster of input values. We call this approach $\textit{Sparse Expansion}$. We show that, for models such as Llama 2 70B, as we increase the number of sparse experts, Sparse Expansion outperforms all other one-shot sparsification approaches for the same inference FLOP budget per token, and that this gap grows as sparsity increases, leading to inference speedups. But why? To answer this, we provide strong evidence that the mixture of sparse experts is effectively $\textit{disentangling}$ the input-output relationship of every individual neuron across clusters of inputs. Specifically, sparse experts approximate the dense neuron output distribution with fewer weights by decomposing the distribution into a collection of simpler ones, each with a separate sparse dot product covering it. Interestingly, we show that the Wasserstein distance between a neuron's output distribution and a Gaussian distribution is an indicator of its entanglement level and contribution to the accuracy of the model. Every layer of an LLM has a fraction of highly entangled Wasserstein neurons, and model performance suffers more when these are sparsified as opposed to others.

我们展示了如何通过将LLM扩展为稀疏专家的混合体来提高其推理效率，其中每个专家是原始权重的副本，经过一次性修剪以特定输入值簇的方式修剪。我们称这种方法为'稀疏扩展'。我们展示了对于像LLama 270B这样的模型，随着稀疏专家的数量增加，稀疏扩展在相同推理FLOP预算下胜过所有其他一次性稀疏化方法，并且随着稀疏性的增加，这种差距加大，导致推理加速。

稀疏展开和神经元解缠