It has been shown that multilingual BERT (mBERT) yields high quality multilingual representations and enables effective zero-shot transfer. This is suprising given that mBERT does not use any kind of crosslingual signal during training. While recent literature has studied this effect, the exact reason for mBERT's multilinguality is still unknown. We aim to identify architectural properties of BERT as well as linguistic properties of languages that are necessary for BERT to become multilingual. To allow for fast experimentation we propose an efficient setup with small BERT models and synthetic as well as natural data. Overall, we identify six elements that are potentially necessary for BERT to be multilingual. Architectural factors that contribute to multilinguality are underparameterization, shared special tokens (e.g., "[CLS]"), shared position embeddings and replacing masked tokens with random tokens. Factors related to training data that are beneficial for multilinguality are similar word order and comparability of corpora.

该研究通过实现小型BERT模型的混合合成数据和自然数据训练，试图从语言学和结构特征两个方面，探究多语BERT能实现无监督跨语言转移的原因。其结果表明，在lexical、syntactic以及阅读理解方面，mBERT已实现了高质量的多语言表征和跨语言转移功能。

为BERT多语能力识别必要元素