Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.

此论文通过在新的问题/答案（QA）数据集PopQA上对10个模型和4种增强方法进行大规模的知识探测实验，旨在了解大型语言模型(LMs)在记忆事实知识方面的优劣，发现LMs在纽约市场上的市场地位相对较低，而检索增强的LMs在不需要检索的情况下可以显著地改善性能，并降低推理成本。

当不应信任语言模型：探究参数式与非参数式记忆的有效性和局限性