Mathematics is a highly specialized domain with its own unique set of challenges. Despite this, there has been relatively little research on natural language processing for mathematical texts, and there are few mathematical language resources aimed at NLP. In this paper, we aim to provide annotated corpora that can be used to study the language of mathematics in different contexts, ranging from fundamental concepts found in textbooks to advanced research mathematics. We preprocess the corpora with a neural parsing model and some manual intervention to provide part-of-speech tags, lemmas, and dependency trees. In total, we provide 182397 sentences across three corpora. We then aim to test and evaluate several noteworthy natural language processing models using these corpora, to show how well they can adapt to the domain of mathematics and provide useful tools for exploring mathematical language. We evaluate several neural and symbolic models against benchmarks that we extract from the corpus metadata to show that terminology extraction and definition extraction do not easily generalize to mathematics, and that additional work is needed to achieve good performance on these metrics. Finally, we provide a learning assistant that grants access to the content of these corpora in a context-sensitive manner, utilizing text search and entity linking. Though our corpora and benchmarks provide useful metrics for evaluating mathematical language processing, further work is necessary to adapt models to mathematics in order to provide more effective learning assistants and apply NLP methods to different mathematical domains.

本文旨在提供可用于研究数学语言的不同背景下的带有注释的文献资料，并使用神经解析模型和人工干预预处理这些资料，以提供词性标签、词形还原和依赖树。我们评估了几种自然语言处理模型，在从文献资料中提取的基准数据上测试它们的性能，并展示它们在数学领域中的适应性和对于探索数学语言的有用性。虽然我们提供了学习助手以在特定环境中访问这些资料内容，进一步的工作仍然需要进行以使模型更好地适应数学，并提供更有效的学习助手以及将自然语言处理方法应用于不同的数学领域。

数学实体：语料库与基准