Detecting online hate is a difficult task that even state-of-the-art models struggle with. In previous research, hate speech detection models are typically evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model quality due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a first suite of functional tests for hate speech detection models. We specify 29 model functionalities, the selection of which we motivate by reviewing previous research and through a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate data quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer detection models as well as a popular commercial model, revealing critical model weaknesses.

介绍 HateCheck，一个用于针对仇恨言论检测模型的功能测试套件，其中包括 29 个模型功能，为每个功能编写测试用例，并通过结构化注释过程验证其质量。测试表明，近最先进的变换器模型以及两个流行的商业模型存在关键的模型弱点。

HateCheck：仇恨言论检测模型的功能测试