Modern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons, or manifestations of implicit bias. Building a large annotated dataset for such veiled toxicity can be very expensive. In this work, we propose a framework aimed at fortifying existing toxic speech detectors without a large labeled corpus of veiled toxicity. Just a handful of probing examples are used to surface orders of magnitude more disguised offenses. We augment the toxic speech detector's training data with these discovered offensive examples, thereby making it more robust to veiled toxicity while preserving its utility in detecting overt toxicity.

针对现代有毒言语检测器在辨识出具有隐蔽性的攻击语言（如故意避开已知有毒词汇表的对抗性攻击或内在偏见的表现）方面的无能，本文提出了一种框架，该框架旨在强化现有的有毒言语检测器，同时又不需要进行大规模的隐蔽性毒性标注语料库训练。只需用极少量的探测示例就可以揭示出更多隐蔽性攻击言辞，然后将这些发现的攻击性示例用于增强有毒言语检测器的训练数据，从而使其在保留检测显性毒性的效用的同时更加鲁棒。

加强有毒言论检测器以抵御隐晦的有毒言论