The spread of various forms of offensive speech online is an important concern in social media. While platforms have been investing heavily in ways of coping with this problem, the question of privacy remains largely unaddressed. Models trained to detect offensive language on social media are trained and/or fine-tuned using large amounts of data often stored in centralized servers. Since most social media data originates from end users, we propose a privacy preserving decentralized architecture for identifying offensive language online by introducing Federated Learning (FL) in the context of offensive language identification. FL is a decentralized architecture that allows multiple models to be trained locally without the need for data sharing hence preserving users' privacy. We propose a model fusion approach to perform FL. We trained multiple deep learning models on four publicly available English benchmark datasets (AHSD, HASOC, HateXplain, OLID) and evaluated their performance in detail. We also present initial cross-lingual experiments in English and Spanish. We show that the proposed model fusion approach outperforms baselines in all the datasets while preserving privacy.

通过引入联邦学习（FL）在辱骂语言识别中的上下文中，我们提出了一种保护用户隐私的去中心化架构，用于辨别网上的辱骂语言。在四个公开可用的英语基准数据集（AHSD、HASOC、HateXplain、OLID）上，我们对多个深度学习模型进行了训练，并进行了详细的性能评估。同时，我们也展示了初步的英语和西班牙语跨语言实验。我们证明了所提出的模型融合方法在所有数据集上优于基准方法，并且能够保护隐私。

一种隐私保护冒犯性语言识别的联邦学习方法