In times of emergency, crisis response agencies need to quickly and accurately assess the situation on the ground in order to deploy relevant services and resources. However, authorities often have to make decisions based on limited information, as data on affected regions can be scarce until local response services can provide first-hand reports. Fortunately, the widespread availability of smartphones with high-quality cameras has made citizen journalism through social media a valuable source of information for crisis responders. However, analyzing the large volume of images posted by citizens requires more time and effort than is typically available. To address this issue, this paper proposes the use of state-of-the-art deep neural models for automatic image classification/tagging, specifically by adapting transformer-based architectures for crisis image classification (CrisisViT). We leverage the new Incidents1M crisis image dataset to develop a range of new transformer-based image classification models. Through experimentation over the standard Crisis image benchmark dataset, we demonstrate that the CrisisViT models significantly outperform previous approaches in emergency type, image relevance, humanitarian category, and damage severity classification. Additionally, we show that the new Incidents1M dataset can further augment the CrisisViT models resulting in an additional 1.25% absolute accuracy gain.

利用最新的深度神经模型，通过将基于Transformer的架构应用于危机图像分类（CrisisViT），以解决利用社交媒体的公民新闻来帮助危机响应的问题，并通过实验证明，CrisisViT模型在紧急类型、图像相关性、人道主义类别和损害严重性分类方面明显优于以前的方法。此外，新的Incidents1M数据集进一步增强了CrisisViT模型，使其准确率提高了1.25%。

CrisisViT：一种适用于危机图像分类的稳健视觉Transformer