The recent success of general-domain large language models (LLMs) has significantly changed the natural language processing paradigm towards a unified foundation model across domains and applications. In this paper, we focus on assessing the performance of GPT-4, the most capable LLM so far, on the text-based applications for radiology reports, comparing against state-of-the-art (SOTA) radiology-specific models. Exploring various prompting strategies, we evaluated GPT-4 on a diverse range of common radiology tasks and we found GPT-4 either outperforms or is on par with current SOTA radiology models. With zero-shot prompting, GPT-4 already obtains substantial gains ($\approx$ 10% absolute improvement) over radiology models in temporal sentence similarity classification (accuracy) and natural language inference ($F_1$). For tasks that require learning dataset-specific style or schema (e.g. findings summarisation), GPT-4 improves with example-based prompting and matches supervised SOTA. Our extensive error analysis with a board-certified radiologist shows GPT-4 has a sufficient level of radiology knowledge with only occasional errors in complex context that require nuanced domain knowledge. For findings summarisation, GPT-4 outputs are found to be overall comparable with existing manually-written impressions.

本论文评估了目前最先进的大型语言模型GPT-4在放射学报告的文本应用中的表现，探索了各种提示策略，并发现GPT-4在常见放射学任务中表现要优于或与目前最先进的放射学模型相媲美。针对需要学习特定样式或架构的任务，GPT-4通过基于示例的提示得到改进并与监督的最先进模型相匹配。通过与一名获得认证的放射科医生的广泛错误分析表明，GPT-4在放射学知识方面具备足够水平，只偶尔在需要微妙领域知识的复杂上下文中出现错误。针对发现的总结，GPT-4的输出整体上与现有的人工编写印象相当。

探索 GPT-4 在放射学领域的边界