The rapid advancements in large language models and generative artificial intelligence (AI) capabilities are making their broad application in the high-stakes testing context more likely. Use of generative AI in the scoring of constructed responses is particularly appealing because it reduces the effort required for handcrafting features in traditional AI scoring and might even outperform those methods. The purpose of this paper is to highlight the differences in the feature-based and generative AI applications in constructed response scoring systems and propose a set of best practices for the collection of validity evidence to support the use and interpretation of constructed response scores from scoring systems using generative AI. We compare the validity evidence needed in scoring systems using human ratings, feature-based natural language processing AI scoring engines, and generative AI. The evidence needed in the generative AI context is more extensive than in the feature-based NLP scoring context because of the lack of transparency and other concerns unique to generative AI such as consistency. Constructed response score data from standardized tests demonstrate the collection of validity evidence for different types of scoring systems and highlights the numerous complexities and considerations when making a validity argument for these scores. In addition, we discuss how the evaluation of AI scores might include a consideration of how a contributory scoring approach combining multiple AI scores (from different sources) will cover more of the construct in the absence of human ratings.

本研究旨在解决生成人工智能在构建反应评分中的有效性证据不足的问题。文章提出一种新的比较方法，分析了基于特征的人工智能评分与生成人工智能评分系统之间的差异，并建议了收集有效性证据的最佳实践。研究发现，生成人工智能的有效性证据要求比基于特征的自然语言处理评分更为广泛，这显示了在高风险测试中应用生成AI的潜在影响和复杂性。

基于生成人工智能应用的构建反应评分的有效性论证