In task-oriented conversational AI evaluation, unsupervised methods poorly
correlate with human judgments, and supervised approaches lack generalization.
Recent advances in large language models (LLMs) show robust zeroshot and
few-shot capabilities across NLP tasks. This paper explores using LLMs for
automated dialogue quality evaluation, experimenting with various
configurations on public and proprietary datasets. Manipulating factors such as
model size, in-context examples, and selection techniques, we examine
"chain-of-thought" (CoT) reasoning and label extraction procedures. Our results
show that (1) larger models yield more accurate dialogue labels; (2)
algorithmic selection of in-context examples outperforms random selection; (3)
CoT reasoning where an LLM is asked to provide justifications before outputting
final labels improves performance; and (4) fine-tuned LLMs outperform
out-of-the-box ones. Our results indicate that LLMs that are suitably
fine-tuned and have sufficient reasoning capabilities can be leveraged for
automated dialogue evaluation.

该论文探讨了使用大型语言模型（LLMs）进行自动对话质量评估的方法，并在公共和专有数据集上尝试了各种配置。结果表明，更大的模型产生了更准确的对话标签；算法选择背景上下文示例优于随机选择；在输出最终标签之前，使用 “思维链”（CoT）推理和标签提取过程进行合理化，可以提高性能；精细调整的 LLMs 优于开箱即用的模型。研究结果表明，合适地调整和具有足够推理能力的 LLMs 可以用于自动对话评估。