This study aims to address the pervasive challenge of quantifying uncertainty in large language models (LLMs) without logit-access. Conformal Prediction (CP), known for its model-agnostic and distribution-free features, is a desired approach for various LLMs and data distributions. However, existing CP methods for LLMs typically assume access to the logits, which are unavailable for some API-only LLMs. In addition, logits are known to be miscalibrated, potentially leading to degraded CP performance. To tackle these challenges, we introduce a novel CP method that (1) is tailored for API-only LLMs without logit-access; (2) minimizes the size of prediction sets; and (3) ensures a statistical guarantee of the user-defined coverage. The core idea of this approach is to formulate nonconformity measures using both coarse-grained (i.e., sample frequency) and fine-grained uncertainty notions (e.g., semantic similarity). Experimental results on both close-ended and open-ended Question Answering tasks show our approach can mostly outperform the logit-based CP baselines.

本研究旨在解决大型语言模型中无法访问 logits 的不确定性量化的普遍挑战。我们提出了一种面向 API-only 语言模型的新型 CP 方法，通过同时利用粗粒度（如样本频率）和细粒度（如语义相似性）的不确定性概念来构建不确定度量，实现了更好的预测性能。实验证明，我们的方法在封闭式和开放式问答任务中大多能够胜过基于 logits 的 CP 对照组。

API已足够：大型语言模型的无需访问逻辑函数的符合预测