1
Return

Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph

delete2025-03-19
delete0
delete
OA
AI
R
Roman Vashurin *
E
Ekaterina Fadeeva
A
Artem Vazhentsev
L
Lyudmila Rvanova
D
Daniil Vasilev
A
Akim Tsvigun
S
Sergey Petrakov
R
Rui Xing
A
Abdelrahman Sadallah
K
Kirill Grishchenkov
A
Alexander Panchenko
T
Timothy Baldwin
P
Preslav Nakov
M
Maxim Panov
A
Artem Shelmanov
DOI:10.1162/tacl_a_00737delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The rapid proliferation of large language models (LLMs) has stimulated researchers to seek effective and efficient approaches to deal with LLM hallucinations and low-quality outputs. Uncertainty quantification (UQ) is a key element of machine learning applications in dealing with such challenges. However, research to date on UQ for LLMs has been fragmented in terms of techniques and evaluation methodologies. In this work, we address this issue by introducing a novel benchmark that implements a collection of state-of-the-art UQ baselines and offers an environment for controllable and consistent evaluation of novel UQ techniques over various text generation tasks. Our benchmark also supports the assessment of confidence normalization methods in terms of their ability to provide interpretable scores. Using our benchmark, we conduct a large-scale empirical investigation of UQ and normalization techniques across eleven tasks, identifying the most effective approaches.

Journal

T
Transactions of the Association for Computational Linguistics
IF:
6.9
Papers:
486
Citations:
5.7K

Organization

N
nebius
Scholars:
1
Papers: 1
Citations: 0
C
ctr artificial intelligence technol
Scholars:
3
Papers: 1
Citations: 0
M
mbzuai
Scholars:
9
Papers: 4
Citations: 2
S
Swiss Fed Inst Technol
Scholars:
2.3K
Papers: 1.1K
Citations: 517
W
weakly supervised nlp grp
Scholars:
1
Papers: 1
Citations: 0
H
HSE Univ
Scholars:
79
Papers: 54
Citations: 20
Cited Papers

Cited Papers

Citing Papers

Citing Papers