arrow
Return

Efficient Benchmarking via Bias-Bounded Subset Selection

delete2025-08-12
delete0
PRE
AI
Y
Yan Zhuang
J
Junhao Yu
刘琦 (Qi Liu)
Y
Yuxuan Sun
J
Jiatong Li
黄振亚 (Zhenya Huang)
陈恩红 (Enhong Chen)
DOI:10.1109/TPAMI.2025.3598031delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Evaluating AI systems, particularly large models, is an essential yet computationally expensive task. The use of extensive benchmarks often leads to substantial computational/human costs that may even exceed those of pretraining. The efficiency of AI model evaluation focuses on estimating the model’s score on the full benchmark based on its responses to a smaller subset. Various empirical selection methods have been proposed to identify valuable subsets within these benchmarks. In this paper, we formally define and approximate the subset selection problem inherent in efficient evaluation. We prove that this problem actually optimizes a submodular function and that a unified subset can be identified using a simple greedy algorithm. Importantly, this approach is the first to provide theoretical guarantees of bias control and generalizability in score estimation. Using language models as a case study, experimental results across 11 different benchmarks validate its superiority in estimating model scores and maintaining ranking consistency. It can achieve accurate score estimation using no more than 30% of the full benchmark, thus facilitating efficient and sparse benchmark design.
Keywords:
AI evaluation
benchmark
metric
performance prediction
subset selection

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

U
university of science and technology of china
Scholars:
1.0W
Papers: 3.9K
Citations: 3