arrow
Return

Can Large Language Models Serve as Evaluators for Code Summarization?

delete2025-08-19
delete0
delete
OA
AI
Y
Yang Wu
Y
Yao Wan
Z
Zhaoyang Chu
W
Wenting Zhao
Y
Ye Liu
H
Hongyu Zhang
X
Xuanhua Shi
金海 (Hai Jin)
P
Philip S. Yu
DOI:10.1109/TSE.2025.3595283delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Code summarization facilitates program comprehension and software maintenance by converting code snippets into natural-language descriptions. Over the years, numerous methods have been developed for this task, but a key challenge remains: effectively evaluating the quality of generated summaries. While human evaluation is effective for assessing code summary quality, it is labor-intensive and difficult to scale. Commonly used automatic metrics, such as BLEU, ROUGE-L, METEOR, and BERTScore, often fail to align closely with human judgments. In this paper, we explore the potential of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Large Language Models (LLMs)</i> for evaluating code summarization. We propose <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CodeRPE</small> (Role-Player for Code Summarization Evaluation), a novel method that leverages role-player prompting to assess the quality of generated summaries. Specifically, we prompt LLM-based evaluators to take on diverse roles, such as code reviewer, code author, code editor, and system analyst. Each role evaluates the quality of code summaries across key dimensions, including coherence, consistency, fluency, and relevance. We further explore the robustness of LLMs as evaluators by employing various prompting strategies, including chain-of-thought reasoning, in-context learning, and tailored rating form designs. The results demonstrate that LLMs serve as effective evaluators for code summarization. Notably, our LLM-based evaluator, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CodeRPE</small>, achieves an 80.18% Spearman correlation with human evaluations, outperforming the existing BERTScore metric by 10.39%.
Keywords:
Code summarization
large language models
role player
model evaluation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Software Engineering cover
IEEE Transactions on Software Engineering
IF:
5.6
Papers:
2.8K
Citations:
1.1W

Organization

S
salesforce research, palo alto, ca, usa
Scholars:
2
Papers: 1
Citations: 0
C
Chongqing University
Scholars:
5.1W
Papers: 4.1W
Citations: 6.0W
U
University of Illinois Chicago
Scholars:
1.7W
Papers: 1.4W
Citations: 3.0W
H
huazhong university of science and technology
Scholars:
2.5W
Papers: 7.5K
Citations: 5
researcher View more organizations