arrow
返回

Do Code Summarization Models Process Too Much Information? Function Signature May Be All That Is Needed

delete2024-06-27
delete3
delete
OA
AI
X
Xi Ding
R
Rui Peng
X
Xiangping Chen
黄袁 封面图
黄袁 (Yuan Huang) *
J
Jing Bian
Z
Zibin Zheng
DOI:10.1145/3652156delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the fast development of large software projects, automatic code summarization techniques, which summarize the main functionalities of a piece of code using natural languages as comments, play essential roles in helping developers understand and maintain large software projects. Many research efforts have been devoted to building automatic code summarization approaches. Typical code summarization approaches are based on deep learning models. They transform the task into a sequence-to-sequence task, which inputs source code and outputs summarizations in natural languages. All code summarization models impose different input size limits, such as 50 to 10,000, for the input source code. However, how the input size limit affects the performance of code summarization models still remains under-explored. In this article, we first conduct an empirical study to investigate the impacts of different input size limits on the quality of generated code comments. To our surprise, experiments on multiple models and datasets reveal that setting a low input size limit, such as 20, does not necessarily reduce the quality of generated comments. Based on this finding, we further propose to use function signatures instead of full source code to summarize the main functionalities first and then input the function signatures into code summarization models. Experiments and statistical results show that inputs with signatures are, on average, more than 2 percentage points better than inputs without signatures and thus demonstrate the effectiveness of involving function signatures in code summarization. We also invite programmers to do a questionnaire to evaluate the quality of code summaries generated by two inputs with different truncation levels. The results show that function signatures generate, on average, 9.2% more high-quality comments than full code.
Keyword:
Code summarization
function signature
empirical study

期刊

A
ACM Transactions on Software Engineering and Methodology
IF:
6.2
论文数:
1.2K
被引数:
3.4K

机构

S
Sun Yat Sen University
学者数:
9.9W
论文数: 7.2W
被引数: 95
引用论文

引用论文

err
IF0
err
err0
errOAAI
err
err分享
err收藏
SCALING WITH KNOWN UNCERTAINTY: A SYNTHESIS
err2006-01-01
err0
PREAI
errJIANGUO WU; HARBIN LI; K. BRUCE JONES; ORIE L. LOUCKS
err分享
err收藏
Impurity effects on ionic-liquid-based supercapacitors
err2016-12-27
err0
errOAAI
errKun Liu; Cheng Lian; Douglas Henderson; Jianzhong Wu
err分享
err收藏
A Comparative Study on Method Comment and Inline Comment方法注释与内联注释的比较研究
err2023-07-22
err8
errOAAI
errHuang, Yuan; Guo, Hanyang; Ding, Xi; Shu, Junhuai; Chen, Xiangping; Luo, Xiapu; Zheng, Zibin; Zhou, Xiaocong
err分享
err收藏
Establishment and characterization of a new triple-negative canine mammary cancer cell line
err2018-10-01
err0
PREAI
errHong Zhang; Shimin Pei; Bin Zhou; Huanan Wang; Hongchao Du; Di Zhang; Degui Lin
err分享
err收藏
Measuring Program Comprehension: A Large-Scale Field Study with Professionals
err2018-10-01
err195
errOAAI
errXia, Xin; Bao, Lingfeng; Lo, David; Xing, Zhenchang; Hassan, Ahmed E.; Li, Shanping
err分享
err收藏
学者 查看更多内容