arrow
返回

How well do pre-trained contextual language representations recommend labels for GitHub issues?

delete2021-11-01
delete20
PRE
AI
J
Jun Wang
章晓芳 封面图
章晓芳 (Xiaofang Zhang) *
L
Lin Chen
DOI:10.1016/j.knosys.2021.107476delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Motivation: Open-source organizations use issues to collect user feedback, software bugs, and feature requests in GitHub. Many issues do not have labels, which makes labeling time-consuming work for the maintainers. Recently, some researchers used deep learning to improve the performance of automated tagging for software objects. However, these researches use static pre-trained word vectors that cannot represent the semantics of the same word in different contexts. Pre-trained contextual language representations have been shown to achieve outstanding performance on lots of NLP tasks. Description: In this paper, we study whether the pre-trained contextual language models are really better than other previous language models in the label recommendation for the GitHub labels scenario. We try to give some suggestions in fine-tuning pre-trained contextual language representation models. First, we compared four deep learning models, in which three of them use traditional pretrained word embedding. Furthermore, we compare the performances when using different corpora for pre-training. Results: The experimental results show that: (1) When using large training data, the performance of BERT model is better than other deep learning language models such as Bi-LSTM, CNN and RCNN. While with a small size training data, CNN performs better than BERT. (2) Further pre-training on domain-specific data can indeed improve the performance of models. Conclusions: When recommending labels for issues in GitHub, using pre-trained contextual language representations is better if the training dataset is large enough. Moreover, we discuss the experimental results and provide some implications to improve label recommendation performance for GitHub issues. (C) 2021 Elsevier B.V. All rights reserved.
Keyword:
Deep learning
Issue labeling
Data analysis
Language model

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

N
nanjing university
学者数:
7.8W
论文数: 5.6W
被引数: 87
S
soochow university - china
学者数:
5.2W
论文数: 3.6W
被引数: 82
引用论文

引用论文

EnTagRec++: An enhanced tag recommendation system for software information sites
err2017-07-21
err60
errOAAI
errWang, Shaowei; Lo, David; Vasilescu, Bogdan; Serebrenik, Alexander
err分享
err收藏
YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
err分享
err收藏
Single Versus Double Anatomic Site Intraosseous Blood Transfusion in a Swine Model of Hemorrhagic Shock
err2021-11-01
err0
errOAAI
errEric Sulava; William Bianchi; Christian S. McEvoy; Paul J. Roszko; Gregory J. Zarow; Micah J. Gaspary; Ramesh Natarajan; Jonathan D. Auten
err分享
err收藏
Carotid Body Inflammation: Role in Hypoxia and in the Anti-inflammatory Reflex
err2022-05-01
err0
PREAI
errRodrigo Iturriaga; Rodrigo Del Rio; Julio Alcayaga
err分享
err收藏
学者 查看更多内容