返回
Semi-supervised multitask learning using convolutional autoencoder for faulty code detection with limited data
DOI:10.1007/s10489-022-03663-5.png)
摘要
En 中文
Detecting faults in source code to fix is an important task in the software quality assurance. Building automated detectors using machine learning has been faced two big challenges of data imbalance and shortages. To address the issues, this paper proposes a deep neural network and training procedures to allow learning with limited annotated data. The network is composed of an unsupervised auto-encoder and a supervised classifier. The two components share some first layers that plays as a program feature extractor. Interestingly, we can leverage a large amount of unlabeled data from various sources to train the auto-encoder independently then transfer to the target domain. Additionally, sharing layers, and jointly training the reconstruction and the classification tasks stimulate the generation of the sophisticated features. We conducted the experiments on four real datasets with different amount of labeled data and with adding more unlabeled data. The results have confirmed that the multi-task outperforms the single-task and leveraging the unlabeled data is beneficial. Specifically, when reducing the labeled data from 100% to 75%, 50%, 25%, the performance of several deep networks drops sharply, while it reduces gradually for our model.
Keyword:
Faulty code detection
Semi-supervised learning
Multitask learning
Self-supervised learning
Convolutional autoencoder
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Automatically identifying code features for software defect prediction: Using AST N-grams用于软件缺陷预测的自动识别代码特征: 使用AST n-gram
Machine Learning Investigation For Tri-Magnetized Sutterby Nanofluidic Model with Joule Heating In Agrivoltaics Technology
Nano
IF0
Deep neural network based hybrid approach for software defect prediction using software metrics基于深度神经网络的软件度量软件缺陷预测混合方法

