arrow
返回

Automatic detection of Feature Envy and Data Class code smells using machine learning

delete2024-06-01
delete1
delete
OA
AI
M
Milica Škipina
Ј
Јелена Сливка
N
Nikola Luburić
A
Aleksandar Kovačević *
DOI:10.1016/j.eswa.2023.122855delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Code smells in software indicate poor design and implementation choices. Detecting and removing them is critical for sustainable software development. Machine learning (ML) can automate code smell detection. Most ML solutions train models from scratch on code smell datasets, using handcrafted source code metrics as features. Pretrained language models, like BERT, fueled a paradigm shift in natural language processing: from handcrafted features to automatically inferred features and from training models from scratch to using pretrained models. Code embeddings offer the potential to bring a similar paradigm shift to code analysis. Nevertheless, the potential of using pretrained neural code embeddings for code smell detection has yet to be fully explored. To this end, we evaluated ML models trained using different code representations: code metrics and state-of-the-art neural code embeddings (CodeT5 and CuBERT). We experimented with CodeT5 variants (base and small) and explored multiple ways of embedding code snippets (by combining line-level embeddings or passing the entire code snippet as input). We tested our approaches on the tasks of detecting Data Class and Feature Envy on the MLCQ dataset. Considering the results of this study and our previous research, performance-wise, there is no clear winner between using code metrics or code embeddings for different code smell types and programming languages. However, given that, in contrast to code metrics, code embeddings can automatically adapt to new programming constructs and are expected to scale better with dataset size, these models are likely to become the future state-of-the-art feature generation technique for code smell detection.
Keyword:
Code smell detection
Neural source code embeddings
Code metrics
Machine learning
Software engineering
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

U
University of Novi Sad
学者数:
9.3K
论文数: 6.2K
被引数: 5.0K
引用论文

引用论文

Automatic detection of Long Method and God Class code smells through neural source code embeddings
err2022-10-01
err25
errOAAI
errKovacevic, Aleksandar; Slivka, Jelena; Vidakovic, Dragan; Grujic, Katarina-Glorija; Luburic, Nikola; Prokic, Simona; Sladic, Goran
err分享
err收藏
err分享
err收藏
err分享
err收藏
Competition of Multi-Platform Ecosystems in the IoT
err2020-01-01
err0
PREAI
errFrank MacCrory; Evangelos Katsamakas
err分享
err收藏
Bagging predictorsBagging预测器
err1996-08-01
err1.0W
PREAI
errBreiman, L
err分享
err收藏
A large empirical assessment of the role of data balancing in machine-learning-based code smell detection
err2020-11-01
err62
errOAAI
errPecorelli, Fabiano; Di Nucci, Dario; De Roover, Coen; De Lucia, Andrea
err分享
err收藏
学者 查看更多内容