arrow
Return

Detecting code smells using industry-relevant data

delete2023-03-01
delete16
PRE
AI
L
Lech Madeyski *
T
Tomasz Lewowski
DOI:10.1016/j.infsof.2022.107112delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Context Code smells are patterns in source code associated with an increased defect rate and a higher maintenance effort than usual, but without a clear definition. Code smells are often detected using rules hard -coded in detection tools. Such rules are often set arbitrarily or derived from data sets tagged by reviewers without the necessary industrial know-how. Conclusions from studying such data sets may be unreliable or even harmful, since algorithms may achieve higher values of performance metrics on them than on models tagged by experts, despite not being industrially useful. Objective Our goal is to investigate the performance of various machine learning algorithms for auto-mated code smell detection trained on code smell data set(MLCQ) derived from actively developed and industry-relevant projects and reviews performed by experienced software developers.Method We assign the severity of the smell to the code sample according to a consensus between the severities assigned by the reviewers, use the Matthews Correlation Coefficient (MCC) as our main performance metric to account for the entire confusion matrix, and compare the median value to account for non-normal distributions of performance. We compare 6720 models built using eight machine learning techniques. The entire process is automated and reproducible.Results Performance of compared techniques depends heavily on analyzed smell. The median value of our performance metric for the best algorithm was 0.81 for Long Method, 0.31 for Feature Envy, 0.51 for Blob, and 0.57 for Data Class.Conclusions Random Forest and Flexible Discriminant Analysis performed the best overall, but in most cases the performance difference between them and the median algorithm was no more than 10% of the latter. The performance results were stable over multiple iterations. Although the F-score omits one quadrant of the confusion matrix (and thus may differ from MCC), in code smell detection, the actual differences are minimal.
Keywords:
Reproducible research
Software engineering
Machine learning
Code smells

Journal

Information and Software Technology cover
Information and Software Technology
IF:
4.3
Papers:
3.8K
Citations:
7.7K

Organization

W
wroclaw university of science & technology
Scholars:
7.4K
Papers: 7.1K
Citations: 2
Cited Papers

Cited Papers

Automatic detection of Long Method and God Class code smells through neural source code embeddings
err2022-10-01
err25
errOAAI
errKovacevic, Aleksandar; Slivka, Jelena; Vidakovic, Dragan; Grujic, Katarina-Glorija; Luburic, Nikola; Prokic, Simona; Sladic, Goran
errShare
errSave
Popping Balloons: A Case Study of Dynamical Fragmentation
err2015-10-30
err0
PREAI
errSébastien Moulinet; Mokhtar Adda-Bedia
errShare
errSave
Comparing and experimenting machine learning techniques for code smell detection
err2015-06-06
err277
PREAI
errFontana, Francesca Arcelli; Mantyla, Mika V.; Zanoni, Marco; Marino, Alessandro
errShare
errSave
Recombination between an expressed immunoglobulin heavy-chain gene and a germline variable gene segment in a Ly 1+ B-cell lymphoma
err1986-08-01
err0
PREAI
errRobert Kleinfield; Richard R. Hardy; David Tarlinton; Jeffery Dangl; Leonard A. Herzenberg; Martin Weigert
errShare
errSave
Single-lung and double-lung transplantation: technique and tips
err2018-04-01
err0
errOAAI
errLucile Gust; Xavier-Benoit D’Journo; Geoffrey Brioude; Delphine Trousse; Stephanie Dizier; Christophe Doddoli; Marc Leone; Pascal-Alexandre Thomas
errShare
errSave
researcher View more