arrow
Return

Can Public Code Smells Datasets Be Trusted?

delete2025-12-01
delete0
PRE
AI
R
Ruchin Gupta *
J
Jitendra Kumar Seth
A
Abhishek Goyal
DOI:10.24138/jcomss-2025-0131delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
smells signal potential issues in a codebase and indicate technical debt. Early detection is crucial for maintaining code quality. Researchers often rely on public datasets to automate and enhance smell detection, but their trustworthiness is frequently assumed rather than verified. While these datasets are valuable for developing detection tools, key questions arise: Can they be fully trusted? Are the labels accurate? Do they reflect realworld software development? Recent studies reveal inconsistencies, biases, and misclassifications, raising concerns about their reliability. This paper explores the integrity of widely used 2 sets of public code smells datasets namely Group A dataset and Group B dataset by examining their internal consistency, alignment with established facts. Through this investigation, we aim to determine whether these datasets can be confidently utilized in research and practical applications, or if their inherent issues undermine the validity of the results they produce. Group A datasets are smaller, balanced, and factually aligned but lack industry relevance, while Group B deviates from known facts. The study acknowledges academic-industry differences, viewing divergence as a reflection of real-world variability rather than a flaw, and emphasizes the need for rigorous validation of public datasets to ensure reliable research outcomes.
Keywords:
Code smell
code smell datasets
validation.

Journal

J
Journal of Communications Software and Systems
IF:
0.7
Papers:
38
Citations:
171

Organization

K
kiet group of institutions
Scholars:
21
Papers: 15
Citations: 0