arrow
返回

Detecting Java software similarities by using different clustering techniques

delete2020-06-01
delete11
delete
OA
AI
A
Andrea Capiluppi *
D
Davide Di Ruscio
J
Juri Di Rocco
P
Phuong T. Nguyen
N
Nemitari Ajienka
DOI:10.1016/j.infsof.2020.106279delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Background: Research on empirical software engineering has increasingly been conducted by analysing and measuring vast amounts of software systems. Hundreds, thousands and even millions of systems have been (and are) considered by researchers, and often within the same study, in order to test theories, demonstrate approaches or run prediction models. A much less investigated aspect is whether the collected metrics might be context-specific, or whether systems should be better analysed in clusters. Objective: The objectives of this study are (i) to define a set of clustering techniques that might be used to group similar software systems, and (ii) to evaluate whether a suite of well-known object-oriented metrics is context-specific, and its values differ along the defined clusters. Method: We group software systems based on three different clustering techniques, and we collect the values of the metrics suite in each cluster. We then test whether clusters are statistically different between each other, using the Kolgomorov-Smirnov (KS) hypothesis testing. Results: Our results show that, for two of the used techniques, the KS null hypothesis (e.g., the clusters come from the same population) is rejected for most of the metrics chosen: the clusters that we extracted, based on application domains, show statistically different structural properties. Conclusions: The implications for researchers can be profound: metrics and their interpretation might be more sensitive to context than acknowledged so far, and application domains represent a promising filter to cluster similar systems.
Keyword:
FOSS (Free and open-source software)
Application Domains
Latent Dirichlet Allocation
Machine Learning
Expert Opinions
OO (object-oriented)
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Information and Software Technology 封面图
Information and Software Technology
IF:
4.3
论文数:
3.7K
被引数:
7.7K

机构

University of LAquila 封面图
University of LAquila
学者数:
7.4K
论文数: 6.6K
被引数: 6.7K
B
brunel university
学者数:
5.8K
论文数: 7.1K
被引数: 9
E
Edge Hill University
学者数:
1.2K
论文数: 1.3K
被引数: 958
学者 查看更多机构
引用论文

引用论文

A survey on Test Suite Reduction frameworks and tools
err2016-12-01
err16
errOAAI
errKhan, Saif Ur Rehman; Lee, Sai Peck; Ahmad, Raja Wasim; Akhunzada, Adnan; Chang, Victor
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
Linked Data - The Story So Far
err2009-07-01
err2.2K
errOAAI
errBizer, Christian; Heath, Tom; Berners-Lee, Tim
err分享
err收藏
VARIATION OF EUTRIC CAMBISOLS CHEMICAL PROPERTIES BASED ON ALTITUDINAL AND GEOMORPHOLOGIC ZONING
err2017-01-01
err0
PREAI
errLucian-Constantin Dinca; Gheorghe Sparchez; Gheorghe Marin; Maria Dinca; Raluca-Elena Enescu
err分享
err收藏
学者 查看更多内容