arrow
返回

Curating GitHub for engineered software projects

delete2017-04-18
delete249
PRE
AI
M
Munaiah, Nuthan *
S
Steven Kroh
C
Craig Cabrey
M
Meiyappan Nagappan
DOI:10.1007/s10664-017-9512-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Software forges like GitHub host millions of repositories. Software engineering researchers have been able to take advantage of such a large corpora of potential study subjects with the help of tools like GHTorrent and Boa. However, the simplicity in querying comes with a caveat: there are limited means of separating the signal (e.g. repositories containing engineered software projects) from the noise (e.g. repositories containing home work assignments). The proportion of noise in a random sample of repositories could skew the study and may lead to researchers reaching unrealistic, potentially inaccurate, conclusions. We argue that it is imperative to have the ability to sieve out the noise in such large repository forges. We propose a framework, and present a reference implementation of the framework as a tool called reaper, to enable researchers to select GitHub repositories that contain evidence of an engineered software project. We identify software engineering practices (called dimensions) and propose means for validating their existence in a GitHub repository. We used reaper to measure the dimensions of 1,857,423 GitHub repositories. We then used manually classified data sets of repositories to train classifiers capable of predicting if a given GitHub repository contains an engineered software project. The performance of the classifiers was evaluated using a set of 200 repositories with known ground truth classification. We also compared the performance of the classifiers to other approaches to classification (e.g. number of GitHub Stargazers) and found our classifiers to outperform existing approaches. We found stargazers-based classifier (with 10 as the threshold for number of stargazers) to exhibit high precision (97%) but an inversely proportional recall (32%). On the other hand, our best classifier exhibited a high precision (82%) and a high recall (86%). The stargazer-based criteria offers precision but fails to recall a significant portion of the population.
Keyword:
Mining software repositories
GitHub
Data curation
Curation tools
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Empirical Software Engineering 封面图
Empirical Software Engineering
IF:
3.6
论文数:
2.0K
被引数:
5.3K

机构

R
Rochester Institute of Technology
学者数:
3.8K
论文数: 3.3K
被引数: 45
U
University of Waterloo
学者数:
2.2W
论文数: 2.3W
被引数: 3.3W
引用论文

引用论文

Predicting patients' exposure to cyclosporin
err1996-08-01
err0
PREAI
errA. Johnston; J. M. Kovarik; E. A. Mueller; D. W. Holt
err分享
err收藏
err分享
err收藏
Programmed sample delivery on a pressurized paper
err2014-10-24
err0
errOAAI
errJoong Ho Shin; Juhwan Park; Seung Hoon Kim; Je-Kyun Park
err分享
err收藏
Does code decay? Assessing the evidence from change management data
err2001-01-01
err344
PREAI
errEick, SG; Graves, TL; Karr, AF; Marron, JS; Mockus, A
err分享
err收藏
Single Versus Double Anatomic Site Intraosseous Blood Transfusion in a Swine Model of Hemorrhagic Shock
err2021-11-01
err0
errOAAI
errEric Sulava; William Bianchi; Christian S. McEvoy; Paul J. Roszko; Gregory J. Zarow; Micah J. Gaspary; Ramesh Natarajan; Jonathan D. Auten
err分享
err收藏
Immunoregulation by interleukin‐12 in MB49.1 tumor‐bearing mice: Cellular and cytokine‐mediated effector mechanisms
err2005-12-06
err0
PREAI
errSharon E. Hunter; Kristine E. Waldburger; Deborah K. Thibodeaux; Robert G. Schaub; Samuel J. Goldman; John P. Leonard
err分享
err收藏
学者 查看更多内容