返回
Boa: Ultra-Large-Scale Software Repository and Source-Code Mining
DOI:10.1145/2803171.png)
摘要
En 中文
In today's software-centric world, ultra-large-scale software repositories, such as SourceForge, GitHub, and Google Code, are the new library of Alexandria. They contain an enormous corpus of software and related information. Scientists and engineers alike are interested in analyzing this wealth of information. However, systematic extraction and analysis of relevant data from these repositories for testing hypotheses is hard, and best left for mining software repository (MSR) experts! Specifically, mining source code yields significant insights into software development artifacts and processes. Unfortunately, mining source code at a large scale remains a difficult task. Previous approaches had to either limit the scope of the projects studied, limit the scope of the mining task to be more coarse grained, or sacrifice studying the history of the code. In this article we address mining source code: (a) at a very large scale; (b) at a fine-grained level of detail; and (c) with full history information. To address these challenges, we present domain-specific language features for source-code mining in our language and infrastructure called Boa. The goal of Boa is to ease testing MSR-related hypotheses. Our evaluation demonstrates that Boa substantially reduces programming efforts, thus lowering the barrier to entry. We also show drastic improvements in scalability.
Keyword:
Boa
mining software repositories
domain-specific language
scalable
ease of use
lower barrier to entry
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
A
IF:
6.2
论文数:
1.2K
被引数:
3.4K
机构
引用论文
Homogeneous catalytic carbonylation of nitroaromatics Part III. Discovery and in situ high pressure FTIR spectral studies of a novel binuclear ruthenium catalyst硝基芳族化合物的均相催化羰基化部分III.新型双核钌催化剂的发现和原位高压FTIR光谱研究

