返回
A universal cross language software similarity detector for open source software categorization
DOI:10.1016/j.jss.2019.110491.png)
摘要
En 中文
While there are novel approaches for detecting and categorizing similar software applications, previous research focused on detecting similarity in applications written in the same programming language and not on detecting similarity in applications written in different programming languages. Cross-language software similarity detection is inherently more challenging due to variations in language, application structures, support libraries used, and naming conventions. In this paper we propose a novel model, CroLSim, to detect similar software applications across different programming languages. We define a semantic relationship among cross-language libraries and API methods (both local and third party) using functional descriptions and a word-vector learning model. Our experiments show that CroLSim can successfully detect cross-language similar software applications, which outperforms all existing approaches (mean average precision rate of 0.65, confidence rate of 3.6, and 75% highly rated successful queries). Furthermore, we applied CroLSim to a source code repository to see whether our model can recommend cross-language source code fragments if queried directly with source code. From our experiments we found that CroLSim can recommend cross-language functional similar source code when source code is directly used as a query (average precision=0.28, recall=0.85, and F-Measure=0.40). (C) 2019 Published by Elsevier Inc.
Keyword:
API Calls
Doc2Vec
Cross-Language software similarity detection
Singular value decomposition
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.1
论文数:
5.5K
被引数:
8.4K
机构
引用论文
The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update用于可访问,可复制和协作生物医学分析的Galaxy平台: 2016更新
NUCLEIC ACIDS RESEARCH
IF13.1
Synthesizing Data Using Variational Autoencoders for Handling Class Imbalanced Deep Learning使用变分自编码器合成数据以处理类不平衡深度学习

