arrow
Return

Improving Cross-Language Code Clone Detection via Code Representation Learning and Graph Neural Networks

delete2023-11-01
delete3
PRE
AI
N
Nikita Mehrotra *
A
Akash Sharma
A
Anmol Jindal
R
Rahul Purandare *
DOI:10.1109/TSE.2023.3311796delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Code clone detection is an important aspect of software development and maintenance. The extensive research in this domain has helped reduce the complexity and increase the robustness of source code, thereby assisting bug detection tools. However, the majority of the clone detection literature is confined to a single language. With the increasing prevalence of cross-platform applications, functionality replication across multiple languages is common, resulting in code fragments having similar functionality but belonging to different languages. Since such clones are syntactically unrelated, single language clone detection tools are not applicable in their case. In this article, we propose a semi-supervised deep learning-based tool Rubhus, capable of detecting clones across different programming languages. Rubhus uses the control and data flow enriched abstract syntax trees (ASTs) of code fragments to leverage their syntactic and structural information and then applies graph neural networks (GNNs) to extract this information for the task of clone detection. We demonstrate the effectiveness of our proposed system through experiments conducted over datasets consisting of Java, C, and Python programs and evaluate its performance in terms of precision, recall, and F1 score. Our results indicate that Rubhus outperforms the state-of-the-art cross-language clone detection tools.
Keywords:
Codes
Cloning
Syntactics
Semantics
Java
Task analysis
Source coding
Program representation learning
cross-language code clone detection
graph-based neural networks
abstract syntax trees

Journal

IEEE Transactions on Software Engineering cover
IEEE Transactions on Software Engineering
IF:
5.6
Papers:
2.9K
Citations:
1.1W

Organization

I
Indraprastha Institute of Information Technology Delhi
Scholars:
933
Papers: 689
Citations: 558
University of Nebraska System cover
University of Nebraska System
Scholars:
2.7W
Papers: 2.3W
Citations: 58
Cited Papers

Cited Papers

Titanium ketimide complexes as α-olefin homo- and copolymerisation catalysts. X-ray diffraction structures of [TiCp′(NCtBu2)Cl2] (Cp′=Ind, Cp*)
err2004-01-01
err0
PREAI
errAlberto R. Dias; M. Teresa Duarte; Anabela C. Fernandes; Susete Fernandes; Maria M. Marques; Ana M. Martins; João F. da Silva; Sandra S. Rodrigues
errShare
errSave
errShare
errSave
Modeling Functional Similarity in Source Code With Graph-Based Siamese Networks
err2022-10-01
err17
errOAAI
errMehrotra, Nikita; Agarwal, Navdha; Gupta, Piyush; Anand, Saket; Lo, David; Purandare, Rahul
errShare
errSave
Lys63/Met1-hybrid ubiquitin chains are commonly formed during the activation of innate immune signalling
err2016-06-01
err0
errOAAI
errChristoph H. Emmerich; Siddharth Bakshi; Ian R. Kelsall; Juanma Ortiz-Guerrero; Natalia Shpiro; Philip Cohen
errShare
errSave
Recombinase polymerase amplification in the molecular diagnosis of microbiological targets and its applications
err2022-06-01
err0
PREAI
errD.S. Mota; J.M. Guimarães; A.M.D. Gandarilla; J.C.B.S. Filho; W.R. Brito; L.A.M. Mariúba
errShare
errSave
SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair
err2021-01-01
err217
errOAAI
errChen, Zimin; Kommrusch, Steve; Tufano, Michele; Pouchet, Louis-Noel; Poshyvanyk, Denys; Monperrus, Martin
errShare
errSave
XCODE: Towards Cross-Language Code Representation with Large-Scale Pre-Training
err2022-04-09
err8
PREAI
errLin, Zehao; Li, Guodun; Zhang, Jingfeng; Deng, Yue; Zeng, Xiangji; Zhang, Yin; Wan, Yao
errShare
errSave
researcher View more