arrow
返回

Transformer-based networks over tree structures for code classification

delete2021-11-09
delete9
PRE
AI
W
Wei Hua *
G
Guangzhong Liu
DOI:10.1007/s10489-021-02894-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In software engineering (SE), code classification and related tasks, such as code clone detection are still challenging problems. Due to the elusive syntax and complicated semantics in software programs, existing traditional SE approaches still have difficulty differentiating between the functionalities of code snippets at the semantic level with high accuracy. As artificial intelligence (AI) techniques have increased in recent years, exploring different machine/deep learning techniques for code classification algorithms has become important. However, most existing machine/deep learning-based approaches often consider using convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to process code texts. However, the two networks inevitably suffer from gradient vanishing problems and fail to capture long-distance dependencies from code statements, resulting in poor performance in downstream tasks. In this paper, we propose the TBCC (Transformer-Based Code Classifier), a novel transformer-based neural network for programming language processing, which can avoid these two problems. Moreover, to capture the important syntactical features from programming languages, we split the deep abstract syntax trees (ASTs) into smaller subtrees that, aim to exploit syntactical information in code statements. We have applied the TBCC to two different common program comprehension tasks to verify its effectiveness: a code classification task for C programs and a code clone detection task for Java programs. The experimental results show that the TBCC achieves state-of-the-art performance, outperforming the baseline methods in terms of accuracy, recall, and F1 score. For subsequent research, the code of TBCC has been released *.
Keyword:
Code clone detection
Code classification
Code representation
Deep neural network
NLP
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

S
Shanghai Maritime University
学者数:
4.8K
论文数: 4.2K
被引数: 4.7K
引用论文

引用论文

err
IF0
err
err0
errOAAI
err
err分享
err收藏
SCALING WITH KNOWN UNCERTAINTY: A SYNTHESIS
err2006-01-01
err0
PREAI
errJIANGUO WU; HARBIN LI; K. BRUCE JONES; ORIE L. LOUCKS
err分享
err收藏
err分享
err收藏
学者 查看更多内容