返回
Heterogeneous tree structure classification to label Java programmers according to their expertise level
DOI:10.1016/j.future.2019.12.016.png)
摘要
En 中文
Open-source code repositories are a valuable asset to creating different kinds of tools and services, utilizing machine learning and probabilistic reasoning. Syntactic models process Abstract Syntax Trees (AST) of source code to build systems capable of predicting different software properties. The main difficulty of building such models comes from the heterogeneous and compound structures of ASTs, and that traditional machine learning algorithms require instances to be represented as n-dimensional vectors rather than trees. In this article, we propose a new approach to classify ASTs using traditional supervised-learning algorithms, where a feature learning process selects the most representative syntax patterns for the child subtrees of different syntax constructs. Those syntax patterns are used to enrich the context information of each AST, allowing the classification of compound heterogeneous tree structures. The proposed approach is applied to the problem of labeling the expertise level of Java programmers. The system is able to label expert and novice programs with an average accuracy of 99.6%. Moreover, other code fragments such as types, fields, methods, statements and expressions could also be classified, with average accuracies of 99.5%, 91.4%, 95.2%, 88.3% and 78.1%, respectively. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Big code
Machine learning
Syntax patterns
Abstract syntax trees
Programmer expertise
Decision trees
Big data
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
F
IF:
6.1
论文数:
6.9K
被引数:
2.3W
机构
引用论文
Lys63/Met1-hybrid ubiquitin chains are commonly formed during the activation of innate immune signallingLys63/Met1-hybrid泛素链通常在先天免疫信号激活过程中形成
Synthesizing Data Using Variational Autoencoders for Handling Class Imbalanced Deep Learning使用变分自编码器合成数据以处理类不平衡深度学习

