arrow
返回

Leveraging pre-trained language models for code generation

delete2024-02-29
delete2
delete
OA
AI
A
Ahmed Soliman *
S
Samir I. Shaheen
M
Mayada Hadhoud
DOI:10.1007/s40747-024-01373-8delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Code assistance refers to the utilization of various tools, techniques, and models to help developers in the process of software development. As coding tasks become increasingly complex, code assistant plays a pivotal role in enhancing developer productivity, reducing errors, and facilitating a more efficient coding workflow. This assistance can manifest in various forms, including code autocompletion, error detection and correction, code generation, documentation support, and context-aware suggestions. Language models have emerged as integral components of code assistance, offering developers the capability to receive intelligent suggestions, generate code snippets, and enhance overall coding proficiency. In this paper, we propose new hybrid models for code generation by leveraging pre-trained language models BERT, RoBERTa, ELECTRA, and LUKE with the Marian Causal Language Model. Selecting these models based on their strong performance in various natural language processing tasks. We evaluate the performance of these models on two datasets CoNaLa and DJANGO and compare them to existing state-of-the-art models. We aim to investigate the potential of pre-trained transformer language models to revolutionize code generation, offering improved precision and efficiency in navigating complex coding scenarios. Additionally, conducting error analysis and refining the generated code. Our results show that these models, when combined with the Marian Decoder, significantly improve code generation accuracy and efficiency. Notably, the RoBERTaMarian model achieved a maximum BLEU score of 35.74 and an exact match accuracy of 13.8% on CoNaLa, while LUKE-Marian attained a BLEU score of 89.34 and an exact match accuracy of 78.50% on DJANGO. Implementation of this work is available at https://github.com/AhmedSSoliman/Leveraging-Pretrained-Language-Models-for-Code-Generation.
Keyword:
Code generation
Code assistant
Language models
Marian model
RoBERTaMarian
LukeMarian

期刊

Complex and Intelligent Systems 封面图
Complex and Intelligent Systems
IF:
4.6
论文数:
2.1K
被引数:
6.6K

机构

E
egyptian knowledge bank (ekb)
学者数:
11.6W
论文数: 9.3W
被引数: 84
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Establishment and characterization of a new triple-negative canine mammary cancer cell line
err2018-10-01
err0
PREAI
errHong Zhang; Shimin Pei; Bin Zhou; Huanan Wang; Hongchao Du; Di Zhang; Degui Lin
err分享
err收藏
YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
Mutation analysis for evaluating code translation用于评价代码翻译的突变分析
err2023-12-06
err1
errOAAI
errGuizzo, Giovani; Zhang, Jie M.; Sarro, Federica; Treude, Christoph; Harman, Mark
err分享
err收藏
Interleukin-6 and cachexia inApcMin/+miceApcMin/ 小鼠的Interleukin-6和恶病质
err2008-02-01
err0
PREAI
errKristen A. Baltgalvis; Franklin G. Berger; Maria Marjorette O. Pena; J. Mark Davis; Stephanie J. Muga; James A. Carson
err分享
err收藏
学者 查看更多内容