arrow
返回

A Survey on Large Language Models for Code Generation

delete2026-02-01
delete33
PRE
AI
J
Jiang, Juyong
W
Wang, Fan
S
Shen, Jiasi
K
Kim, Sungju *
K
Kim, Sunghun *
DOI:10.1145/3747588delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Large Language Models (LLMs) have garnered remarkable advancements across diverse code-related tasks, known as Code LLMs, particularly in code generation that generates source code with LLM from natural language descriptions. This burgeoning field has captured significant interest from both academic researchers and industry professionals due to its practical significance in software development, e.g., GitHub Copilot. Despite the active exploration of LLMs for a variety of code tasks, either from the perspective of Natural Language Processing (NLP) or Software Engineering (SE) or both, there is a noticeable absence of a comprehensive and up-to-date literature review dedicated to LLM for code generation. In this survey, we aim to bridge this gap by providing a systematic literature review that serves as a valuable reference for researchers investigating the cutting-edge progress in LLMs for code generation. We introduce a taxonomy to categorize and discuss the recent developments in LLMs for code generation, covering aspects such as data curation, latest advances, performance evaluation, ethical implications, environmental impact, and real-world applications. In addition, we present a historical overview of the evolution of LLMs for code generation and provide a quantitative and qualitative comparative analysis of experimental results of code LLMs, sourced from their original papers to ensure a fair comparison on the HumanEval, MBPP, and BigCodeBench benchmarks, across various levels of difficulty and types of programming tasks, to highlight the progressive enhancements in LLM capabilities for code generation. We identify critical challenges and promising opportunities regarding the gap between academia and practical development. Furthermore, we have established a dedicated resource GitHub page (https://github.com/juyongjiang/CodeLLMSurvey) to continuously document and disseminate the most recent advances in the field.
Keyword:
Large Language Models
Code Large Language Models
Code Generation

期刊

A
ACM Transactions on Software Engineering and Methodology
IF:
6.2
论文数:
1.2K
被引数:
3.4K

机构

H
hong kong university of science & technology (guangzhou)
学者数:
214
论文数: 92
被引数: 0
引用论文

引用论文

Taking Flight with Copilot与副驾驶一起飞行
err2023-01-26
err0
errOAAI
errChristian Bird; Denae Ford; Thomas Zimmermann; Nicole Forsgren; Eirini Kalliamvakou; Travis Lowdermilk; Idan Gazit
err分享
err收藏
Unveiling Memorization in Code Models
err2024-04-12
err0
errOAAI
errZhou Yang; Zhipeng Zhao; Chenyu Wang; Jieke Shi; Dongsun Kim; Donggyun Han; David Lo
err分享
err收藏
A Methodology for Controlling Bias and Fairness in Synthetic Data Generation
err2022-05-04
err0
errOAAI
errEnrico Barbierato; Marco L. Della Vedova; Daniele Tessera; Daniele Toti; Nicola Vanoli
err分享
err收藏
学者 查看更多内容