arrow
Return

Towards multi-language repository-level code generation: From-scratch to guided tasks

delete2026-02-27
delete0
PRE
AI
J
Jingjing Liu
S
Silin Li
Z
Zeming Liu *
Z
Zihao Cheng
Y
Yuhang Guo
Y
Yuanfang Guo
Y
Yunhong Wang
H
Haifeng Wang
DOI:10.1016/j.neucom.2026.133204delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We introduce a repository-level code generation benchmark with multi-language support and varying task complexities which injects requirement comments as positive noise. • We propose RepoGenesis, a GRPO-based reinforcement learning framework with 3 distinct reward signals: (i) structure reward, (ii) syntactic reward, and (iii) semantic reward. These rewards respectively guide the repository structure, enforce syntactic correctness, and ensure the semantic correctness of the generated code. • Experiment results highlight the limitations of current LLMs and demonstrate that RepoGenesis enables Qwen2.5-Coder-7B to achieves the performance comparable to Claude Sonnet 4 (>100B).
Keywords:
code generation
repository-level
multi-language
reinforcement learning
semantic correctness

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

B
beihang university
Scholars:
5.2K
Papers: 2.0K
Citations: 21
I
inc.
Scholars:
286
Papers: 101
Citations: 3
B
Beijing Institute of Technology
Scholars:
5.2K
Papers: 2.1K
Citations: 6.0W
researcher View more organizations