Return
Towards multi-language repository-level code generation: From-scratch to guided tasks
DOI:10.1016/j.neucom.2026.133204.png)
Abstract
En 中文
• We introduce a repository-level code generation benchmark with multi-language support and varying task complexities which injects requirement comments as positive noise. • We propose RepoGenesis, a GRPO-based reinforcement learning framework with 3 distinct reward signals: (i) structure reward, (ii) syntactic reward, and (iii) semantic reward. These rewards respectively guide the repository structure, enforce syntactic correctness, and ensure the semantic correctness of the generated code. • Experiment results highlight the limitations of current LLMs and demonstrate that RepoGenesis enables Qwen2.5-Coder-7B to achieves the performance comparable to Claude Sonnet 4 (>100B).
Keywords:
code generation
repository-level
multi-language
reinforcement learning
semantic correctness
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

