arrow
Return

SynthCoder: Anti-pattern identification and model training for FIM mode code completion

delete2026-09-28
delete0
PRE
AI
D
Dongjun Yu
X
Xiao Yan
Z
Zhenrui Li
J
Jipeng Xiao
H
Haochuan He
Y
Yongda Yu
H
Hao Zhang
G
Guoping Rong *
X
Xiaobo Huang
DOI:10.1007/s10664-026-10957-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As a leading application of large language models (LLMs) in software engineering, Fill-in-the-Middle (FIM) mode code completion has drawn wide attention. Training such models requires masking code corpora, yet common strategies tend to be problematic. For example, character-delimited random masking may yield many unrealistic cases (e.g., cutting keywords or identifiers), while purely AST-based masking cannot mask concurrent elements (subtrees) that span across multiple subtrees. These limitations easily create anti-patterns that rarely occur in real-world code completion scenarios, diminishing FIM performance. We introduce SynthCoder, which adopts optimized masking strategies better aligned with developer’s expectation in FIM. Specifically, we first refine AST-level node masking and add heuristics that better mimic developers’ expectation to construct the training corpora. Subsequently, SynthCoder-Seed and SynthCoder-Qwen, built upon Seed-Coder-8B-Base and Qwen2.5-Coder-7B respectively, employ a two-stage training pipeline, i.e., curriculum-based fine-tuning stage followed by Direct Preference Optimization (DPO) alignment stage with rejected code sampled preference data. Besides, to suppress erroneous context repetition, we include negative samples that duplicate existing code during DPO, mitigating code-echo failures where the model copies neighboring context instead of producing a valid completion. Extensive experiments on Santacoder-fim-task, aiXcoder-FIM-Evaluation and CrossCodeEval benchmarks show that the models trained with our mitigation strategies improve over mainstream baselines on Exact Match (EM) and Edit Similarity (ES) for text-based FIM benchmarks, and on Pass@1 for Santacoder-fim-task, a benchmark with test cases. SynthCoder also yields less code-echo and consumes fewer tokens at inference, yielding higher practical efficiency. Ablation studies further support the contribution of our optimized masking and repetition-suppression mechanisms.
Keywords:
LLM
Code completion
FIM
Post-training

Journal

Empirical Software Engineering cover
Empirical Software Engineering
IF:
3.6
Papers:
2.0K
Citations:
5.3K

Organization

No organization information available
Cited Papers

Cited Papers

err1999-01-01
err0
PREAI
errKhaled El Emam
errShare
errSave
Reliability in software engineering qualitative research through Inter-Coder Agreement
err2023-08-01
err3
errOAAI
errGonzalez-Prieto, Angel; Perez, Jorge; Diaz, Jessica; Lopez-Fernandez, Daniel
errShare
errSave
Training Language Models to Follow Instructions with Human Feedback
err2022-01-01
err0
PREAI
errAgarwal,Sandhini; Almeida,Diogo; Askell,Amanda; Christiano,Paul; Hilton,Jacob; Jiang,Xu; Kelton,Fraser; Leike,Jan; Lowe,Ryan; Miller,Luke; Mishkin,Pamela; Ouyang,Long; Ray,Alex; Schulman,John; Simens,Maddie; Slama,Katarina; Wainwright,Carroll; Welinder,Peter; Wu,Jeffrey; Zhang,Chong
errShare
errSave
errShare
errSave
CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion
err2023-01-01
err0
PREAI
errDing,Yangruibo; Wang,Zijian; Ahmad,Wasi; Ding,Hantian; Tan,Ming; Jain,Nihal; Ramanathan,Murali Krishna; Nallapati,Ramesh; Bhatia,Parminder; Roth,Dan; Xiang,Bing
errShare
errSave
Large Language Models for Software Engineering: A Systematic Literature Review
err2024-12-03
err6
PREAI
errHou, Xinyi; Zhao, Yanjie; Liu, Yue; Yang, Zhou; Wang, Kailong; Li, Li; Luo, Xiapu; Lo, David; Grundy, John; Wang, Haoyu
errShare
errSave
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
err2023-01-01
err0
PREAI
errErmon,Stefano; Finn,Chelsea; Manning,Christopher D; Mitchell,Eric; Rafailov,Rafael; Sharma,Archit
errShare
errSave
Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models
err2022-01-01
err0
PREAI
errBosma,Maarten; Chi,Ed; Ichter,Brian; Le,Quoc V; Schuurmans,Dale; Wang,Xuezhi; Wei,Jason; Xia,Fei; Zhou,Denny
errShare
errSave
no more