1
Return

Understanding Chain-of-Thought effectiveness in code generation: an empirical and information-theoretic analysis

delete2026-08-10
delete0
PRE
AI
N
Naizhu Jin
Z
Zhong Li *
G
Guang Yang
T
Tian Zhang
Q
Qingkai Zeng
DOI:10.1007/s10664-026-10947-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) achieve strong performance on code generation, but the mechanisms by which Chain-of-Thought (CoT) prompting helps remain unclear. We present a systematic empirical and proxy-based information-theoretic study of CoT effectiveness in neural code generation, evaluating five paradigms (Zero-Shot, Zero-Shot CoT, Self-Planning, Structured CoT, Reasoning-CoT) across six Python benchmarks, a multilingual benchmark with 12 programming languages, and six models from 7B to 480B parameters, using conditional mutual information I(Y; C|X) as a conceptual lens. Our results show that externally guided CoT consistently outperforms direct generation, with structured methods improving Pass@1 by 5–12% on average while using substantially fewer tokens than reflective reasoning, and that CoT benefits depend on language type systems and model capacity. We further find that reasoning quality is critical: high-quality structured CoT from strong generators yields consistently higher accuracy than lightweight alternatives with the same template, whereas naive Zero-Shot CoT can even degrade performance. These findings provide practical guidance for choosing CoT strategies based on model capacity, language characteristics, and task complexity.
Keywords:
Chain-of-Thought
Code generation
Empirical study

Journal

Empirical Software Engineering cover
Empirical Software Engineering
IF:
3.6
Papers:
1.9K
Citations:
5.3K

Organization

S
State Key Laboratory for Novel Software Technology
Scholars:
48
Papers: 18
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers