Return
MathCoRL: Structured function-prototype prompting with policy-guided exemplars for efficient mathematical reasoning
DOI:10.1016/j.compeleceng.2026.111103.png)
Abstract
En 中文
The proficiency of Large Language Models (LLMs) in mathematical reasoning is evident; however, their reasoning traces often combine symbolic reasoning and computation, resulting in traces that are complex, lengthy, and difficult to verify. To overcome this challenge, this work proposes MathCoRL, which uses Function-Prototype Prompting (FPP) and a lightweight policy-guided exemplar selection approach. FPP limits the reasoning process to a predefined set of typed Python functions, making reasoning traces short and easily verifiable, while a policy network trained using reinforcement learning in context fine-tunes exemplar selection to optimize mathematical reasoning with shorter reasoning traces. Empirical evaluations of MathCoRL were conducted on several arithmetic-intensive mathematical reasoning benchmarks, including GSM8K, SVAMP, TabMWP, TAT-QA, and FinQA. The results show that MathCoRL achieves high accuracy and improved interpretability compared with representative program-aided baselines such as PAL and PoT, attaining 95.58% on GSM8K, 87.21% on TAT-QA, and 77.62% on FinQA, while producing substantially shorter reasoning traces. Additional efficiency analyses demonstrate that MathCoRL maintains competitive latency while reducing effective output token usage relative to strong baseline methods. MathCoRL therefore provides a lightweight and interpretable framework for mathematical reasoning with LLMs, emphasizing concise reasoning traces that improve inference efficiency and latency.
Keywords:
Mathematical reasoning
Large Language Models
Function-Prototype Prompting
Policy-guided exemplar selection
Efficient reasoning
Journal
C
IF:
4.9
Papers:
6.7K
Citations:
1.3W

