Return
A Program Synthesis Dataset for LLM Temperature Analysis
DOI:10.1109/ACCESS.2025.3625443.png)
Abstract
En 中文
Large Language Models (LLMs) play an increasingly critical role in software engineering research, aiding tasks such as program synthesis, automated program repair, and test case generation. While extensive evaluations of LLMs exist, those results are mostly based on well-known and frequently used benchmarks, therefore, the available generated texts are mostly the same. This means the published generated texts do not provide additional data to the researcher community. Generated texts on well-known and frequently used benchmarks also introduce the risk that models may be inadvertently or deliberately optimized for those specific benchmarks, thereby distorting the assessment of their true generalization capabilities. This work introduces a dataset containing 18,900 raw and 18,896 processed LLM-generated outputs from nine open-source models across three model families (Llama, Qwen, DeepSeek). These models were evaluated on seven programming tasks sourced from the Sapientia ECN competition, ensuring diversity in available LLM-generated content. To capture the stochastic variability of LLM generation, inference was conducted three times per model while systematically varying the temperature parameter across 100 settings per run. Using all model and temperature variations on every task for the evaluation, we created a JSON file that contains information about 29.882 test cases that can be used simply by loading the JSON file.
Keywords:
Benchmark testing
Biological system modeling
Software engineering
Computational modeling
Analytical models
Maintenance engineering
Large language models
Training
Prompt engineering
Programming profession
evaluation
data

