arrow
Return

Evaluating large language models for software testing

delete2025-04-01
delete0
PRE
AI
Y
Yihao Li *
P
Pan Liu *
H
Haiyang Wang
J
Jie Chu
W
W. Eric Wong *
DOI:10.1016/j.csi.2024.103942delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) have demonstrated significant prowess in code analysis and natural language processing, making them highly valuable for software testing. This paper conducts a comprehensive evaluation of LLMs applied to software testing, with a particular emphasis on test case generation, error tracing, and bug localization across twelve open-source projects. The advantages and limitations, as well as recommendations associated with utilizing LLMs for these tasks, are delineated. Furthermore, we delve into the phenomenon of hallucination in LLMs, examining its impact on software testing processes and presenting solutions to mitigate its effects. The findings of this work contribute to a deeper understanding of integrating LLMs into software testing, providing insights that pave the way for enhanced effectiveness in the field.
Keywords:
Large language models
LLM-driven testing
Evaluation
Hallucination

Journal

C
Computer Standards and Interfaces
IF:
3.1
Papers:
2.3K
Citations:
2.0K

Organization

L
Ludong University
Scholars:
5.6K
Papers: 3.3K
Citations: 3.7K
U
University of Texas Dallas
Scholars:
5.6K
Papers: 5.0K
Citations: 15
U
university of texas system
Scholars:
18.5W
Papers: 15.6W
Citations: 210
S
Shanghai Business School
Scholars:
302
Papers: 452
Citations: 481
researcher View more organizations