返回
ChatAssert: LLM-Based Test Oracle Generation With External Tools Assistance
DOI:10.1109/TSE.2024.3519159.png)
摘要
En 中文
Test oracle generation is an important and challenging problem. Neural-based solutions have been recently proposed for oracle generation but they are still inaccurate. For example, the accuracy of the state-of-the-art technique teco is only 27.5% on its dataset including 3,540 test cases. We propose ChatAssert, a prompt engineering framework designed for oracle generation that uses dynamic and static information to iteratively refine prompts for querying large language models (LLMs). ChatAssert uses code summaries and examples to assist an LLM in generating candidate test oracles, uses a lightweight static analysis to assist the LLM in repairing generated oracles that fail to compile, and uses dynamic information obtained from test runs to help the LLM in repairing oracles that compile but do not pass. Experimental results using an independent publicly-available dataset show that ChatAssert improves the state-of-the-art technique, teco, on key evaluation metrics. For example, it improves Acc@1 by 15%. Overall, results provide initial yet strong evidence that using external tools in the formulation of prompts is an important aid in LLM-based oracle generation.
Keyword:
Chatbots
Codes
Measurement
Prompt engineering
Maintenance engineering
Large language models
Accuracy
Static analysis
Standards
Semantics
Test oracle generation
large language models (LLMs)
tool-augmented LLMs
prompt engineering framework
期刊
IF:
5.6
论文数:
2.8K
被引数:
1.1W
机构
引用论文
ATG24 Represses Autophagy and Differentiation and Is Essential for Homeostasy of the Flagellar Pocket in Trypanosoma brucei
PLOS ONE
IF0

