Return
An empirical investigation of large language models for automated software compliance testing: Evidence from web accessibility
J
J
DOI:10.1016/j.infsof.2026.108215.png)
Abstract
En 中文
• Empirical evaluation of LLMs for automated compliance testing using 384 test cases across 39 WCAG criteria as a representative complex domain. • LLMs achieved accuracies ranging from 53.34% to 71.52% under provider-default configurations, compared to traditional tools’ 22.9%–33.9% real accuracy. • Systematic analysis of prompt engineering impact on software testing reliability, demonstrating 8.2% improvement through structured examples. • Identification of semantic complexity and context dependency as key factors determining amenability to AI-augmented testing. • Evidence for hybrid human-AI testing strategies, with LLMs addressing semantic gaps in current automated software quality assurance.
Keywords:
Software quality assurance
Automated compliance testing
Large language models
Web accessibility evaluation
Empirical software engineering
AI-augmented testing
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.3
Papers:
3.7K
Citations:
7.7K
Organization
No organization information available
