1
Return

An empirical investigation of large language models for automated software compliance testing: Evidence from web accessibility

delete2026-06-05
delete0
delete
OA
AI
J
Juan Miguel López *
J
Juanan Pereira
DOI:10.1016/j.infsof.2026.108215delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• Empirical evaluation of LLMs for automated compliance testing using 384 test cases across 39 WCAG criteria as a representative complex domain. • LLMs achieved accuracies ranging from 53.34% to 71.52% under provider-default configurations, compared to traditional tools’ 22.9%–33.9% real accuracy. • Systematic analysis of prompt engineering impact on software testing reliability, demonstrating 8.2% improvement through structured examples. • Identification of semantic complexity and context dependency as key factors determining amenability to AI-augmented testing. • Evidence for hybrid human-AI testing strategies, with LLMs addressing semantic gaps in current automated software quality assurance.
Keywords:
Software quality assurance
Automated compliance testing
Large language models
Web accessibility evaluation
Empirical software engineering
AI-augmented testing
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Information and Software Technology cover
Information and Software Technology
IF:
4.3
Papers:
3.7K
Citations:
7.7K

Organization

No organization information available
Cited Papers

Cited Papers

Citing Papers

Citing Papers