Return
Comparative analysis of design pattern implementation validity in LLM-based code refactoring
DOI:10.1016/j.jss.2025.112519.png)
Abstract
En 中文
Design patterns are essential in software engineering, providing proven solutions for recurring design challenges, thereby enhancing maintainability, flexibility, and reusability of code. Despite their significance, the ability of Large Language Models (LLMs) to accurately implement these patterns has not been thoroughly explored. This research introduces a novel assessment framework that combines predicate logic specifications with quantitative metrics to evaluate pattern implementation quality. Using two case studies - a Point of Sale System (POSS) and Smart Wallet System (SWS) - we assess the LLMs’ capabilities in implementing design patterns including the Factory Method, Strategy, Composite, Observer, and Singleton patterns. The evaluation framework employs three metrics: Property Satisfaction Rate (PSR), Critical Property Coverage (CPC), and Pattern Implementation Quality Score (PIQS). The results demonstrate varying levels of effectiveness across the LLMs, with Claude achieving the highest average PIQS of 89.51, followed by Meta (88.98), ChatGPT (87.75), Copilot (82.69), and Gemini (71.04). These findings suggest that while LLMs show promise as refactoring tools, they are best utilized as assistive technologies rather than replacements for human developers.
Journal
IF:
4.1
Papers:
5.4K
Citations:
8.4K
Organization
No organization information available

