arrow
Return

A data-centric chip design agent framework for Verilog code generation

delete2025-11-01
delete0
PRE
AI
K
Kaiyan Chang *
W
Wenlong Zhu
K
Kun Wang
X
Xinyang He
N
Nan Yang
Z
Zhirong Chen
D
Dantong Jin
C
Cangyuan Li
Y
Yunhao Zhou
H
Hao Yan
Z
Zhuoliang Zhao
Y
Yuan Cheng
M
Mengdi Wang
S
Shengwen Liang
Y
Yinhe Han
X
Xiaowei Li
H
Huawei Li
Y
Ying Wang
DOI:10.1145/3727980delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advances in large language models (LLMs) have demonstrated significant potential for automated hardware description language (HDL) code generation from high-level specifications. However, two critical challenges limit further progress in this domain: the scarcity of quality Verilog training data and the inability of current approaches to generate RTL code optimized for power, performance, and area (PPA) metrics. This article presents a comprehensive data-centric framework that addresses these limitations through innovations in both pre-fine-tuning data preparation and after-fine-tuning optimization strategies. In the pre-fine-tuning phase, we tackle the data scarcity problem with an automated design-data augmentation framework that generates high-volume, high-quality natural language specifications aligned with corresponding Verilog code and EDA scripts. Our approach creates a complete RTL-level feedback loop by augmenting EDA scripts, RTL code, and EDA tool feedback. In the after-fine-tuning phase, we focus on generating PPA-aware RTL code through a novel search and prompt framework. Our approach implements iterative filtering and selection of LLM-generated Verilog variants while providing high-quality predefined prompts, including composition and interface specifications. To evaluate the effectiveness of our data augmentation method, we fine-tune Llama 2-13B and Llama 2-7B models using the dataset generated by our augmentation framework. The results demonstrate a significant improvement in the Verilog generation tasks with LLMs. Moreover, the accuracy of Verilog generation surpasses that of the current state-of-the-art open-source Verilog generation model, increasing from 58.8% to 70.6% with the same benchmark. Our 13B model has a pass rate improvement compared with GPT-3.5 in Verilog generation and outperforms in EDA script (i.e., SiliconCompiler) generation with only 200 EDA script data. Additionally, to evaluate the effectiveness of the our agent framework, we compare the PPA on the GPT-3.5, where the results show that the agent refined RTL code can have a better quality than the generated RTL code only with GPT-3.5.
Keywords:
Large language model
hardware generation
data augmentation

Journal

A
ACM Transactions on Design Automation of Electronic Systems
IF:
2
Papers:
112
Citations:
1.2K

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 75
I
institute of computing technology, cas
Scholars:
1.0K
Papers: 877
Citations: 1
C
Chinese Academy of Sciences
Scholars:
3.9W
Papers: 1.5W
Citations: 58.4W
researcher View more organizations