返回
Holistic Optimization Framework for FPGA Accelerators
DOI:10.1145/3769307.png)
摘要
En 中文
Customized accelerators have revolutionized modern computing by delivering substantial gains in energy efficiency and performance through hardware specialization. Field-Programmable Gate Arrays (FPGAs) play a crucial role in this paradigm, offering unparalleled flexibility and high-performance potential. High-Level Synthesis (HLS) and source-to-source compilers have simplified FPGA development by translating high-level programming languages into hardware descriptions enriched with directives. However, achieving high Quality of Results (QoR) remains a significant challenge, requiring intricate code transformations, strategic directive placement, and optimized data communication. This article presents Prometheus, a holistic optimization framework that integrates key optimizations-including task fusion, tiling, loop permutation, computation-communication overlap, and concurrent task execution-into a unified design space. By leveraging Non-Linear Programming (NLP) methodologies, Prometheus explores the optimization space under strict resource constraints, enabling automatic bitstream generation. Unlike existing frameworks, Prometheus considers interdependent transformations and dynamically balances computation and memory access. We evaluate Prometheus across multiple benchmarks, demonstrating its ability to maximize parallelism, minimize execution stalls, and optimize data movement. The results showcase its superior performance compared to state-of-the-art FPGA optimization frameworks, highlighting its effectiveness in delivering high QoR while reducing manual tuning efforts.
Keyword:
High-level synthesis
non-linear programming
compiler
期刊
A
IF:
2
论文数:
112
被引数:
1.2K
机构
引用论文
IronMan-Pro: Multiobjective Design Space Exploration in HLS via Reinforcement Learning and Graph Neural Network-Based ModelingIronMan-Pro:基于强化学习和图神经网络建模的高层次综合多目标设计空间探索
A pipelined and scalable dataflow implementation of convolutional neural networks on FPGA一种基于FPGA的卷积神经网络的流水线化和可扩展数据流实现

