arrow
Return

Algorithm-hardware Co-optimization for Energy-efficient Drone Detection on Resource-constrained FPGA

delete2023-05-10
delete3
PRE
AI
H
Han-Sok Suh *
J
Jian Meng
T
Ty Nguyen
V
Vijay Kumar
Y
Yu Cao
J
Jae-sun Seo
DOI:10.1145/3583074delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Convolutional neural network (CNN)-based object detection has achieved very high accuracy; e.g., singleshot multi-box detectors (SSDs) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multidrone self-localization and formation control. In this article, we designed and co-optimized an algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained an SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, and throughput optimization. We evaluated the FPGA hardware for a custom drone dataset, Pascal VOC, and COCO2017. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy efficiency of 79 GOPS/W and throughput of 158 GOPS using the Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 1.1 to 8.7x higher energy efficiency than prior works that used the same Pascal VOC dataset, using the same FPGA device, but at a low-power consumption of 2.54W. For the COCO dataset, our MobileNet-V1 implementation achieved an mAP of 16.8, and 4.9 FPS/W for energy-efficiency, which is similar to 1.9x higher than prior FPGA works or other commercial hardware platforms.
Keywords:
FPGA accelerator
object detection
algorithm-hardware co-design
neural networks

Journal

ACM Transactions on Reconfigurable Technology and Systems cover
ACM Transactions on Reconfigurable Technology and Systems
IF:
2.8
Papers:
597
Citations:
810

Organization

A
Arizona State University
Scholars:
2.7W
Papers: 2.5W
Citations: 4.2W
A
arizona state university-tempe
Scholars:
1.5W
Papers: 1.2W
Citations: 13