返回
Binarized Encoder-Decoder Network and Binarized Deconvolution Engine for Semantic Segmentation
DOI:10.1109/ACCESS.2020.3048375.png)
摘要
En 中文
Recently, semantic segmentation based on deep neural network (DNN) has attracted attention as it exhibits high accuracy, and many studies have been conducted on this. However, DNN-based segmentation studies focused mainly on improving accuracy, thus greatly increasing the computational demand and memory footprint of the segmentation network. For this reason, the segmentation network requires a lot of hardware resources and power consumption, and it is difficult to be applied to an environment where they are limited, such as an embedded system. In this paper, we propose a binarized encoder-decoder network (BEDN) and a binarized deconvolution engine (BiDE) accelerating the network to realize low-power, real-time semantic segmentation. BiDE implements a binarized segmentation network with custom hardware, greatly reducing the hardware resource usage and greatly increasing the throughput of network implementation. The deconvolution used for upsampling in a segmentation network includes zero padding. In order to enable deconvolution in a binarized segmentation network that cannot express zero, we introduce zero-aware binarized deconvolution which skips padded zero activations and zero-aware batch normalization embedded binary activation considering zero-skipped convolution. The BEDN, which is a binarized segmentation network proposed to be accelerated on BiDE, has acceptable accuracy while greatly reducing the computational and memory demands of the segmentation network through full-binarization and simple structure. BEDN has a network size of 0.21 MB, and its maximum memory usage is 1.38 MB. BiDE was implemented on Xilinx ZU7EV field-programmable gate array (FPGA) to operate at 187.5 MHz. BiDE accelerated the proposed BEDN within CamVid11 images of 480 x 360 size at 25.89 frames per second (FPS) achieving a performance of 1.682 Tera operations per second (TOPS) and 824 Giga operations per second per watt (GOPS/W).
Keyword:
Deconvolution
Image segmentation
Hardware
Semantics
Memory management
Acceleration
Training
Binarized neural network
binarized deconvolution
binarized segmentation network
zero-aware deconvolution
zero-skip deconvolution
neural network accelerator
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Neurocognitive enhancement in older adults: Comparison of three cognitive training tasks to test a hypothesis of training transfer in brain connectivity
NeuroImage
IF0
Electrochemically assisted micro localized grafting of aptamers in a microchannel engraved in fluorinated thermoplastic polymer Dyneon THV
RSC Advances
IF0
Distant regulatory elements in a Sox10‐βGEO BAC transgene are required for expression of Sox10 in the enteric nervous system and other neural crest‐derived tissuesSox10-βgeo BAC转基因中的远距离调控元件是肠神经系统和其他神经源性组织中 Sox10 表达所必需的

