arrow
返回

FSS: algorithm and neural network accelerator for style transfer

delete2023-10-10
delete0
PRE
AI
L
Ling Yi
Y
Yujie Huang
Y
Yujie Cai
Z
Zhaojie Li
M
Mingyu Wang *
W
Wenhong Li
X
Xiaoyang Zeng
DOI:10.1007/s11432-022-3676-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Neural networks (NNs), owing to their impressive performance, have gradually begun to dominate multimedia processing. For resource-constrained and energy-sensitive mobile devices, an efficient NN accelerator is necessary. Style transfer is an important multimedia application. However, existing arbitrary style transfer networks are complex and not well supported by current NN accelerators, limiting their application on mobile devices. Moreover, the quality of style transfer needs improvement. Thus, we design the FastStyle system (FSS), where a novel algorithm and an NN accelerator are proposed for style transfer. In FSS, we first propose a novel arbitrary style transfer algorithm, FastStyle. We propose a light network that contributes to high quality and low computational complexity and a prior mechanism to avoid retraining when the style changes. Then, we redesign an NN accelerator for FastStyle by applying two improvements to the basic NVIDIA deep learning accelerator (NVDLA) architecture. First, a flexible dat FSM and wt FSM are redesigned to enable the original data path to perform other operations (including the GRAM operation) by software programming. Moreover, statistics and judgment logic are designed to utilize the continuity of a video stream and remove the data dependency in the instance normalization, which improves the accelerator performance by 18.6%. The experimental results demonstrate that the proposed FastStyle can achieve higher quality with a lower computational cost, making it more suitable for mobile devices. The proposed NN accelerator is implemented on the Xilinx VCU118 FPGA under a 180-MHz clock. Experimental results show that the accelerator can stylize 512x512-pixel video with 20 FPS, and the measured performance reaches up to 306.07 GOPS. The ASIC implementation in TSMC 28 nm achieves about 22 FPS in the case of a 720-p video.
Keyword:
neural network accelerator
style transfer
neural network
deep learning

期刊

Science China Information Sciences 封面图
Science China Information Sciences
IF:
7.6
论文数:
4.9K
被引数:
8.9K

机构

F
fudan university
学者数:
11.8W
论文数: 7.7W
被引数: 121
引用论文

引用论文

Background K2P Channels KCNK3/9/15 Limit the Budding of Cell Membrane-derived Vesicles
err2011-07-15
err0
errOAAI
errDaniel Tsung-Ning Huang; Naiwen Chi; Shiou-Ching Chen; Ting-Ying Lee; Kate Hsu
err分享
err收藏
Factional Politics
err
IF0
err2012-01-01
err0
PREAI
errFrançoise Boucek
err分享
err收藏
Small towns shrinkage in the Jilin Province: A comparison between China and developed countries
err2020-04-13
err0
errOAAI
errYao Tong; Wei Liu; Chenggu Li; Jing Zhang; Zuopeng Ma
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
NMR studies in orthorhombic Fe3B1−xCx (0.1≤x≤0.4)
err1987-04-15
err0
PREAI
errY. D. Zhang; J. I. Budnick; F. H. Sanchez; W. A. Hines; D. P. Yang; J. D. Livingston
err分享
err收藏
学者 查看更多内容