arrow
返回

Learning to compress videos without computing motion

delete2022-04-01
delete3
delete
OA
AI
M
Meixu Chen *
T
Todd Goodall
A
Anjul Patney
A
Alan C. Bovik
DOI:10.1016/j.image.2022.116633delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Video has become an increasingly important part of our daily digital communication. With the development of higher resolution contents and displays, its significant volume poses significant challenges to the goals of acquiring, transmitting, compressing and displaying high quality video content. In this paper, we propose a new deep learning video compression architecture that does not require motion estimation, which is the most expensive element of modern hybrid video compression codecs like H.264 and HEVC. Our framework exploits the regularities inherent to video motion, which we capture by using displaced frame differences as video representations to train the neural network. In addition, we propose a new space-time reconstruction network based on both an LSTM model and a UNet model, which we call LSTM-UNet. The combined network is able to efficiently capture both temporal and spatial video information, making it highly amenable for our purposes. The new video compression framework has three components: a Displacement Calculation Unit (DCU), a Displacement Compression Network (DCN), and a Frame Reconstruction Network (FRN), all of which are jointly optimized against a single perceptual loss function. The DCU removes the need for motion estimation found in hybrid codecs, and is less expensive. In the DCN, an RNN-based network is utilized to compress displaced frame differences as well as retain temporal information between frames. The LSTMUNet is used in the FRN to learn space time differential representations of videos. Our experimental results show that our compression model, which we call the MOtionless VIdeo Codec (MOVI-Codec), learns how to efficiently compress videos without computing motion. Our experiments show that MOVI-Codec outperforms the Low-Delay P (LDP) veryfast setting of the video coding standard H.264 and exceeds the performance of the modern global standard HEVC codec, using the same setting, as measured by MS-SSIM, especially on higher resolution videos. In addition, our network outperforms the latest H.266 (VVC) codec at higher bitrates, when assessed using MS-SSIM, on high resolution videos. The MOVI-Codec project page can be found at https://github.com/Meixu-Chen/MOVI-Codec.
Keyword:
Video compression
Deep learning
Motion
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

S
Signal Processing and Image Communication
IF:
2.7
论文数:
2.8K
被引数:
4.2K

机构

U
university of texas austin
学者数:
2.4W
论文数: 2.0W
被引数: 54
U
university of texas system
学者数:
18.5W
论文数: 15.6W
被引数: 210
引用论文

引用论文

Judgment Capacity, Fear of Falling, and the Risk of Falls in Community-Dwelling Older Adults: The Progetto Veneto Anziani Longitudinal Study社区居住的老年人的判断能力,对跌倒的恐惧和跌倒的风险: Progetto Veneto Anziani纵向研究
err2020-06-01
err0
errOAAI
errCaterina Trevisan; Bruno M. Zanforlini; Stefania Maggi; Marianna Noale; Federica Limongi; Marina De Rui; Maria Chiara Corti; Egle Perissinotto; Anna-Karin Welmer; Enzo Manzato; Giuseppe Sergi
err分享
err收藏
Solvent complexes of the type [FeIII(CN)5L]n−
err1985-01-01
err0
PREAI
errGrażyna Stochel; Zofia Stasicka
err分享
err收藏
err分享
err收藏
Voice matters in a dictator game
err2007-07-14
err0
PREAI
errTetsuo Yamamori; Kazuhiko Kato; Toshiji Kawagoe; Akihiko Matsui
err分享
err收藏
err分享
err收藏
Identification and functional characterization of three type III polyketide synthases from Aquilaria sinensis calli
err2017-05-01
err0
PREAI
errXiaohui Wang; Zhongxiu Zhang; Xianjuan Dong; Yingying Feng; Xiao Liu; Bowen Gao; Jinling Wang; Le Zhang; Juan Wang; Shepo Shi; Pengfei Tu
err分享
err收藏
学者 查看更多内容