arrow
返回

Feature Maps Correlation-based Video Quality Assessment

delete2024-01-11
delete1
PRE
AI
A
Amir Hossein Bakhtiari
A
Azadeh Mansouri *
DOI:10.1007/s11042-023-18068-wdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Blind video quality assessment (BVQA) techniques try to assess the perceived quality of a degraded video with no prior knowledge of the reference. Deep learning-based techniques have been used in different approaches so far. These methods frequently pool frame-level features to create a video representation and assess quality. The features are conventionally taken from the final convolutional layers of the network, or the mid-layers at times. Regardless of the details and information about the frames' appearance, such approaches generally assume that degradations affect the high-level features and general patterns taken from the last layers. The methods mentioned above mainly have to utilize ensemble techniques because of the relatively poor correlation between video quality and such features. We introduce a novel method in this study to acquire frame-level deep features for assessing the quality of videos. To accomplish this, we look at the deep feature maps correlations of specific layers of a pre-trained network, or more specifically, their similarities as helpful features for assessing video quality. The covariance matrix i.e. the Gram matrix, which depicts the correlation between all feature maps of a specific mid-layer, can be stated as deep feature relationships. The structural details of each frame's texture and color, in other words, frame's appearance, are reflected in these relations and significantly correlate with the perceived quality of a given video. In fact, the extracted feature maps relations in different granularities can effectively illustrate the influence of various distortions. The experimental results on three UGC video quality benchmarks, including YouTube-UGC, KoNViD-1k, and LIVE-VQC individual datasets depict acceptable results. As one can see, the resultant SROCCs using the proposed features extracted from the EfficientNet B4 network, show improvements of around 10%, 10%, and 7%, on YouTube-UGC, KoNViD-1k, and LIVE-VQC respectively, compared to typical features using last convolutional layers (avgpool). Moreover, the average SROCC results in 4 out of 6 cross-dataset tests is around 0.22% higher compared to the state-of-the-art where the SVR is trained on YouTube-UGC or KoNViD-1k. Thus, employing feature maps correlation of mid-layers of a pre-trained network as frame-level feature provides better cross-dataset results using the proposed computationally efficient method. The implementation of our method is available at https://github.com/amirh-bakhtiari/FMC-VQA.
Keyword:
Gram matrix
no-reference video quality assessment
convolutional neural network
Correlation of feature maps

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
2.0W
被引数:
3.2W

机构

K
Kharazmi University
学者数:
2.3K
论文数: 2.1K
被引数: 2.1K
引用论文

引用论文

err分享
err收藏
Making a Completely Blind Image Quality Analyzer制作全盲图像质量分析仪
err2013-03-01
err4.2K
PREAI
errMittal, Anish; Soundararajan, Rajiv; Bovik, Alan C.
err分享
err收藏
Non-invasive prenatal testing for Down syndrome
err2014-02-01
err0
PREAI
errPhilip Twiss; Melissa Hill; Rebecca Daley; Lyn S. Chitty
err分享
err收藏
err分享
err收藏
Kono-S Anastomosis for Surgical Prophylaxis of Anastomotic Recurrence in Crohn’s Disease: an International Multicenter Study
err2016-04-01
err0
PREAI
errToru Kono; Alessandro Fichera; Koutarou Maeda; Yoshiharu Sakai; Hiroki Ohge; Mukta Krane; Hidetoshi Katsuno; Mikihiro Fujiya
err分享
err收藏
学者 查看更多内容