返回
A Novel Spatio-Temporal 3D Convolutional Encoder-Decoder Network for Dynamic Saliency Prediction
DOI:10.1109/ACCESS.2021.3063372.png)
摘要
En 中文
As human beings are living in an always changing environment, predicting saliency maps from dynamic visual stimulus is of importance for modeling human visual system. Compared with human behavior, recent models based on LSTM and 3DCNN are still not good enough due to the limitation in spatio-temporal feature representation. In this paper, a novel 3D convolutional encoder-decoder architecture is proposed for saliency prediction on dynamic scenes. The encoder consists of two subnetworks to extract both spatial and temporal features in parallel with intermediate fusion, respectively. The saliency map is produced in decoder by firstly enlarging features in spatial dimensions and then aggregating temporal information. Specially designed structures can transfer pooling indices from encoder to decoder, which helps the generation of location-aware saliency maps. The proposed network can be trained and inferred in an end-to-end manner. Experimental results on benchmark DHF1K show that the proposed model achieves the state-of-the-art performance on key metrics including both normalized scanpath saliency and Pearson's correlation coefficient.
Keyword:
Feature extraction
Predictive models
Decoding
Computational modeling
Three-dimensional displays
Visualization
Solid modeling
Visual attention
dynamic saliency prediction
3D fully convolutional networks
spatio-temporal features
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Electrochemically assisted micro localized grafting of aptamers in a microchannel engraved in fluorinated thermoplastic polymer Dyneon THV
RSC Advances
IF0
Global and local sensitivity guided key salient object re-augmentation for video saliency detection
PATTERN RECOGNITION
IF7.6

