1
Return

Semantics-Aware Spatial-Temporal Binaries for Cross-Modal Video Retrieval

delete2021-01-01
delete41
PRE
AI
M
Mengshi Qi
Q
Qin, Jie
Y
Yi Yang
Y
Yunhong Wang *
J
Jiebo Luo
DOI:10.1109/TIP.2020.3048680delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the current exponential growth of video-based social networks, video retrieval using natural language is receiving ever-increasing attention. Most existing approaches tackle this task by extracting individual frame-level spatial features to represent the whole video, while ignoring visual pattern consistencies and intrinsic temporal relationships across different frames. Furthermore, the semantic correspondence between natural language queries and person-centric actions in videos has not been fully explored. To address these problems, we propose a novel binary representation learning framework, named Semanticsaware Spatial-temporal Binaries (S(2)Bin), which simultaneously considers spatial-temporal context and semantic relationships for cross-modal video retrieval. By exploiting the semantic relationships between two modalities, S(2)Bin can efficiently and effectively generate binary codes for both videos and texts. In addition, we adopt an iterative optimization scheme to learn deep encoding functions with attribute-guided stochastic training. We evaluate our model on three video datasets and the experimental results demonstrate that S(2)Bin outperforms the state-of-the-art methods in terms of various cross-modal video retrieval tasks.
Keywords:
Semantics
Binary codes
Feature extraction
Visualization
Task analysis
Natural languages
Stochastic processes
Cross-modal hashing
video retrieval
binary representation
spatial-temporal features
natural language
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

B
Beihang University
Scholars:
5.0W
Papers: 4.0W
Citations: 37
U
University of Rochester
Scholars:
2.6W
Papers: 2.1W
Citations: 2.2W
E
Ecole Polytechnique Federale de Lausanne
Scholars:
1.7W
Papers: 1.3W
Citations: 25
U
university of technology sydney
Scholars:
1.6W
Papers: 2.0W
Citations: 25
Cited Papers

Cited Papers

Citing Papers

Citing Papers