arrow
Return

Predicting Visual Features From Text for Image and Video Caption Retrieval

delete2018-12-01
delete172
delete
OA
AI
D
Dong, Jianfeng
李锡荣 cover
李锡荣 (Xirong Li) *
C
Cees G. M. Snoek
DOI:10.1109/TMM.2018.2832602delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This paper strives to find amidst a set of sentences the one best describing the content of a given image or video. Different from existing works, which rely on a joint subspace for their image and video caption retrieval, we propose to do so in a visual space exclusively. Apart from this conceptual novelty, we contribute Word2VisualVec, a deep neural network architecture that learns to predict a visual feature representation from textual input. Example captions are encoded into a textual embedding based on multiscale sentence vectorization and further transferred into a deep visual feature of choice via a simple multilayer perceptron. We further generalize Word2VisualVec for video caption retrieval, by predicting from text both three-dimensional convolutional neural network features as well as a visual-audio representation. Experiments on Flickr8k, Flickr30k, the Microsoft Video Description dataset, and the very recent NIST TrecVid challenge for video caption retrieval detail Word2VisualVec's properties, its benefit over textual embeddings, the potential for multimodal query composition, and its state-of-the-art results.
Keywords:
Image and video
caption retrieval
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

U
university of amsterdam
Scholars:
6.0W
Papers: 5.1W
Citations: 94
R
Renmin University of China
Scholars:
8.1K
Papers: 7.7K
Citations: 1.1W
Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152
researcher View more organizations