arrow
返回

Hypercube Pooling for Visual Semantic Embedding

delete2024-11-18
delete0
delete
OA
AI
王宏镔 封面图
王宏镔 (Hongbin Wang)
R
Rui Tang
李
李凡 (Fan Li) *
DOI:10.1145/3689637delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Visual Semantic Embedding (VSE) is a primary model for cross-modal retrieval, wherein the global feature aggregator is a crucial component of the VSE model. In recent research, the General Pooling Operator (GPO) aggregator, which weighs the features reconstructed from the local feature set to aggregate, facilitates the related models to achieve good retrieval performance. However, the reason for the effectiveness remains to be explored. To enhance the rationality of aggregator designs, we analyze the reason from the perspective of feature space. Indeed, for each data, the local feature set forms a hypercube containing abundant data information, and the feature learned by GPO measures the hypercube, thereby representing the data. The geometric structure of the hypercube implies that the set containing all points within the hypercube is a convex set, so the feature learned by weighted aggregation is an interior point of the hypercube. However, using the interior point to measure the hypercube leads to some problems in feature representation and model optimization, as well as the reduction of retrieval efficiency caused by weight computation. For example, the related pair's features may be far, while the unrelated ones may be close. To measure the hypercube more clearly and alleviate the problems mentioned above, we propose Hypercube Pooling (HCP) aggregator. Specifically, HCP concatenates the Max and Min Pooling features as the global features. This aggregation method has multiple advantages, e.g., the learned global feature represents all hyperplanes of the hypercube that contain critical information and hypercube geometric structure. Moreover, HCP adds normalizationbefore-concatenation and reduces the usual setting of margin in the loss function by half to avoid gradient loss caused by the difference in the feature value and dimensionality doubling. The experimental results on the Flickr30K and MSCOCO datasets show that the HCP model has excellent performance with high efficiency, confirming the correctness of the spatial analysis and the effectiveness of the HCP aggregator.
Keyword:
Cross-Modal Retrieval
Visual Semantic Embedding
Global Pooling
Hypercube

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

暂无机构信息
引用论文

引用论文

The Social Impact of Events in Social Media Conversation
err2014-12-27
err0
PREAI
errAlessandro Inversini; Rogan Sage; Nigel Williams; Dimitrios Buhalis
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
学者 查看更多内容