返回
Discriminability objective for training descriptive captions
DOI:10.1109/CVPR.2018.00728.png)
摘要
En 中文
One property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a loss component directly related to ability (by a machine) to disambiguate image/caption matches, we obtain systems that produce much more discriminative caption, according to human evaluation. Remarkably, our approach leads to improvement in other aspects of generated captions, reflected by a battery of standard scores such as BLEU, SPICE etc. Our approach is modular and can be applied to a variety of model/loss combinations commonly proposed for image captioning.
期刊
机构
引用论文
Substrate specificity and inhibitor analyses of human steroid 5β-reductase (AKR1D1)人类类固醇5 β-还原酶 (AKR1D1) 的底物特异性和抑制剂分析
Steroids
IF0
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉

