返回
Deep Learning for Image-to-Text Generation A technical overview
DOI:10.1109/MSP.2017.2741510.png)
摘要
En 中文
Generating a natural language description from an image is an emerging interdisciplinary problem at the intersection of computer vision, natural language processing, and artificial intelligence ( AI). This task, often referred to as image or visual captioning, forms the technical foundation of many important applications, such as semantic visual search, visual intelligence in chatting robots, photo and video sharing in social media, and aid for visually impaired people to perceive surrounding visual content. Thanks to the recent advances in deep learning, the AI research community has witnessed tremendous progress in visual captioning in recent years. In this article, we will first summarize this exciting emerging visual captioning area. We will then analyze the key development and the major progress the community has made, their impact in both research and industry deployment, and what lies ahead in future breakthroughs.
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.6
论文数:
1.1W
被引数:
1.7W
机构
引用论文
Nerve growth factor and its receptors TrkA and p75 are upregulated in the brain of mdx dystrophic mouse
Neuroscience
IF0
The role of a critical left fronto-temporal network with its right-hemispheric homologue in syntactic learning based on word category information基于单词类别信息的关键左前-时间网络及其右半球同源物在句法学习中的作用

