Return
Brain-inspired learning to deeper inductive reasoning for video captioning
DOI:10.1007/s13042-023-01876-9.png)
Abstract
En 中文
Video captioning requires deeply understanding video content, describing the video concisely and accurately in one sentence. Since the video usually contains multiple atomic events, conventional methods using the attention mechanism or alignment between frame and word, lack deep inductive reasoning for multiple motions and appearances. Considering the inductive reasoning mechanism of the human brain, a brain-inspired deeper inductive reasoning model(DIR) is proposed in this paper. The DIR model discusses the inductive reasoning to presents the semantic similarity and dissimilarity of multiple atomic events, describing the video concisely and accurately. We evaluate the effectiveness of our method on public benchmarks (MSVD and MSR-VTT). Extensive experiments demonstrate that DIR outperforms general state-of-the-art methods, and show the advantages in deep reasoning compared with traditional captioning models.
Keywords:
Video captioning
Brain-inspired
Inductive reasoning
Semantic analysis
Journal
IF:
2.7
Papers:
3.1K
Citations:
5.6K

