Return
Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph Captioning
S
X
X
L
盛
A
DOI:10.1109/tmm.2026.3668632.png)
Abstract
En 中文
Video paragraph captioning (VPC) aims to generate coherent, detailed narratives that accurately reflect a video's content. However, existing methods typically depend on coarse-grained event correlations and neglect the nuanced spatio-temporal interactions critical for comprehensive understanding. Refined verbs and prepositions, encoding actions and spatial relations, are essential for clear, consistent descriptions. To address these issues, we propose the Fine-Grained Lexical-Centric Semantic Network (FLS-Net), which emphasizes verbs and prepositions linked to salient objects to improve spatio-temporal coherence across events. FLS-Net integrates a multi-lexical synergy mechanism, leveraging nouns obtained via multi-modal matching, and employs a Verb-Guided Event Consistency Module (VECM) alongside a Preposition-Driven Relation Representation Module (PRRM). A cyclic encoder-decoder architecture further enforces event consistency, significantly boosting VPC performance. Extensive experiments on <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">ActivityNet Captions</small> and <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">YouCook2</small> demonstrate FLS-Net's superiority over state-of-the-art approaches.
Keywords:
Video paragraph captioning
multi-lexical synergy
verb-guided event consistency
preposition-driven relation representation
spatio-temporal interactions
Journal
IF:
9.7
Papers:
4.4K
Citations:
2.4W
