Return
PointDiff: Conditional Diffusion Model for Point-Supervised Video Moment Retrieval
DOI:10.1109/tmm.2026.3676163.png)
Abstract
En 中文
Point-supervised Video Moment Retrieval (PS-VMR) aims to retrieve a video moment that is semantically aligned with a given natural language query through the use of single-point annotations. Existing approaches suffer from sub-optimal performance mainly due to two aspects: i) the generated proposals are centered on the annotation frames which are actually randomly selected, and ii) the usage of multiple Gaussian distributions leads to boundary overstepped. To address these issues, we propose a generative framework for the PS-VMR task with the spirit of diffusion model. Specifically, we introduce a Gaussian prior constraint module to explore salient regions and leverage Gaussian weights to achieve compact pseudo-segments. Subsequently, the multi-modal conditional diffusion module diffuses the pseudo-segment into random noise, which is further denoised back to the initial segment with the guidance of similarity between query and video. It boasts an enhanced ability to capture video-query relationships and alleviate reliance on the annotation frame. Moreover, a noise intensity prediction task is integrated to assist the denoising process in the multi-modal conditional diffusion module. We conduct extensive experiments on three benchmarks of Charades-STA, ActivityNet Captions and TACoS, where the proposed network outperforms state-of-the-art methods, demonstrating its effectiveness for this task.
Keywords:
Diffusion model
point-supervised video moment retrieval (PS-VMR)
video understanding
Journal
IF:
9.7
Papers:
4.5K
Citations:
2.4W
Organization
Cited Papers
No cited papers available

