Return
RAD: A retrieval-augmented distillation framework for lightweight video captioning
B
卫
S
W
DOI:10.1016/j.neucom.2026.134630.png)
Abstract
En 中文
Video captioning aims to automatically generate natural language descriptions from video content. To address the unstable supervision caused by the diversity and noise of multi-reference annotations in existing video captioning datasets, as well as the difficulty of lightweight models in capturing fine-grained temporal semantics under limited model capacity, this paper proposes a Retrieval-Augmented Distillation framework (RAD), which unifies high-quality global knowledge transfer with fine-grained semantic enhancement from within the video itself. Specifically, we first design a teacher-guided high-quality distillation target construction strategy that leverages the consensus semantics embedded in multi-reference annotations to select reliable global distillation anchors from teacher-generated candidate captions, thereby mitigating the effect of annotation noise during lightweight model training. We then construct a Temporal Semantic Memory Retrieval (TSMR) mechanism that explicitly organizes local semantic cues within the video via an integrated visual-semantic-temporal representation, thereby enhancing the model’s ability to capture fine-grained details. On this basis, we further introduce a retrieval-augmented distillation training strategy that incorporates retrieved auxiliary context into both the training and inference stages of the student model, thereby improving semantic completeness and narrative coherence in generated captions. Experimental results on the MSVD and MSR-VTT benchmarks show that RAD significantly improves generation performance while preserving the lightweight student backbone and maintaining relatively stable peak GPU memory. The additional inference overhead mainly comes from local caption generation in TSMR and can be controlled by adjusting the number of sampled frames, making RAD suitable for offline or latency-tolerant lightweight video captioning scenarios. Qualitative results further demonstrate its advantages in fine-grained action description and temporal reasoning.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
No organization information available
