1
Return

RAD: A retrieval-augmented distillation framework for lightweight video captioning

delete2026-08-01
delete0
PRE
AI
B
Bo Li
卫莹莹 (Yingying Wei)
S
Si Su
W
Wenti Huang *
DOI:10.1016/j.neucom.2026.134630delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video captioning aims to automatically generate natural language descriptions from video content. To address the unstable supervision caused by the diversity and noise of multi-reference annotations in existing video captioning datasets, as well as the difficulty of lightweight models in capturing fine-grained temporal semantics under limited model capacity, this paper proposes a Retrieval-Augmented Distillation framework (RAD), which unifies high-quality global knowledge transfer with fine-grained semantic enhancement from within the video itself. Specifically, we first design a teacher-guided high-quality distillation target construction strategy that leverages the consensus semantics embedded in multi-reference annotations to select reliable global distillation anchors from teacher-generated candidate captions, thereby mitigating the effect of annotation noise during lightweight model training. We then construct a Temporal Semantic Memory Retrieval (TSMR) mechanism that explicitly organizes local semantic cues within the video via an integrated visual-semantic-temporal representation, thereby enhancing the model’s ability to capture fine-grained details. On this basis, we further introduce a retrieval-augmented distillation training strategy that incorporates retrieved auxiliary context into both the training and inference stages of the student model, thereby improving semantic completeness and narrative coherence in generated captions. Experimental results on the MSVD and MSR-VTT benchmarks show that RAD significantly improves generation performance while preserving the lightweight student backbone and maintaining relatively stable peak GPU memory. The additional inference overhead mainly comes from local caption generation in TSMR and can be controlled by adjusting the number of sampled frames, making RAD suitable for offline or latency-tolerant lightweight video captioning scenarios. Qualitative results further demonstrate its advantages in fine-grained action description and temporal reasoning.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

No organization information available
Cited Papers

Cited Papers

Citing Papers

Citing Papers