Return
Modality Interaction Decoupling: Spatiotemporal Memory-Driven Unified Multimodal Tracking Framework
DOI:10.1109/JIOT.2026.3664965.png)
Abstract
En 中文
Current designs of multimodal tracking networks primarily focus on spatial feature interaction within the backbone and lack the exploitation of temporal information. Although some approaches incorporate temporal cues by introducing sequential information from adjacent frames or employ updated temporal features during the feature extraction stage, they struggle to capture dynamic object variations and motion information in complex scenarios. To address these limitations, this article proposes a spatiotemporal memory-driven unified multimodal tracking framework (SMMTrack). Unlike existing works relying on feature interaction paradigms within the backbone network, this study innovatively introduces a decoupled backbone feature extraction framework. It deploys parallel, independent Vision Transformer (ViT) networks dedicated to extracting information from RGB and X modalities (RGB-T, RGB-D, and RGB-E), abandoning the conventional intrabackbone feature interaction. Furthermore, we introduce a memory mechanism during the feature extraction stage to enable long-term object modeling. In addition, a long-term memory storage and retrieval module is designed to dynamically update the Memory-List, thereby allowing the model to capture object appearance variations and motion trends comprehensively. SMMTrack is a unified framework across three tasks (RGB-T, RGB-D, and RGB-E tracking). Experimental results demonstrate that SMMTrack outperforms the state-of-the-art (SOTA) models, achieving outstanding performance in diverse multimodal tracking scenarios. The codes and results are released on https://github.com/qfxb/SMMTrack
Keywords:
Decoupled modality interaction
long-term object modeling
multimodal object tracking
spatiotemporal memory mechanism
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

