Return
Memory Experts Aggregation for Visual Place Recognition
DOI:10.1142/S2737480725500402.png)
Abstract
En 中文
Visual Place Recognition (VPR) poses significant challenges due to the need for simultaneous comprehension of macro-level semantic layouts and micro-level discriminative details. Traditional single-scale feature representations struggle to meet these multi-granularity cognitive demands. To address this, we propose a novel multi-scale feature fusion strategy that effectively integrates high-level semantic context with spatially precise shallow-layer features, significantly enhancing recognition accuracy in structurally similar environments. Additionally, we overcome the computational inefficiency inherent in conventional Vision Transformers (ViTs) by introducing a specialized cross-attention mechanism augmented with memory expert modules. Inspired by human visual cognition, these modules selectively attend to key visual landmarks, progressively accumulating and transferring discriminative knowledge across tasks. This approach achieves superior recognition performance while substantially reducing computational complexity.
Keywords:
Visual place recognition
cross attention
visual location
feature fusion
Journal
G
IF:
1.5
Papers:
17
Citations:
0

