Return
CVAF: A CLIP-Based View-Consistent Alignment Framework for Aerial-Ground Person Re-Identification
D
S
L
DOI:10.1145/3785482.png)
Abstract
En 中文
With the increasing adoption of UAV platforms in areas such as public safety and smart cities, Aerial-Ground Person Re-Identification (AGPReID) has emerged as a crucial yet highly challenging task, garnering growing interest from the research community. While existing approaches have leveraged identity attributes and viewpoint disentanglement strategies to improve cross-view matching, their heavy reliance on prior knowledge often compromises model generalization. Furthermore, some methods that explicitly separate viewpoints may unintentionally discard identity-related, view-invariant features, leading to incomplete identity representations. To address these limitations, we propose a CLIP-based View-Consistent Alignment Framework (CVAF) with two training stages. In the first stage, learnable text tokens are employed to represent identity-aware textual descriptions. To promote consistent alignment across varying viewpoints, we introduce a Text Consistency Loss (TCL) that regularizes the stability of text-token interactions with multi-view images. In the second stage, we present a Semantic Filtering Module (SFM) that jointly modulates image patch tokens along spatial and channel dimensions. A text-guided cross-attention mechanism generates spatial attention maps to explicitly emphasize identity-relevant regions, while semantic matching between textual features and visual tokens enables adaptive reweighting of image representations, effectively suppressing background clutter and view-specific noise. Extensive experiments on multiple AGPReID datasets demonstrate that our CVAF outperforms the state-of-the-art methods.
Keywords:
Vision-language Learning
Aerial-Ground View
Person Re-Identification
Image Retrieval
Journal
IF:
6
Papers:
2.0K
Citations:
5.4K
