Return
Hypergraph-Guided Multimodal Prototype for Remote Sensing Scene Understanding
DOI:10.1109/TGRS.2025.3534288.png)
Abstract
En 中文
Noticeable achievements have been made in entity-level perception tasks (e.g., object detection) in remote sensing (RS) image interpretation. But for RS images carrying rich content, individual perception cannot well obtain the interaction patterns between entities. The recognition of relationships between entities is the key to deeply understanding RS scenes. In this article, we propose a hypergraph-guided multimodal prototype network (HMPNet), which performs relation recognition by matching relation representations with multimodal predicate prototypes. To overcome the imbalance of modal information in the matching process, a multimodal calibration strategy is devised, taking into account the image subprototype and text subprototype, which makes prediction results more reliable. Meanwhile, to align image and text subprototypes and explore relevant semantic patterns, the multimodal hypergraph is constructed to efficiently capture the associations between heterogeneous prototypes. Experimental results show that the performance of our model can reach the state-of-the-art (SOTA) level on the RS scene graph generation (SGG) task.
Keywords:
Remote sensing
Prototypes
Semantics
Cognition
Calibration
Reliability
Marine vehicles
Correlation
Accuracy
Object detection
Hypergraph network
multimodal prototypes
remote sensing (RS) scene understanding
scene graph generation (SGG)
Journal
IF:
8.6
Papers:
2.1W
Citations:
10.7W

