Return
Self-Supervised Aggregation Framework for Text-Attributed Heterogeneous Graphs Representation
DOI:10.1109/TKDE.2026.3683100.png)
Abstract
En 中文
Text-Attributed Heterogeneous Graphs (TAHGs) integrate topological relationships with rich textual node attributes, offering expressive representations for complex multi-faceted data. While recent methods jointly leverage textual and structural information, they still face two critical limitations: (i) existing approaches are constrained to neighborhood modeling, failing to capture semantic dependencies in higher-order topologies; (ii) current techniques exhibit inadequate unified alignment strategies, limiting dynamic interaction between cross modalities. To address these challenges, we propose SATH, a self-supervised information aggregation model for TAHGs, designed to effectively leverage textual and structural information within TAHGs. SATH aggregates higher-order neighbor textual attributes through comparative learning, and dynamically aligns these attributes to higher-order topologies through a unified strategy. This approach integrates both types of information effectively, enhancing the expressiveness and discriminative capability of the learned node representations in downstream tasks. Extensive experiments on real-world datasets demonstrate that SATH significantly outperforms baseline models while eliminating the need for manual meta-path design or text feature concatenation. It also improves efficiency and scalability on large-scale TAHGs, achieving superior representation quality in TAHG-based tasks.
Keywords:
Node representation
contrastive learning
self-supervised learning
text-attribute heterogeneous graph
Journal
IF:
10.4
Papers:
6.7K
Citations:
3.2W

