Return
Retrieval augmented generation for open-set entity alignment with large language models
DOI:10.1016/j.eswa.2026.132489.png)
Abstract
En 中文
Entity alignment aims to identify equivalent entities across knowledge graphs (KGs) and is essential for multi-source knowledge fusion. Recent embedding-based methods achieve strong performance by matching entities via learned representations, but they typically assume that every source entity has a counterpart in the target KG-a condition rarely satisfied in practice. They also struggle with many-to-one alignments and offer limited interpretability. To address these issues, we propose REALLM, a retrieval-augmented large language model (LLM) framework for open-set entity alignment. REALLM combines embedding-based retrieval with LLM-driven symbolic reasoning to both identify equivalent entity pairs with explanations and detect unmatchable entities. The framework first retrieves top-k candidates using semantic and structural similarities, then iteratively prompts an LLM to reason over each pair, filtering out non-equivalent matches and recognizing unmatched entities. To further mitigate many-to-one errors, REALLM incorporates a memory mechanism that records previous alignments and alerts the LLM when a candidate has already been matched. Experiments across multiple benchmark datasets show that REALLM outperforms state-of-the-art baselines, demonstrating the promise of LLMs for advancing open-set entity alignment.
Keywords:
Entity alignment
Knowledge graph
Large language model
Knowledge fusion
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

