Return
Multi-source knowledge enhancement for multimodal semantic representation
DOI:10.1007/s13042-025-02888-3.png)
Abstract
En 中文
Semantic similarity is a fundamental task in natural language processing. Despite the impressive performances achieved by the existing knowledge enhancement methods in semantic similarity, they predominantly target at the unimodal information or depend solely on knowledge derived from a singular source. How to deal with the refined knowledge enhancement while well considering the complement between multimodal information is still a challenging problem, especially in the era of rapidly evolving large language models (LLMs), which offer new opportunities for knowledge enrichment and reasoning. To address this, we focus on semantic similarity representation in Chinese and introduce MuS $$^2$$ , a novel Multi-Source knowledge enhancement for Multimodal Semantic representation model, which integrates generative textual representations with retrieval-based visual features. By leveraging LLMs to enrich retrieved content and employing retrieval mechanisms to counteract LLM-induced hallucinations, MuS $$^2$$ achieves a synergistic balance between generation and retrieval. Specifically, the MuS $$^2$$ model comprises four core components: (1) A multi-source knowledge enhancement module that leverages LLMs for semantic textual enrichment and integrates image knowledge through search engine retrieval; (2) A dual-stream multimodal encoder processing textual and image modalities separately; (3) A multimodal fusion module aligning and combining the encoded representations; and (4) A semantic similarity calculation module generating pairwise lexical similarity metrics. Extensive experiments on four public datasets demonstrate the effectiveness of MuS $$^2$$ .
Keywords:
Semantic similarity
Multi-source knowledge enhancement
Multimodal fusion
Semantic representation
Journal
IF:
2.7
Papers:
3.1K
Citations:
5.6K

