Return
DeepMEL: A multi-agent collaboration framework for multimodal entity linking
DOI:10.1016/j.ipm.2025.104507.png)
Abstract
En 中文
• We propose DeepMEL, a multi-agent framework based on Role-Specialized, whichfuses visual and context information through collaboration between LLMs andLVMs. To our knowledge, this is the first multi-agent approach applied to MEL. • The modality conversion alignment strategy for cross-modal representationalignmentthrough two paths: using a Context Summary Prompt with LLMs to generatefine-grained descriptions, and a Visual Q¥&A Prompt with LVMs to extract structuredimage representations. • We propose an Adaptive Iteration strategy that dynamically optimizes candidatesbycombining tool-based retrieval with LLM-based semantic reasoning. • We propose a unified prompt that reformulates MEL into a structuredcloze-styletask, simplifying input parsing and improving LLMs’ semantic understandingof multimodal disambiguation. • Substantial experiments on many datasets validate the superiority of our method
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W

