arrow
Return

Named entity recognition and coreference resolution using prompt-based generative multimodal

delete2025-11-08
delete0
delete
OA
AI
Q
Qingchuan Zhang
Z
Zexi Song *
D
Delong Wang
Y
Yuanyuan Cai
M
Mingwen Bi *
左敏 cover
左敏 (Min Zuo) *
DOI:10.1007/s40747-025-02122-1delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Named Entity Recognition (NER) and Coreference Resolution (CR) are both of pivotal significance in the domain of natural language processing. The accurate detection of pronouns and noun phrases has been shown to have a direct impact on cross-sentence and even cross-modal semantic understanding. However, existing methods often suffer from inadequate semantic fusion in multimodal scenarios, leading to issues such as text-image mismatches and incorrect referential resolution. The present paper proposes a multimodal generation framework that integrates object detection with prompt-based strategies. First, a prompt mechanism is employed to embed entity types into the visual channel using object detection results and hierarchical features extracted by a visual Transformer. This design enhances multimodal NER. Then, the PGMCR model is constructed, which aligns the pronouns and noun phrases identified by NER with the image regions based on prompt information to achieve generative multimodal co-reference resolution. The experimental findings on the public CIN dataset demonstrate that PGMNER attains an F1 score of 96.55% in the NER task, which is notably higher than that of representative multimodal models. PGMCR attains a CoNLL F1 score of 62.53% on the CR task, thereby surpassing existing state-of-the-art methods. These results provide validation of the effectiveness of the proposed framework in cross-modal named entity recognition and co-reference resolution, thus providing new insights for multimodal semantic understanding in complex scenarios.
Keywords:
Named entity recognition
Coreference resolution
Generative
Bert
Object detection
Multimodal
Visual grounding
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Complex and Intelligent Systems cover
Complex and Intelligent Systems
IF:
4.6
Papers:
2.1K
Citations:
6.6K

Organization

B
Beijing Wuzi University
Scholars:
498
Papers: 473
Citations: 374
Cited Papers

Cited Papers

Multi-view semantic understanding for visual dialog
err2023-05-01
err1
PREAI
errJiang, Tianling; Zhang, Zefan; Li, Xin; Ji, Yi; Liu, Chunping
errShare
errSave
Multi-granularity cross-modal representation learning for named entity recognition on social media
err2024-01-01
err7
PREAI
errLiu, Peipei; Wang, Gaosheng; Li, Hong; Liu, Jie; Ren, Yimo; Zhu, Hongsong; Sun, Limin
errShare
errSave
Improved relation span detection in question answering systems over extracted knowledge bases
err2023-08-01
err5
PREAI
errBehmanesh, Somayyeh; Talebpour, Alireza; Shamsfard, Mehrnoush; Jafari, Mohammad Mahdi
errShare
errSave
Efficient relation extraction via quantum reinforcement learning
err2024-02-29
err0
errOAAI
errZhu, Xianchao; Mu, Yashuang; Wang, Xuetao; Zhu, William
errShare
errSave
Enriched entity representation of knowledge graph for text generation
err2022-11-04
err2
errOAAI
errShi, Kaile; Cai, Xiaoyan; Yang, Libin; Zhao, Jintao
errShare
errSave
researcher View more