arrow
Return

Semantic-assisted report generation with memory enhanced transformer using context-aware visual extractor

delete2025-12-05
delete0
PRE
AI
P
Peketi Divya *
P
Partha Sarathi Chakraborty
C
C. Krishna Mohan
Y
Yen‐Wei Chen
DOI:10.1007/s10489-025-07031-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Medical imaging is crucial for clinical decisions, but manual report writing is time-consuming and error-prone. Automated generation helps professionals avoid errors and save time. Transformer models excel in capturing long-term dependencies, but face challenges in detailed image representation and modeling lengthy descriptions in medical reports. Addressing these challenges, we propose two modules. First, a transformable and context-aware attentive visual extractor (TCAVE) enhances image features by selecting informative spatial-semantic details across scales, improving report quality. TCAVE involves two novel networks in the ResNet 101 architecture, i.e., (i) incorporating a transformable network into ResNet 101’s intermediate layer for spatially invariant features and (ii) integrating a context-aware network into the final ResNet 101 layer for rich multi-scale contextual features. The transformer encoder encodes fine-grained radiology image details using these representations. Second, a selective memory module with relational gating SMMRG is integrated into the transformer decoder to efficiently model long-term dependencies between input-output sequences while maintaining contextual memory. Our model outperforms existing works on the IU X-Ray and MIMIC-CXR datasets.
Keywords:
Medical report generation
Transformers
Transformable and context-aware attentive visual extractor (TCAVE)
Selective memory module with relational gating (SMMRG)

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.5K
Citations:
1.7W

Organization

C
College of Information Science and Engineering
Scholars:
276
Papers: 131
Citations: 0