arrow
Return

Multi-Grained Radiology Report Generation With Sentence-Level Image-Language Contrastive Learning

delete2024-07-01
delete1
PRE
AI
A
Aohan Liu
Y
Yuchen Guo *
雍俊海 (Jun‐Hai Yong)
徐锋 cover
徐锋 (Feng Xu) *
DOI:10.1109/TMI.2024.3372638delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The automatic generation of accurate radiology reports is of great clinical importance and has drawn growing research interest. However, it is still a challenging task due to the imbalance between normal and abnormal descriptions and the multi-sentence and multi-topic nature of radiology reports. These features result in significant challenges to generating accurate descriptions for medical images, especially the important abnormal findings. Previous methods to tackle these problems rely heavily on extra manual annotations, which are expensive to acquire. We propose a multi-grained report generation framework incorporating sentence-level image-sentence contrastive learning, which does not require any extra labeling but effectively learns knowledge from the image-report pairs. We first introduce contrastive learning as an auxiliary task for image feature learning. Different from previous contrastive methods, we exploit the multi-topic nature of imaging reports and perform fine-grained contrastive learning by extracting sentence topics and contents and contrasting between sentence contents and refined image contents guided by sentence topics. This forces the model to learn distinct abnormal image features for each specific topic. During generation, we use two decoders to first generate coarse sentence topics and then the fine-grained text of each sentence. We directly supervise the intermediate topics using sentence topics learned by our contrastive objective. This strengthens the generation constraint and enables independent fine-tuning of the decoders using reinforcement learning, which further boosts model performance. Experiments on two large-scale datasets MIMIC-CXR and IU-Xray demonstrate that our approach outperforms existing state-of-the-art methods, evaluated by both language generation metrics and clinical accuracy.
Keywords:
Self-supervised learning
Decoding
Radiology
Biomedical imaging
Task analysis
Training
Reinforcement learning
Medical report generation
contrastive learning
multi-grained

Journal

IEEE Transactions on Medical Imaging cover
IEEE Transactions on Medical Imaging
IF:
9.8
Papers:
6.2K
Citations:
3.7W

Organization

T
tsinghua university
Scholars:
11.7W
Papers: 10.0W
Citations: 137