1
Return

Multi-View Chest X-Ray Vision-Language Pre-Training via Semantic-Aware Masked Language Modeling and High-Order Alignment

delete2026-06-05
delete0
PRE
AI
乔丽红 (Lihong Qiao)
J
Jingya Gong
Y
Yucheng Shu
L
Lifang Zhou
X
Ximing Xu
B
Baobin Li
W
Weisheng Li
雷柏英 (Baiying Lei)
DOI:10.1109/tmi.2026.3700856delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Chest X-Ray Vision-Language pretraining (VLP) leverages large-scale radiograph-report pairs to develop joint image-text representations, demonstrating significant potential for medical image diagnosis. However, existing VLP approaches often overlook the multi-view nature of chest X-Rays, and some multi-view methods apply uniform feature fusion, neglecting view-key semantic contributions. Moreover, random cross-modal Masked Language Modeling (MLM) fails to facilitate effective interactions, impeding representation alignment. Additionally, global alignment in VLP may lead to the false-negative problem. To address these limitations, we propose a novel medical VLP framework comprising three core components. First, a Key Semantics-enhanced Multi-view MLM module aggregates pathology-relevant patches across views, providing semantically rich supervision for MLM. A local semantics enhancing approach, which identifies and aggregates pathology-relevant key patches across views to guide MLM. Second, a Frontal-Lateral Alignment module extracts view-specific pathological features, ensuring semantic consistency and preserving critical information during aggregation. This module independently extracts pathological features from both views to preserve view-specific information while ensuring semantic consistency, which mitigates the loss of crucial information during aggregation. Third, a High-order Semantic Alignment approach mitigates false-negative issues by aligning features with semantically consistent clusters, enhancing global alignment through prototype-level semantics. Extensive experiments across seven public datasets demonstrate that our framework outperforms state-of-the-art methods in four downstream tasks, validating its efficacy. The code is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/sajiutea/F-L</uri>
Keywords:
Chest X-ray vision-language pre-training
multi-view representation learning
masked language modeling (MLM)
high-order semantic alignment

Journal

IEEE Transactions on Medical Imaging cover
IEEE Transactions on Medical Imaging
IF:
9.8
Papers:
6.2K
Citations:
3.7W

Organization

C
children's hospital of chongqing medical university
Scholars:
93
Papers: 31
Citations: 0
U
University of Chinese Academy of Sciences
Scholars:
5.7K
Papers: 2.3K
Citations: 24.6W
S
shenzhen university
Scholars:
4.4W
Papers: 3.4W
Citations: 72
C
Chongqing University of Posts and Telecommunications
Scholars:
2.2K
Papers: 876
Citations: 3.8K
Cited Papers

Cited Papers

Citing Papers

Citing Papers