Return
DAVLMF-Seg: Vision-language model guided latent frequency-aware diffusion for semi-supervised medical image segmentation
J
Q
N
H
J
C
DOI:10.1016/j.neunet.2026.109438.png)
Abstract
En 中文
Medical image segmentation is a fundamental task in computer-aided diagnosis and treatment planning. Fully supervised methods achieve strong performance but rely on large-scale annotated datasets. Semi-supervised learning (SSL) alleviates this limitation by leveraging unlabeled data. However, existing SSL methods often lack effective semantic modeling and suffer from domain gaps between natural image pretraining and medical imaging. In this paper, we propose DAVLMF-Seg, a domain-adaptive vision-language model guided frequency-aware SSL framework. The method aligns medical images and textual descriptions in a shared latent space via parameter-efficient adaptation, providing semantic priors for pseudo-label refinement. We further introduce a frequency-domain conditioned diffusion module to progressively enhance feature fusion and reduce decoding ambiguity. An uncertainty-aware regularization strategy is also designed to improve confidence calibration. Extensive experiments on multiple benchmarks demonstrate consistent improvements over state-of-the-art methods. On ACDC, our method achieves 90.25% Dice with 10% labels (+1.2%) and 90.48% Dice with 20% labels (+0.8%), while HD95 is reduced by 2.3 and 1.0, respectively. On M&Ms and MyoPS, it attains 84.57% (+2.0%) and 75.23% (+1.1%) Dice, with HD95 reduced by 1.1 and 4.9, respectively. These results highlight the effectiveness of the proposed framework, especially under limited supervision and cross-domain scenarios.
Keywords:
Medical image segmentation
Semi-supervised learning
Vision-language model
Diffusion model
Frequency-aware fusion
Journal
IF:
6.3
Papers:
7.7K
Citations:
3.0W
