Return
3D Vision-Language Models with Segmentation-Guided Multimodal Data for Spinal MRI Report Generation
DOI:10.1007/978-3-032-07502-4_14.png)
Abstract
En 中文
Automated radiology report generation using visionlanguage models (VLMs) holds significant promise for improving clinical workflow and diagnostic consistency. However, most existing approaches are limited to 2D image inputs and lack explicit incorporation of anatomical priors. In this study, we present a 3D-aware VLM framework designed to generate radiology reports from volumetric spine MRI scans while integrating anatomical segmentation masks to guide clinical relevance. We evaluate four input configurations: a baseline model using unsegmented MRI volumes, and three segmentation-aware variantsV1 (T1-weighted + segmentation), V2 (T2-weighted + segmentation), and V3 (T1- and T2-weighted + segmentation). Quantitative results across five evaluation metrics (BLEU, ROUGE-1, ROUGE-L, METEOR, and BERTScore) show that all segmentation-based variants significantly outperform the baseline, with V3 achieving the highest lexical accuracy and overall report quality. These findings underscore the value of spatial priors and multimodal fusion in improving the generation of structured, clinically meaningful spinal MRI reports.
Keywords:
Multimodal Language Models
Vision Language Model
Report Generation
Phi3
Spine Cord MRI
Journal
E
IF:
0
Papers:
14
Citations:
0

