Return
TIMF: TabPFN-integrated multimodal framework for robust tabular-image learning
DOI:10.1016/j.neucom.2026.134975.png)
Abstract
En 中文
Tabular–image multimodal learning, which integrates structured tabular data with imaging data, holds significant promise for real-world applications, particularly in healthcare. However, two fundamental challenges remain: (1) the absence of standardized, pretrained representations for tabular data, in contrast to vision and language domains; and (2) the prevalence of missing values in tabular inputs, which complicates reliable multimodal integration in practice. To address these challenges, we propose the TabPFN-Integrated Multimodal Framework (TIMF), a novel framework that leverages TabPFN as a pretrained tabular foundation encoder and integrates it with visual representations from pretrained image backbones. By using TabPFN as a frozen tabular encoder, TIMF produces robust and informative tabular embeddings that naturally handle missing values and generalize well in data-limited settings. We systematically study different fusion strategies and tabular encoders, and evaluate TIMF on both natural and medical multimodal datasets. Experimental results demonstrate that TIMF consistently outperforms strong baselines across varying levels of missingness, highlighting the effectiveness of pretrained tabular representations for multimodal learning.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

