Return
FM2: Fusing multiple foundation models for pathology image analysis via disentangled consensus-divergence representation
DOI:10.1016/j.inffus.2025.103840.png)
Abstract
En 中文
Foundation models (FMs) have emerged and achieved good performance on numerous downstream tasks. However, different FMs, like CLIP, DINOv2, and SAM, are trained on diverse datasets with varying methodologies, exhibiting model-specific characteristics and encoding scenario-specific knowledge. Efforts to unify the strengths of these different FMs through knowledge distillation show promise but remain challenging due to the inconsistencies in feature distributions, which can lead to suboptimal convergence and reduced generalizability. In this paper, we propose a novel aggregation framework, FM2 (Fusing Multiple Foundation Models), which leverages disentangled representation learning to address these challenges. Specifically, our approach effectively disentangles consensus and divergence features from multiple expert FMs and then aligns them into a unified and robust representation. Extensive experiments on datasets with over 1,000,000 pathology images across various tasks, including zero-shot and few-shot classification, cross-modal retrieval, and survival analysis, demonstrate that our method consistently outperforms state-of-the-art models, delivering superior accuracy and reliability across various clinical scenarios. Additionally, the visualizations offer insights into the model’s ability to harmonize knowledge across different FMs, highlighting its potential for enhancing diagnostic precision in medical imaging. The significant advancements demonstrated in our work underscore the promise of effectively aligning FMs, showing potential for broadening their application not only in pathology but also in other medical imaging domains.
Journal
IF:
15.5
Papers:
4.1K
Citations:
2.7W

