Return
Data augmentation in multimodal frameworks: a survey
DOI:10.1007/s10462-026-11648-w.png)
Abstract
En 中文
Training machine learning models with more than one data modality has enhanced predictive performance in most contexts. Thus, many recent applications of machine learning use data from different sources and forms. Multimodal data augmentation (MMDA) addresses critical challenges in multimodal learning, such as data scarcity, modality imbalance, and cross-modal alignment. This survey systematically reviews 68 state-of-the-art MMDA approaches, and, as result, proposes a taxonomy for the area. For each revised paper, this survey analyzes its methodology, applications, and predictive performance gains, while highlighting key challenges, such as scalability and evaluation metrics. The proposed taxonomy provides a unified framework for understanding MMDA methods, their strengths, and limitations. This survey also identifies emerging trends, including the integration of large language models and diffusion processes, and outlines future research directions to advance multimodal learning.
Keywords:
Data augmentation
Multimodal
Generative AI
LLM
VLM
Survey
Journal
IF:
13.9
Papers:
6.1K
Citations:
1.9W

