arrow
Return

Data augmentation in multimodal frameworks: a survey

delete2026-08-14
delete0
delete
OA
AI
D
Davi Fileti
M
Márcio Basgalupp *
J
João Gama
A
André Carvalho *
DOI:10.1007/s10462-026-11648-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Training machine learning models with more than one data modality has enhanced predictive performance in most contexts. Thus, many recent applications of machine learning use data from different sources and forms. Multimodal data augmentation (MMDA) addresses critical challenges in multimodal learning, such as data scarcity, modality imbalance, and cross-modal alignment. This survey systematically reviews 68 state-of-the-art MMDA approaches, and, as result, proposes a taxonomy for the area. For each revised paper, this survey analyzes its methodology, applications, and predictive performance gains, while highlighting key challenges, such as scalability and evaluation metrics. The proposed taxonomy provides a unified framework for understanding MMDA methods, their strengths, and limitations. This survey also identifies emerging trends, including the integration of large language models and diffusion processes, and outlines future research directions to advance multimodal learning.
Keywords:
Data augmentation
Multimodal
Generative AI
LLM
VLM
Survey

Journal

Artificial Intelligence Review cover
Artificial Intelligence Review
IF:
13.9
Papers:
6.1K
Citations:
1.9W

Organization

I
instituto de ciência e tecnologia
Scholars:
12
Papers: 6
Citations: 0
I
F
faculdade de economia
Scholars:
12
Papers: 8
Citations: 0
researcher View more organizations