Return
An audio-visual multimodal adaptive balanced learning method based on gradient modulation
DOI:10.1007/s11227-025-07311-w.png)
Abstract
En 中文
To address the challenge encountered in audio-visual multimodal learning where heterogeneous learning rates result in one modality dominating the learning process and suppressing other modalities, thus weakening the multimodal collaborative decision-making process, a multimodal adaptive balanced learning method based on gradient modulation (AGM-CR) is proposed. This method introduces modulation coefficients to dynamically adjust the learning rates of different modalities based on gradient variations. A gradient balancing strategy is further employed by incorporating the gradient losses of individual modalities into the total loss as a regularization term to mitigate gradient disparities and balance the learning process. Experimental results show that AGM-CR improves classification accuracy by 3.1% and 1.3% on the CREMA-D and RAVDESS datasets, respectively, and reduces gradient fluctuations over multiple iterations, thereby enhancing training stability and accelerating convergence. Moreover, AGM-CR is a plug-and-play approach, offering greater flexibility and generalizability than the existing balancing methods do.
Keywords:
Balanced learning
Multimodal learning
Gradient modulation
Adaptive learning
Journal
IF:
2.7
Papers:
1.1K
Citations:
1.0W

