arrow
返回

MIT-FRNet: Modality-invariant temporal representation learning-based feature reconstruction network for missing modalities

delete2024-09-01
delete0
PRE
AI
李嘉瑶 封面图
李嘉瑶 (Jiayao Li)
S
Saihua Cai
L
Li Li
R
Ruizhi Sun *
袁
袁刚 (Gang Yuan)
R
Rui Zhu
DOI:10.1016/j.eswa.2024.123655delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The investigation of missing modalities aims to extract valuable feature information from missing multi-modal data, it is a focal point in multi-modal learning. Existing missing modalities processing methods primarily focused on multi-modal fusion schemes to achieve optimal performance, but they faced the following two key challenges: (1) how to improve the robustness of incomplete multi-modal sequence representation, (2) how to effectively learn modality-invariant representations to mitigate heterogeneity between modalities. In this paper, we propose a modality-invariant temporal representation learning-based feature reconstruction network called MIT-FRNet for the missing modalities to tackle these challenges. In the MIT-FRNet, we first extract the latent features for each modality considering intra-modality and inter-modality, and then introduce an encoder-decoder framework to address the first challenge, it reconstructs missing element features via taking the incomplete modal sequences as input and implementing the inter-modal and cross-modal attention mechanisms for feature extraction. And then, through treating each timestamp as a single Gaussian distribution, we design a fine-grained similarity constraint based on distribution-level modality-invariant representations to learn effective modalityinvariant representations, thereby addressing the second challenge. Finally, the efficiency of proposed model is validated through the classification results after multi-modal fusion that involves using a gate encoder to pass it and followed by a vector fusion to fuse it. Extensive experiments on public benchmark datasets demonstrate that the proposed MIT-FRNet method achieves promising results under varying missing rates while exhibiting good convergence.
Keyword:
Feature reconstruction
Missing modalities
Modality -invariant representation
Data fusion

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

J
Jiangsu University
学者数:
4.0W
论文数: 2.8W
被引数: 5.5W
C
china agricultural university
学者数:
5.1W
论文数: 3.0W
被引数: 43
引用论文

引用论文

Myelin genes: getting the dosage right
err1995-11-01
err0
PREAI
errSteven S. Scherer; Phillip F. Chance
err分享
err收藏
err分享
err收藏
Sentiment Analysis of Comment Texts Based on BiLSTM
err2019-01-01
err314
errOAAI
errXu, Guixian; Meng, Yueting; Qiu, Xiaoyu; Yu, Ziheng; Wu, Xu
err分享
err收藏
err分享
err收藏
Lack of flagella disadvantages Salmonella enterica serovar Enteritidis during the early stages of infection in the rat
err2003-01-01
err0
errOAAI
errJeanette M. C. Robertson; Norma H. McKenzie; Michelle Duncan; Emma Allen-Vercoe; Martin J. Woodward; Harry J. Flint; George Grant
err分享
err收藏
学者 查看更多内容