arrow
返回

Modality mixer exploiting complementary information for multi-modal action recognition

delete2025-05-01
delete0
PRE
AI
S
Sumin Lee
S
Sangmin Woo
M
Muhammad Adi Nugroho
C
Changick Kim *
DOI:10.1016/j.cviu.2025.104358delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
由于传感器的独特特性,每种模态表现出独特的物理属性。因此,在多模态动作识别的背景下,不仅要考虑整体动作内容,还要考虑不同模态之间的互补性。在本文中,我们提出了一种新型网络,命名为Modality Mixer(M-Mixer)网络,它有效地利用并整合了跨模态的互补信息与动作的时间上下文,用于动作识别。我们提出的M-Mixer的一个关键组件是Multi-modal Contextualization Unit(MCU),这是一个简单而有效的循环单元。我们的MCU负责将一种模态(例如RGB)的序列进行时间编码,并融合其他模态(例如深度和红外模态)的动作内容特征。这一过程促使M-Mixer网络利用全局动作内容,并补充其他模态的互补信息。此外,为了提取与给定模态设置相关的适当互补信息,我们引入了一个新模块,命名为Complementary Feature Extraction Module(CFEM)。CFEM为每种模态集成了单独的可学习查询嵌入,这些嵌入引导CFEM从其他模态中提取互补信息和全局动作内容。因此,我们提出的方法在NTU RGB+D 60、NTU RGB+D 120和NW-UCLA数据集上超越了当前最先进的方法。此外,通过全面的消融研究,我们进一步验证了我们提出方法的有效性。
Keyword:
Multi-modality
Action recognition
Recurrent unit

期刊

Computer Vision and Image Understanding 封面图
Computer Vision and Image Understanding
IF:
3.5
论文数:
450
被引数:
7.3K

机构

K
Korea Advanced Institute of Science and Technology
学者数:
3.7K
论文数: 1.4K
被引数: 254
引用论文

引用论文

D3D: Distilled 3D Networks for Video Action Recognition
err2020-03-01
err0
errOAAI
errJonathan C. Stroud; David A. Ross; Chen Sun; Jia Deng; Rahul Sukthankar
err分享
err收藏
Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding
err2022-12-01
err0
errOAAI
errMathew Monfort; Bowen Pan; Kandan Ramakrishnan; Alex Andonian; Barry A. McNamara; Alex Lascelles; Quanfu Fan; Dan Gutfreund; Rogerio Schmidt Feris; Aude Oliva
err分享
err收藏
Distillation Multiple Choice Learning for Multimodal Action Recognition
err2021-01-01
err0
PREAI
errNuno Cruz Garcia; Sarah Adel Bargal; Vitaly Ablavsky; Pietro Morerio; Vittorio Murino; Stan Sclaroff
err分享
err收藏
学者 查看更多内容