1
Return

Human Activity Recognition using RGB-DVS Cameras: A Multi-modal Heat Conduction Model and A Benchmark Dataset

delete2026-08-01
delete0
PRE
AI
S
Shiao Wang
X
Xiao Wang *
B
Bo Jiang *
L
Lin Zhu
G
Guoqi Li
Y
Yaowei Wang
Y
Yonghong Tian
J
Jin Tang
DOI:10.1007/s11263-026-02936-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Human Activity Recognition (HAR) has long been a fundamental research direction in the field of computer vision. Previous studies have primarily relied on traditional RGB cameras to achieve high-performance activity recognition. However, the challenging factors in real-world scenarios, such as insufficient lighting and rapid movements, inevitably degrade the performance of RGB cameras. To address these challenges, biologically inspired event cameras, with their advantages such as high dynamic range and high temporal resolution, offer a promising solution to overcome the limitations of traditional RGB cameras. In this work, we rethink human activity recognition by combining RGB and event cameras. We first publish the large-scale multi-modal RGB-Event aligned human activity recognition benchmark dataset, termed HARDVS 2.0, which bridges the dataset gaps, in terms of modality diversity and real-world scenario coverage. The existing mainstream HAR methods have been retrained and evaluated on this dataset to provide a new platform for comparison. It contains 300 categories of everyday real-world actions with a total of 107,646 paired videos covering various challenging scenarios. Inspired by the physics-informed heat conduction model, we propose a novel multi-modal heat conduction operation framework for effective activity recognition, termed MMHCO-HAR. More in detail, given the RGB frames and event streams, we first extract the feature embeddings using a stem network (embedding layer). Then, these feature embeddings are then processed by a set of multi-modal heat conduction blocks, where the core component is the Heat Conduction Operation (HCO) layer. In the HCO layer, we fuse RGB and event features through a multi-modal DCT-IDCT layer while adaptively incorporating the thermal conductivity coefficient via Frequency Value Embeddings (FVEs) into this module. After that, we propose an adaptive fusion module based on a policy routing strategy for high-performance classification. We conduct comprehensive experiments comparing our proposed method with baseline methods on the HARDVS 2.0 dataset and other public datasets, achieving Top-1 accuracies of 53.2% on HARDVS 2.0 and 57.4% on the PokerEvent dataset, outperforming existing models such as Vision mamba by +1.4% and +0.9%, respectively. These results demonstrate that our method consistently performs well, validating its effectiveness and robustness. The source code and benchmark dataset will be released on  https://github.com/Event-AHU/HARDVS/tree/HARDVSv2 .
Keywords:
Event Camera
Human Activity Recognition
Multi-modal Learning
Physics-informed Heat Conduction
Signal Processing

Journal

International Journal of Computer Vision cover
International Journal of Computer Vision
IF:
9.3
Papers:
3.9K
Citations:
2.8W

Organization

U
University of Chinese Academy of Sciences
Scholars:
5.7K
Papers: 2.3K
Citations: 24.6W
S
School of Computer Science and Technology
Scholars:
1.3K
Papers: 513
Citations: 0
B
beijing institute of technology
Scholars:
5.3W
Papers: 3.9W
Citations: 63
P
Peng Cheng Laboratory
Scholars:
1.7K
Papers: 1.7K
Citations: 2.0K
H
Harbin Institute of Technology
Scholars:
1.1W
Papers: 3.8K
Citations: 8.5W
Cited Papers

Cited Papers

Citing Papers

Citing Papers