arrow
返回

Temporal Dynamic Concept Modeling Network for Explainable Video Event Recognition

delete2023-07-12
delete0
delete
OA
AI
W
Weigang Zhang
Z
Zhaobo Qi *
S
Shuhui Wang
C
Chi Su
L
Li Su
Q
Qingming Huang *
DOI:10.1145/3568312delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Recently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution (TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution (CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network (TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field.
Keyword:
Event recognition
temporal concept receptive field
dynamic convolution

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
学者 查看更多机构
引用论文

引用论文

Spreading of extrinsic grain boundary dislocations in plastically deformed aluminium
err1978-02-16
err0
PREAI
errR. A. Varin; J. W. Wyrzykowski; W. Łojkowski; M. W. Grabski
err分享
err收藏
Fast Semantic Diffusion for Large-Scale Context-Based Image and Video Annotation
err2012-06-01
err37
errOAAI
errJiang, Yu-Gang; Dai, Qi; Wang, Jun; Ngo, Chong-Wah; Xue, Xiangyang; Chang, Shih-Fu
err分享
err收藏
Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
err分享
err收藏
Discovering Latent Discriminative Patterns for Multi-Mode Event Representation
err2019-06-01
err6
PREAI
errXie, Wenlong; Yao, Hongxun; Sun, Xiaoshuai; Han, Tingting; Zhao, Sicheng; Chua, Tat-Seng
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
学者 查看更多内容