arrow
返回

Video Visual Relation Detection via 3D Convolutional Neural Network

delete2022-01-01
delete3
delete
OA
AI
M
Mingcheng Qu *
J
Jianxun Cui
苏
苏统华 (Tonghua Su)
W
Wenkai Shao
DOI:10.1109/ACCESS.2022.3154423delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video visual relation detection, which aims to detect the visual relations between objects in the form of relation triplet (e.g., person-ride-bike, dog-toward-car, etc.), is a significant and fundamental task in computer vision. However, most of the existing works about visual relation instances are focused on static images. Modeling the non-static relationships in videos has drawn little attention due to lacking large-scale video dataset support. In our work, we propose a video dataset named Video Predicate Detection and Reasoning (VidPDR) for dynamic video visual relation detection, which consists of 1,000 videos with dense manually dynamic labeled annotations on 21 object classes and 37 predicates classes. Moreover, we propose a novel spatio-temporal feature extraction framework with 3D Convolutional Neural Networks (ST3DCNN), which includes three modules 1) object trajectory, 2) short-term relation prediction, and 3) greedy relational association. We conducted appropriate experiments on public datasets and our own dataset (VidPDR). Results demonstrate that our proposed method has a great improvement in comparison to the state-of-the-art baselines.
Keyword:
Visualization
Feature extraction
Trajectory
Three-dimensional displays
Convolutional neural networks
Task analysis
Object detection
Computer vision
3D convolutional neural network
video visual relation detection

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
引用论文

引用论文

Beteiligung des peripheren Nervensystems bei Morbus Crohn
err1999-12-03
err0
PREAI
errB. Moormann; H. Herath; O. Mann; A. Ferbert
err分享
err收藏
err分享
err收藏
Trust in Virtual Teams: A Multidisciplinary Review and Integration
err2019-01-21
err0
errOAAI
errJanine Viol Hacker; Michael Johnson; Carol Saunders; Amanda L. Thayer
err分享
err收藏
The Utility of Combining the IAD and SES Frameworks
err2019-05-07
err0
errOAAI
errDaniel H. Cole; Graham Epstein; Michael D. McGinnis
err分享
err收藏
Interface Network Models for Complex Urban Infrastructure Systems
err2011-12-01
err0
PREAI
errJames Winkler; Leonardo Dueñas-Osorio; Robert Stein; Devika Subramanian
err分享
err收藏
3-D Relation Network for visual relation recognition in videos
err2021-04-01
err18
PREAI
errCao, Qianwen; Huang, Heyan; Shang, Xindi; Wang, Boran; Chua, Tat-Seng
err分享
err收藏
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations视觉基因组: 使用众包密集图像注释连接语言和视觉
err2017-02-06
err3.1K
errOAAI
errKrishna, Ranjay; Zhu, Yuke; Groth, Oliver; Johnson, Justin; Hata, Kenji; Kravitz, Joshua; Chen, Stephanie; Kalantidis, Yannis; Li, Li-Jia; Shamma, David A.; Bernstein, Michael S.; Li Fei-Fei
err分享
err收藏
学者 查看更多内容