返回
Human activity recognition in videos using a single example
DOI:10.1016/j.imavis.2013.08.005.png)
摘要
En 中文
This paper presents a novel approach for action recognition, localization and video matching based on a hierarchical codebook model of local spatio-temporal video volumes. Given a single example of an activity as a query video, the proposed method finds similar videos to the query in a target video dataset. The method is based on the bag of video words (BOV) representation and does not require prior knowledge about actions, background subtraction, motion estimation or tracking. It is also robust to spatial and temporal scale changes, as well as some deformations. The hierarchical algorithm codes a video as a compact set of spatio-temporal volumes, while considering their spatio-temporal compositions in order to account for spatial and temporal contextual information. This hierarchy is achieved by first constructing a codebook of spatio-temporal video volumes. Then a large contextual volume containing many spatio-temporal volumes (ensemble of volumes) is considered. These ensembles are used to construct a probabilistic model of video volumes and their spatio-temporal compositions. The algorithm was applied to three available video datasets for action recognition with different complexities (KTH. Weizmann, and MSR II) and the results were superior to other approaches, especially in the case of a single training example and cross-dataset(1) action recognition. (C) 2013 Elsevier B.V. All rights reserved.
Keyword:
Action recognition
Bag of video words
Hierarchical codebook
Spatio-temporal contextual information
Probabilistic modeling
Context
Ensemble of volumes
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.2
论文数:
4.1K
被引数:
6.7K

