Return
Human action recognition based on multi-layer Fisher vector encoding method
DOI:10.1016/j.patrec.2015.06.029.png)
Abstract
En 中文
In this paper, we propose a new multi layer Fisher vector encoding method based on trajectory descriptors for human action recognition. The proposed method aims at improving the classical shallow Fisher vector (FV) encoding method. Our main contribution resides in considering a progressive representation of the geometric relationships among trajectories. In fact, our presentation is based on three nested layers and provides deep and discriminant structures by local spatial pooling and refining the representation from one layer to the next. To preserve more information in feature encoding process, fine and large spatio-ternporal structures have been applied. Fine structures aim at exploiting the local spatio-temporal information by building graphs of trajectories, while large structures aim at exploiting the global spatio-temporal information by spatio-temporal video subdivision. Our approach is evaluated on three popular and large human action datasets: Hollywood2, Olympic sports and HMDB51. Experiments show that more layers produce higher action classification accuracy, which proves the capability of our multi-layer Fisher vector encoding method. (C) 2015 Elsevier B.V. All rights reserved.
Keywords:
Human action recognition
Geometric relationships
Multi-layer Fisher encoding
Local pooling
Global pooling
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
3.3
Papers:
8.0K
Citations:
1.6W
Organization
Cited Papers
Extending Laplacian sparse coding by the incorporation of the image spatial context
NEUROCOMPUTING
IF6.5
no more

