arrow
Return

HRNeXt: High-Resolution Context Network for Crowd Pose Estimation

delete2023-01-01
delete11
PRE
AI
Q
Qun Li
Z
Ziyi Zhang
F
Feng Zhang
F
Fu Xiao *
DOI:10.1109/TMM.2023.3248144delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Occlusion handling in crowded scenes is an intractable challenge for human pose estimation. To address this problem, we propose two novel feed-forward network structures named Global Feed-Forward Network (GFFN) and Dynamic Feed-Forward Network (DFFN), which are specifically designed for image-based tasks to capture both local and global contextual information within intermediate features and update feature representations with high adaptability for occlusions. By exploiting the context modeling ability of the proposed GFFN and DFFN, we present a novel backbone network, namely High-Resolution Context Network (HRNeXt), which learns high-resolution representations with abundant contextual information to better estimate poses of occluded human bodies. Compared to state-of-the-art pose estimation networks, our HRNeXt absorbs advantages of convolution operation and attention mechanism, and it is more efficient in terms of training data sizes, network parameters and computational costs. Experimental results show that our HRNeXt significantly outperforms state-of-the-art backbone networks on challenging pose estimation datasets with high occurrence of crowds and occlusions.
Keywords:
Pose estimation
Convolution
Task analysis
Kernel
Feature extraction
Context modeling
Transformers
Crowd pose estimation
context modeling
feed-forward network
attention mechanism
high-resolution representation

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

No organization information available