arrow
返回

Rethinking Global Context in Crowd Counting

delete2024-04-15
delete3
PRE
AI
G
Guolei Sun
Y
Yun Liu *
T
Thomas Probst
D
Danda Pani Paudel
N
Nikola Popović
L
Luc Van Gool
DOI:10.1007/s11633-023-1475-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper investigates the role of global context for crowd counting. Specifically, a pure transformer is used to extract features with global information from overlapping image patches. Inspired by classification, we add a context token to the input sequence, to facilitate information exchange with tokens corresponding to image patches throughout transformer layers. Due to the fact that transformers do not explicitly model the tried-and-true channel-wise interactions, we propose a token-attention module (TAM) to recalibrate encoded features through channel-wise attention informed by the context token. Beyond that, it is adopted to predict the total person count of the image through regression-token module (RTM). Extensive experiments on various datasets, including ShanghaiTech, UCF-QNRF, JHU-CROWD++ and NWPU, demonstrate that the proposed context extraction techniques can significantly improve the performance over the baselines.
Keyword:
Crowd counting
vision transformer
global context
attention
density map

期刊

Machine Intelligence Research 封面图
Machine Intelligence Research
IF:
8.7
论文数:
301
被引数:
882

机构

E
ETH Zurich
学者数:
3.0W
论文数: 2.4W
被引数: 8.4W
S
swiss federal institutes of technology domain
学者数:
9.0W
论文数: 8.0W
被引数: 163
A
agency for science technology & research (a*star)
学者数:
2.2W
论文数: 1.9W
被引数: 57
学者 查看更多机构