arrow
Return

An interactive network based on transformer for multimodal crowd counting

delete2023-06-30
delete5
PRE
AI
余鹰 cover
余鹰 (Ying Yu) *
C
Cai Zhen
苗
苗夺谦 (Duoqian Miao)
J
Jin Qian
唐洪 cover
唐洪 (Hong Tang)
DOI:10.1007/s10489-023-04721-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Crowd counting is a task to estimate the total number of pedestrians in an image. In most of the existing research, good vision problems, such as in parks, squares, and bright shopping malls during the day, have been addressed. However, there is little research on complex scenes in darkness. To study this problem, we propose an interactive network based on Transformer for multi-modal crowd counting. First, sliding convolutional encoding is adopted for the image to obtain better encoding features. The features are extracted through the designed primary interaction network, and then channel token attention is used to modulate the features. Then, the FGAF-MLP is used for high and low semantic fusion to enhance the feature expression and fully fuse the data in different modes to improve the accuracy of the method. To verify the effectiveness of our method, we conducted extensive ablation experiments with the latest multimodal benchmark RGBT-CC, and we verified the complementarity between multiple modal data and the effectiveness of the model components. We also verified the effectiveness of our method with the ShanghaiTechRGBD benchmark. The experimental results showed that our proposed method exhibits good results and achieves an improvement of more than 10% in terms of the mean average error and mean squared error for the RGBT-CC benchmark.
Keywords:
Crowd counting
Transformer
Multimodal data
Feature fusion

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.6K
Citations:
1.7W

Organization

T
tongji university
Scholars:
7.9W
Papers: 6.0W
Citations: 98
E
East China Jiaotong University
Scholars:
4.1K
Papers: 2.9K
Citations: 2.9K
Cited Papers

Cited Papers

errShare
errSave
errShare
errSave
errShare
errSave
Pixel-level image fusion: A survey of the state of the art
err2017-01-01
err889
PREAI
errLi, Shutao; Kang, Xudong; Fang, Leyuan; Hu, Jianwen; Yin, Haitao
errShare
errSave
errShare
errSave
A Pilot Study of the Efficacy of the Unified Protocol for Transdiagnostic Treatment of Emotional Disorders in Treating Posttraumatic Psychopathology: A Randomized Controlled Trial
err2021-01-16
err0
errOAAI
errMeaghan L. O'Donnell; Winnie Lau; Katherine Chisholm; James Agathos; Jonathon Little; Sonia Terhaag; Rachel Brand; Andrea Putica; Alexander C. N. Holmes; Lynda Katona; Kim L. Felmingham; Kim Murray; Fardous Hosseiny; Matthew W. Gallagher
errShare
errSave
Multi-source information fusion based on rough set theory: A review
err2021-04-01
err188
PREAI
errZhang, Pengfei; Li, Tianrui; Wang, Guoqiang; Luo, Chuan; Chen, Hongmei; Zhang, Junbo; Wang, Dexian; Yu, Zeng
errShare
errSave
Crystalline‐State Reaction with Allosteric Effect in Spin‐Crossover, Interpenetrated Networks with Magnetic and Optical Bistability
err2003-08-13
err0
PREAI
errVirginie Niel; Amber L. Thompson; M. Carmen Muñoz; Ana Galet; Andrés E. Goeta; José A. Real
errShare
errSave
researcher View more