Return
Mask-guided explicit feature modulation for multispectral pedestrian detection
DOI:10.1016/j.compeleceng.2022.108385.png)
Abstract
En 中文
Multi-task learning of object detection and box-level segmentation is commonly formulated as an implicit feature modulation method, which suffers lacking task interaction. In this paper, a novel explicit feature modulation solution with two different mask infusion methods is proposed. The first modulation method is semantic feature enhancement of backbone, which is achieved by a novel mask-guided mutual attention module (MMA). The proposed MMA module can explicitly guide the feature maps towards a more semantic informative direction for focalizing centrality of pedestrian, which can significantly improve the performance. The second modulation method is confidence score enhancement of detection head, which is benefited from our proposed mask-guided score fusion module (MSF). The proposed MSF module collects information from the classification, IOU, centerness feature map and the learned mask, which can discriminate false and true positives more effectively. It is qualitatively validated that the modulated feature maps in both backbone and detection head become more semantically meaningful and robust to scale and occlusion. Our method achieves a considerable gain over the state-of-the-arts on the KAIST, CVC14 and FLIR datasets. Besides, it runs at 22 FPS in default setting, making it favorable in many practical scenarios.
Keywords:
Multispectral pedestrian detection
Feature modulation
Mutual attention
Score fusion
Journal
C
IF:
4.9
Papers:
6.7K
Citations:
1.3W

