arrow
返回

Multi-level refinement enriched feature pyramid network for object detection

delete2021-11-01
delete16
PRE
AI
L
Lubna Aziz *
M
Md. Sah Bin Haji Salam FC
S
Sara Ayub
DOI:10.1016/j.imavis.2021.104287delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Class Imbalance and scales imbalance are common in object detection. A class imbalance occurs due to insufficient inequality between the number of instances with respect to different classes, while an imbalance in scale occurs when object have different scales and a different number of examples of different scales. In order to solve the problem of scale variance (scale imbalance) and class imbalance together, we propose a simple and effective feature enhancement scheme that explicitly uses all information of a multi-level structure to generate a multilevel contextual features pyramid with multiple scales. We also introduce a cascaded refinement scheme that incorporates multi-scale contextual features into the Single Shot Detector (SSD) predictive layers to improve their distinctiveness for multi-scale detection. A stack of multi-scale contextual feature modules is used in a feature enhancement scheme to merge the multi-level and multi-scale features. Then we collect the equivalent scale features over the Multi-layer Feature Fusion (MLFF) unit to construct a feature pyramid in which each feature map is made up of layers from multiple levels. More robustness and contextual information are integrated into the pyramid through chain parallel pooling operation. To improve classification and regression, a cascaded refinement scheme is proposed that effectively captures a large amount of contextual information and refines the anchors to solve the class imbalance problem. The experiments are carried out on two benchmarks datasets: MS COCO and PASCAL VOC 07/12. Our proposed approach achieves state-of-the-art accuracy with an AP of 40.6 in the case of multi-scale inference on MS COCO Test-dev (input size 320 x 320). For 512 x 512 input on the MS COCO Test-dev, our approach leads in an absolute gain in precision of 1.8% compared to the best reported results of single-stage detector (AP: 45.7). (c) 2021 Elsevier B.V. All rights reserved.
Keyword:
CNN
Object detection
Chained parallel pooling
Computer vision
Feature pyramid
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

U
Universiti Teknologi Malaysia
学者数:
1.4W
论文数: 1.1W
被引数: 85
引用论文

引用论文

err分享
err收藏
The FIGNL1-FIRRM complex is required to complete meiotic recombination in the mouse and prevents massive DNA damage-independent RAD51 and DMC1 loading
err
IF0
err2023-05-18
err0
errOAAI
errAkbar Zainu; Pauline Dupaigne; Soumya Bouchouika; Julien Cau; Julie A. J. Clément; Pauline Auffret; Virginie Ropars; Jean-Baptiste Charbonnier; Bernard de Massy; Raphael Mercier; Rajeev Kumar; Frédéric Baudat
err分享
err收藏
The Pascal Visual Object Classes (VOC) ChallengePascal视觉对象课程 (VOC) 挑战
err2009-09-09
err9.0K
PREAI
errEveringham, Mark; Van Gool, Luc; Williams, Christopher K. I.; Winn, John; Zisserman, Andrew
err分享
err收藏
err分享
err收藏
err分享
err收藏
没有更多内容