arrow
Return

Mixed spatial pyramid pooling for semantic segmentation

delete2020-06-01
delete10
PRE
AI
Z
Zhengyu Xia
J
Joohee Kim *
DOI:10.1016/j.asoc.2020.106209delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Semantic segmentation is a challenging task as each pixel should be labeled accurately in the image. To improve the performance of semantic segmentation, some Fully Convolutional Network (FCN) based semantic segmentation methods adopt a spatial pyramid pooling structure to enrich contextual information. Others employ an encoder-decoder architecture to recover object details gradually. In this paper, we propose a semantic segmentation framework which combines the benefits of these approaches. Specifically, we propose a Mixed Spatial Pyramid Pooling (MSPP) module based on region-based average pooling and dilated convolution to obtain dense multi-level contextual priors. To further refine the details of objects more effectively, we also propose a Global-Attention Fusion (GAF) module to provide global context as guidance for low-level features. Our proposed method achieves mIoU of 84.1% on PASCAL VOC 2012 dataset and 80.4% on Cityscapes dataset without using any post-processing or additional datasets for pretrained model. (C) 2020 Elsevier B.V. All rights reserved.
Keywords:
Semantic segmentation
Convolutional neural network
Spatial pyramid pooling
Dilated convolution
Encoder-decoder architecture
Scene understanding
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

I
Illinois Institute of Technology
Scholars:
3.8K
Papers: 3.9K
Citations: 4.2K