返回
Multi-Granularity Context Network for Efficient Video Semantic Segmentation
DOI:10.1109/TIP.2023.3269982.png)
摘要
En 中文
Current video semantic segmentation tasks involve two main challenges: how to take full advantage of multi-frame context information, and how to improve computational efficiency. To tackle the two challenges simultaneously, we present a novel Multi-Granularity Context Network (MGCNet) by aggregating context information at multiple granularities in a more effective and efficient way. Our method first converts image features into semantic prototypes, and then conducts a non-local operation to aggregate the per-frame and short-term contexts jointly. An additional long-term context module is introduced to capture the video-level semantic information during training. By aggregating both local and global semantic information, a strong feature representation is obtained. The proposed pixel-to-prototype non-local operation requires less computational cost than traditional non-local ones, and is video-friendly since it reuses the semantic prototypes of previous frames. Moreover, we propose an uncertainty-aware and structural knowledge distillation strategy to boost the performance of our method. Experiments on Cityscapes and CamVid datasets with multiple backbones demonstrate that the proposed MGCNet outperforms other state-of-the-art methods with high speed and low latency.
Keyword:
Semantics
Semantic segmentation
Prototypes
Aggregates
Feature extraction
Training
Task analysis
Video semantic segmentation
light-weight networks
non-local operation
期刊
IF:
13.7
论文数:
1.0W
被引数:
8.4W
机构
引用论文
In vitro and in vivo binding of neuroactive steroids to the sigma‐1 receptor as measured with the positron emission tomography radioligand [18F]FPS
Synapse
IF0
Recent Development of Dual-Dictionary Learning Approach in Medical Image Analysis and Reconstruction
CO2 to Terpenes: Autotrophic and Electroautotrophic α‐Humulene Production with Cupriavidus necatorCO2 到萜烯: 使用 Cupriavidus necator 的自养和电自养 α-腐殖质生产

