arrow
Return

CLIP Driven Few-Shot Panoptic Segmentation

delete2023-01-01
delete0
delete
OA
AI
P
Pengfei Xian *
L
Lai-Man Po
Y
Yuzhi Zhao
W
Wing-Yin Yu
K
Kwok-Wai Cheung
DOI:10.1109/ACCESS.2023.3290070delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper presents CLIP Driven Few-shot Panoptic Segmentation (CLIP-FPS), a novel few-shot panoptic segmentation model that leverages the knowledge of Contrastive Language-Image Pre-training (CLIP) model. The proposed method builds upon a center indexing attention mechanism to facilitate knowledge transfer, which entails representing objects in an image as centers along with their pixel offsets. The model comprises a decoder responsible for generating object center-offset groups and a self-attention module tasked with producing a feature attention map. Subsequently, the object centers index the map to acquire the corresponding embeddings, paving the way for matrix multiplication and SoftMax operation to facilitate text embedding matching and the computation of the final panoptic segmentation masks. Quantitative evaluation on datasets such as COCO and Cityscapes shows that our method outperforms existing panoptic segmentation techniques in terms of Panoptic Quality (PQ) metrics.
Keywords:
Panoptic segmentation
CLIP
cityscapes
convolutional neural network
image processing

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

H
Hang Seng University of Hong Kong
Scholars:
350
Papers: 504
Citations: 1
C
City University of Hong Kong
Scholars:
2.3W
Papers: 3.0W
Citations: 6.1W