Return
Information Retrieval-Based Zero-Shot Surface Defect Semantic Segmentation
DOI:10.1109/jsen.2026.3691093.png)
Abstract
En 中文
Surface defect segmentation is a critical task in industrial production. To address the limitations of deep learning-based segmentation methods that heavily rely on large-scale labeled training data, zero-shot segmentation approaches have been proposed. However, existing methods still face two major challenges. First, current zero-shot approaches typically rely on a single modality, either defect-free images or textual prompts, without fully exploiting multimodal information. Second, although those methods based on defect-free images generally achieve better results, obtaining such images is often difficult and may be infeasible in practical industrial environments. To overcome these challenges, we propose a zero-shot segmentation method, termed the information retrieval-based zero-shot segmentation network (IRNet). Specifically, we employ the contrastive language–image pretraining (CLIP) framework to integrate textual and image guidance, and combine fine-tuning with prompt tuning to align text and image representations in industrial contexts. Considering the difficulty of acquiring high-quality defect-free images, our approach adopts a human–machine hybrid intelligence strategy, where the machine conducts automatic search and scoring, and human experts provide supervision and make the final decisions. This design enhances the model’s capability to segment previously unseen defects. Extensive experiments demonstrate that IRNet achieves effective defect segmentation in both seen and unseen scenarios, and further highlight its strong generalization ability and practical applicability in real-world industrial inspection tasks. We will provide the source code publicly: https://github.com/c-haos-onfusion/IRNet
Keywords:
Human–machine hybrid intelligence
multimodal
prompt tuning
surface defect segmentation
zero-shot segmentation
Journal
IF:
4.5
Papers:
2.2W
Citations:
7.3W

