arrow
Return

Multi-level visual-language models feature learning for generalizable anomaly detection

delete2026-08-17
delete0
delete
OA
AI
J
jianfeng Qiu
J
Junfa Li
谢娟 (Juan Xie)
X
Xueliang Ma
K
Ke Xu *
DOI:10.1007/s40747-026-02462-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in target datasets without accessing their samples. Although CLIP and other large-scale vision language models show strong generalization, their potential for multi-level feature extraction in ZSAD remains underexplored. To address this, we propose a Multi-Level Feature Learning (MLFL) framework to enhance the zero-shot capability of CLIP via hierarchical alignment. MLFL adopts a two-stage training paradigm: Multi-Level Text Prompt Tuning (MLTP) and Multi-Level Text-Image Feature Alignment (MLFA). MLTP learns object-agnostic and object-aware prompts tailored to different encoder blocks. MLFA aligns textual and visual features using linear layers for shallow blocks and a Deep Feature Alignment (DFA) module for deep blocks. To compress parameters and preserve semantics, we introduce a Generalized Prompt Distillation (GPD) module that distills object-aware prompts into a unified representation. Experiments on seven industrial datasets achieve state-of-the-art performance, and deployment tests on edge devices demonstrate the potential applicability of the framework in practical industrial scenarios.
Keywords:
Anomaly detection
Visual-language models
CLIP
Prompt tuning
Zero-shot

Journal

C
Complex & Intelligent Systems
IF:
4.6
Papers:
240
Citations:
0

Organization

S
School of Mathematics and Physics
Scholars:
242
Papers: 126
Citations: 0
S
School of Artificial Intelligence
Scholars:
660
Papers: 304
Citations: 0