Skip to main navigation Skip to search Skip to main content

HEPE++: Towards VLM-based Hybrid Explainable Prompt Enhancement for Zero-Shot Industrial Anomaly Detection

  • Weichao Cai
  • , Weiliang Huang
  • , Yunkang Cao
  • , Chao Huang
  • , Bob Zhang
  • , Fei Yuan*
  • , Jie Wen
  • , Weiming Shen
  • *Corresponding author for this work
  • Xiamen University
  • University of Macau
  • Hunan University
  • Sun Yat-Sen University
  • Harbin Institute of Technology
  • Fuyao University of Science & Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Zero-Shot Industrial Anomaly Detection (ZSIAD) aims to identify and localize anomalies in industrial images from unseen categories. Owing to the powerful generalization capabilities, Vision-Language Models (VLMs) have achieved growing interest in ZSIAD. To guide the model toward understanding and localizing the semantically complex industrial anomalies, existing VLM-based methods have attempted to provide additional prompts to the model through learnable text prompt templates. However, these zero-shot methods lack detailed descriptions of specific anomalies, making it difficult to classify and segment the diverse range of industrial anomalies accurately. To address the aforementioned issue, we propose an effective solution, which includes a multi-stage prompt generation agent and a Hybrid Explainable Prompt Enhancement plus (HEPE++) framework. Specifically, we leverage the Multi-modal Language Large Model (MLLM) to articulate the detailed differential information between normal and test samples, which can provide detailed text prompts to the model through further refinement and anti-false alarm constraint. Moreover, we introduce the Visual Fundamental Model (VFM) to generate anomaly-related attention prompts for more accurate localization of anomalies with varying sizes and shapes. Building on this foundation, HEPE++ jointly utilizes generated detailed text and attention prompts to enhance the performance of VLM-based methods in ZSIAD. Furthermore, we introduce a novel gated attention mechanism to augment the model's ability to focus on potential subtle anomalous regions in the foreground. Extensive experiments on seven real-world industrial anomaly detection datasets have shown that the proposed method not only outperforms recent SOTA methods, but also its explainable prompts provide the model with a more intuitive basis for anomaly identification.

Original languageEnglish
JournalIEEE Transactions on Circuits and Systems for Video Technology
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • Industrial Anomaly Detection
  • Multimodal Learning
  • Zero-Shot

Fingerprint

Dive into the research topics of 'HEPE++: Towards VLM-based Hybrid Explainable Prompt Enhancement for Zero-Shot Industrial Anomaly Detection'. Together they form a unique fingerprint.

Cite this