Return
Protecting Your Customized LLM Systems From Backdoored Instructions With Metacognitive Probing
S
Z
X
X
Y
L
A
DOI:10.1109/tifs.2026.3714137.png)
Abstract
En 中文
Imagine a user who lacks any prior experience in artificial intelligence deployment and who, in order to enhance work productivity, is consequently required to customize large language model (LLM) systems provided by third-party institutions. While highly effective in practical applications, these customized LLMs introduce critical security risks. Specifically, customized LLMs are highly vulnerable to backdoored instructions, which are malicious rules stealthily embedded within the system instructions. Unlike traditional backdoor attacks, this threat executes attacks through malicious instructions, eliminating the need for LLM fine-tuning. Despite the existence of several algorithms designed to defend against backdoor attacks, these are not applicable to backdoored instructions. In this paper, we introduce the first Black-box Safety Auditing Agent targeting backdoored system instructions, which is motivated by the malicious tendencies and triggers observed in customized LLM systems. Specifically, we design an auditing agent that leverages metacognitive probing to induce LLMs to reveal predefined triggers. Subsequently, these triggers are sanitized from user queries to ensure that the backdoor remains inactive. The fundamental principle of metacognitive probing lies in leveraging the model’s safety-alignment mechanisms to countervail its instruction-following tendencies. This auditing agent considers two different scenarios, which involve prompt-based approaches for both task-specific probing and broad-spectrum probing, to comprehensively identify triggers. Furthermore, we also discuss the strategy of requiring the model to articulate its reasoning process, enhancing the defense capabilities of the auditing agent. We conduct experiments using three representative backdoor attack algorithms across multiple state-of-the-art LLMs. Our empirical results substantiate the effectiveness of the proposed auditing agent.
Keywords:
Large language models
metacognitive probing
backdoor attack
auditing agent
instructions
Journal
IF:
8
Papers:
5.2K
Citations:
2.3W
