1
Return

Protecting Your Customized LLM Systems From Backdoored Instructions With Metacognitive Probing

delete2026-07-16
delete0
PRE
AI
S
Shuai Zhao
Z
Zhongliang Guo
X
Xinyi Wu
X
Xiaobao Wu
Y
Yanhao Jia
L
Luwei Xiao
A
Anh Tuan Luu
DOI:10.1109/tifs.2026.3714137delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Imagine a user who lacks any prior experience in artificial intelligence deployment and who, in order to enhance work productivity, is consequently required to customize large language model (LLM) systems provided by third-party institutions. While highly effective in practical applications, these customized LLMs introduce critical security risks. Specifically, customized LLMs are highly vulnerable to backdoored instructions, which are malicious rules stealthily embedded within the system instructions. Unlike traditional backdoor attacks, this threat executes attacks through malicious instructions, eliminating the need for LLM fine-tuning. Despite the existence of several algorithms designed to defend against backdoor attacks, these are not applicable to backdoored instructions. In this paper, we introduce the first Black-box Safety Auditing Agent targeting backdoored system instructions, which is motivated by the malicious tendencies and triggers observed in customized LLM systems. Specifically, we design an auditing agent that leverages metacognitive probing to induce LLMs to reveal predefined triggers. Subsequently, these triggers are sanitized from user queries to ensure that the backdoor remains inactive. The fundamental principle of metacognitive probing lies in leveraging the model’s safety-alignment mechanisms to countervail its instruction-following tendencies. This auditing agent considers two different scenarios, which involve prompt-based approaches for both task-specific probing and broad-spectrum probing, to comprehensively identify triggers. Furthermore, we also discuss the strategy of requiring the model to articulate its reasoning process, enhancing the defense capabilities of the auditing agent. We conduct experiments using three representative backdoor attack algorithms across multiple state-of-the-art LLMs. Our empirical results substantiate the effectiveness of the proposed auditing agent.
Keywords:
Large language models
metacognitive probing
backdoor attack
auditing agent
instructions

Journal

IEEE Transactions on Information Forensics and Security cover
IEEE Transactions on Information Forensics and Security
IF:
8
Papers:
5.2K
Citations:
2.3W

Organization

U
university of st andrews
Scholars:
9.2K
Papers: 1.0W
Citations: 15
S
shanghai jiao tong university
Scholars:
15.1W
Papers: 11.5W
Citations: 159
N
Nanyang Technological University
Scholars:
4.8W
Papers: 4.7W
Citations: 8.1W
N
National University of Singapore
Scholars:
7.4W
Papers: 6.4W
Citations: 11.4W
Cited Papers

Cited Papers

Citing Papers

Citing Papers