arrow
Return

BiasAgent: Exploiting Agent Bias for Preference Manipulation Attacks on Model Context Protocol

delete2026-07-17
delete0
PRE
AI
J
Ji Guo
Z
Zhijing Wang
W
Wenbo Jiang
张瑞颖 (Rui Ying Zhang)
J
Jian Xiong
Q
Qiyang Song
H
H B Wu
Y
Yi-Jing Liu
DOI:10.1109/tccn.2026.3714063delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Model Context Protocol (MCP) has recently emerged as a unified framework for connecting Large Language Model (LLM) agents with external tools and applications, enabling more flexible interactions with their environments. However, recent studies have shown that MCP is also vulnerable to preference manipulation attacks, where an attacker deploys a customized MCP server to manipulate an LLM agent, causing it to favor that server over competing alternatives. Prior approaches rely on injecting manipulative keywords or phrases into tool names or descriptions, making them easily detectable. In this paper, we explore leveraging agent bias to directly manipulate tool selection without modifying tool descriptions or user inputs. Specifically, we construct a biased dataset that favors a designated tool and train a biased agent using Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO). The resulting agent selectively prefers the attacker-specified tool for relevant tasks while maintaining normal functionality on other tasks. Experimental results demonstrate that our approach achieves high effectiveness and strong stealthiness across multiple LLM agents and task settings.
Keywords:
Agent security
MCP security
preference manipulation attacks

Journal

I
IEEE Transactions on Cognitive Communications and Networking
IF:
7
Papers:
1.5K
Citations:
5.5K

Organization

U
university of electronic science and technology of china
Scholars:
1.3W
Papers: 4.7K
Citations: 4
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704