Return
BiasAgent: Exploiting Agent Bias for Preference Manipulation Attacks on Model Context Protocol
DOI:10.1109/tccn.2026.3714063.png)
Abstract
En 中文
Model Context Protocol (MCP) has recently emerged as a unified framework for connecting Large Language Model (LLM) agents with external tools and applications, enabling more flexible interactions with their environments. However, recent studies have shown that MCP is also vulnerable to preference manipulation attacks, where an attacker deploys a customized MCP server to manipulate an LLM agent, causing it to favor that server over competing alternatives. Prior approaches rely on injecting manipulative keywords or phrases into tool names or descriptions, making them easily detectable. In this paper, we explore leveraging agent bias to directly manipulate tool selection without modifying tool descriptions or user inputs. Specifically, we construct a biased dataset that favors a designated tool and train a biased agent using Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO). The resulting agent selectively prefers the attacker-specified tool for relevant tasks while maintaining normal functionality on other tasks. Experimental results demonstrate that our approach achieves high effectiveness and strong stealthiness across multiple LLM agents and task settings.
Keywords:
Agent security
MCP security
preference manipulation attacks
Journal
I
IF:
7
Papers:
1.5K
Citations:
5.5K

