arrow
Return

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

delete2026-09-28
delete0
delete
OA
AI
C
Chaojie Wei
Y
Yangyang Geng *
Y
Yunfeng Wang
Q
Qilong Wu
J
Jing Huang
Q
Qianqiong Wu
Q
Qiang Wei *
DOI:10.1186/s42400-026-00649-5delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Automatic Exploit Generation (AEG) plays an important role in proactive assessment of software threats by identifying vulnerabilities and constructing functional payloads. Existing Large Language Model (LLM)-based methods, however, often struggle to reason about complex exploit logic and to perform runtime introspection, leaving a gap between static vulnerability analysis and dynamic memory behavior. We present PwnAgent, an LLM-driven multi-agent framework for end-to-end exploit generation that combines offensive domain knowledge with active runtime introspection. PwnAgent uses a hierarchical knowledge base for multi-stage exploit reasoning and a feedback-driven self-correction engine to calibrate dynamic memory parameters during execution. Because broad Capture The Flag (CTF) benchmarks offer limited binary-exploitation depth and pwn-specific evaluation must balance reproducibility, difficulty progression, and exploit diversity, we construct a 66-task pwn benchmark from public CTF-style challenges. The benchmark is primarily composed of Linux x86/x86-64 ELF binaries and stack-oriented tasks, with smaller format-string, heap, integer-overflow, ARM, and MIPS subsets used as limited probes beyond the dominant setting. Under the same recent Kimi-K2.6 backend, PwnAgent achieves a 62.12% end-to-end success rate, compared with 31.82% for the evaluated PwnGPT baseline, a 30.30 percentage-point gain. These paired results indicate that structured knowledge guidance, execution-grounded measurement, and feedback repair improve LLM-based exploit generation in the evaluated setting, while the absolute success rate shows that fully autonomous exploitation remains challenging.
Keywords:
Automatic exploit generation
Multi-agent systems
Large language models
Knowledge-guided reasoning
Dynamic introspection
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Cybersecurity
IF:
3.7
Papers:
589
Citations:
1.0K

Organization

No organization information available
Cited Papers

Cited Papers

A survey on large language model based autonomous agents
err2024-03-22
err145
errOAAI
errWang, Lei; Ma, Chen; Feng, Xueyang; Zhang, Zeyu; Yang, Hao; Zhang, Jingsen; Chen, Zhiyuan; Tang, Jiakai; Chen, Xu; Lin, Yankai; Zhao, Wayne Xin; Wei, Zhewei; Wen, Jirong
errShare
errSave
ExploitGen: Template-augmented exploit code generation based on CodeBERT
err2023-03-01
err31
PREAI
errYang, Guang; Zhou, Yu; Chen, Xiang; Zhang, Xiangyu; Han, Tingting; Chen, Taolue
errShare
errSave