arrow
Return

Evaluating Large Language Models for Line-Level Vulnerability Localization

delete2026-03-01
delete0
PRE
AI
J
Jian Zhang
C
Chong Wang
A
Anran Li *
W
Weisong Sun
C
C Zhang
W
Wei Ma
Y
Yang Liu
DOI:10.1109/TSE.2025.3649250delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recently, Automated Vulnerability Localization (AVL) has attracted growing attention, aiming to facilitate diagnosis by pinpointing the specific lines of code responsible for vulnerabilities. Large Language Models (LLMs) have shown potential in various domains, yet their effectiveness in line-level vulnerability localization remains underexplored. In this work, we present the first comprehensive empirical evaluation of LLMs for AVL. Our study examines 19 leading LLMs suitable for code analysis, including ChatGPT and multiple open-source models, spanning encoder-only, encoder-decoder, and decoder-only architectures, with model sizes from 60M to 70B parameters. We evaluate three paradigms including few-shot prompting, discriminative fine-tuning, and generative fine-tuning with and without Low-Rank Adaptation (LoRA), on both a BigVul-derived dataset for C/C++ and a smart contract vulnerability dataset. Our results show that discriminative fine-tuning achieves substantial performance gains over existing learning-based AVL methods when sufficient training data is available. In low-data settings, prompting advanced LLMs such as ChatGPT proves more effective. We also identify challenges related to input length and unidirectional context during fine-tuning, and propose two remedial strategies: a sliding window approach and right-forward embedding, both of which yield significant improvements. Moreover, we provide the first assessment of LLM generalizability in AVL, showing that certain models can transfer effectively across Common Weakness Enumerations (CWEs) and projects. However, performance degrades notably for newly discovered vulnerabilities containing unfamiliar lexical or structural patterns, underscoring the need for continual adaptation. These findings offer practical guidance for deploying LLM-based AVL systems in realistic software security workflows.
Keywords:
Codes
Location awareness
Software
Large language models
Adaptation models
Training
Training data
Security
Robustness
Maintenance engineering
Vulnerability localization
large language models
deep learning
software security

Journal

IEEE Transactions on Software Engineering cover
IEEE Transactions on Software Engineering
IF:
5.6
Papers:
2.8K
Citations:
1.1W

Organization

B
Beihang University
Scholars:
5.1W
Papers: 4.1W
Citations: 37
Y
yale university
Scholars:
7.1K
Papers: 3.1K
Citations: 2
N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.7W
Citations: 8.1W
S
singapore management university
Scholars:
318
Papers: 241
Citations: 0
researcher View more organizations