arrow
Return

Learning to detect hardcoded secrets

delete2026-08-13
delete0
delete
OA
AI
F
Farnaz Soltaniani *
M
Mohammad Ghafari
DOI:10.1007/s10664-026-10929-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Numerous tools have been developed to uncover secrets and credentials in source code. However, they often rely on predefined rules that yield a high rate of false positives. We conduct an in-depth investigation into the performance of different AI models in secret detection. In non-generative models, we found that the choice of feature set greatly influences the performance in secret detection. These models demonstrated the best performance using the context feature, i.e., code surrounding the secret value. Particularly, CodeBERT outperforms other models with an MCC and F1-scores of 88% and 89%, respectively. When we applied the models to an unseen dataset of secrets, the top models were Random Forest and CodeBERT, achieving an MCC score of around 65%. In generative models, we observed moderate performance (a maximum MCC of 63%) and a high number of false positives. The investigation of incorrect predictions revealed that many are within test-path files and that LLMs tend to be cautious, flagging potential secrets to promote secure coding practices. We noted that prompt engineering reduces false predictions significantly, yet human intervention remains necessary.
Keywords:
Hardcoded secrets
Credential leakage
AI for security

Journal

Empirical Software Engineering cover
Empirical Software Engineering
IF:
3.6
Papers:
2.0K
Citations:
5.3K

Organization

T
technische universität clausthal
Scholars:
38
Papers: 16
Citations: 0
T
tehran institute for advanced studies
Scholars:
4
Papers: 4
Citations: 0
Cited Papers

Cited Papers

Why secret detection tools are not enough: It's not just about false positives-An industrial case study
err2022-03-17
err7
errOAAI
errRahman, Md Rayhanur; Imtiaz, Nasif; Storey, Margaret-Anne; Williams, Laurie
errShare
errSave
Long Short-Term Memory
err1997-11-01
err0
PREAI
errSepp Hochreiter; Jürgen Schmidhuber
errShare
errSave
A survey on ensemble learning
err2019-08-30
err1.1K
PREAI
errDong, Xibin; Yu, Zhiwen; Cao, Wenming; Shi, Yifan; Ma, Qianli
errShare
errSave
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
err2023-01-01
err0
PREAI
errCao,Yuan; Griffiths,Tom; Narasimhan,Karthik; Shafran,Izhak; Yao,Shunyu; Yu,Dian; Zhao,Jeffrey
errShare
errSave
Finding bugs is easy
err2004-12-01
err0
PREAI
errDavid Hovemeyer; William Pugh
errShare
errSave
Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models
err2022-01-01
err0
PREAI
errBosma,Maarten; Chi,Ed; Ichter,Brian; Le,Quoc V; Schuurmans,Dale; Wang,Xuezhi; Wei,Jason; Xia,Fei; Zhou,Denny
errShare
errSave
On Inter-Dataset Code Duplication and Data Leakage in Large Language Models
err2025-01-01
err0
errOAAI
errLopez, Jose Antonio Hernandez; Chen, Boqi; Saad, Mootez; Sharma, Tushar; Varro, Daniel
errShare
errSave
A Few Billion Lines of Code Later Using Static Analysis to Find Bugs in the Real World
err2010-02-01
err437
PREAI
errBessey, Al; Block, Ken; Chelf, Ben; Chou, Andy; Fulton, Bryan; Hallem, Seth; Henri-Gros, Charles; Kamsky, Asya; McPeak, Scott; Engler, Dawson
errShare
errSave
no more