arrow
Return

Evaluating API-Level Deep Learning Fuzzers: A Comprehensive Benchmarking Study

delete2026-02-01
delete1
PRE
AI
N
Nima Shiri Harzevili *
M
Moshi Wei
M
Mohammad Mahdi Mohajer
S
Song Wang
H
Hung Viet Pham
DOI:10.1145/3729533delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, the practice of fuzzing Deep Learning (DL) APIs has received significant attention in the software engineering community. Many API-level DL fuzzers have been proposed to test individual DL APIs by generating malformed input. Although these fuzzers have been effective in detecting bugs and outperforming prior work, there remains a gap in benchmarking them against ground-truth, real-world bugs in DL libraries. Existing comparisons among these API-level DL fuzzers primarily focus on the bugs detected but do not offer a comprehensive, in-depth evaluation of the fuzzers' effectiveness. In this work, we perform the first in-depth evaluation of state-of-the-art API-level DL fuzzers that generate tests for single DL APIs, focusing on their effectiveness against real-world bugs. We manually created an extensive benchmark dataset, including 517 real-world DL bugs collected from PyTorch and TensorFlow libraries that can be triggered by malformed inputs. We then apply seven state-of-the-art DL fuzzers-FreeFuzz, DeepRel, NablaFuzz, DocTer, ACETest, TitanFuzz, and FuzzGPT-to our benchmark dataset, following their respective instructions. Our results show that these fuzzers detect only 6.5% (34 out of 517) of the unique real-world bugs in the dataset. Our analysis identifies two dominant factors that impact the effectiveness of these fuzzers in detecting real-world bugs. These findings suggest opportunities for improving the performance of fuzzers in future work. Overall, this study extends previous work on DL fuzzers by providing an extensive evaluation and benchmarking platform for fuzzing DL libraries.
Keywords:
fuzzing
deep learning
benchmarking

Journal

A
ACM Transactions on Software Engineering and Methodology
IF:
6.2
Papers:
1.2K
Citations:
3.4K

Organization

Y
york university - canada
Scholars:
8.3K
Papers: 9.0K
Citations: 10
Cited Papers

Cited Papers

Igor: Crash Deduplication Through Root-Cause Clustering
err2021-11-12
err0
PREAI
errJiang,Zhiyuan; Jiang,Xiyue; Hazimeh,Ahmad; Tang,Chaojing; Zhang,Chao; Payer,Mathias
errShare
errSave
Magma
err2020-11-30
err0
errOAAI
errAhmad Hazimeh; Adrian Herrera; Mathias Payer
errShare
errSave
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
researcher View more