返回
Evaluating API-Level Deep Learning Fuzzers: A Comprehensive Benchmarking Study
DOI:10.1145/3729533.png)
摘要
En 中文
In recent years, the practice of fuzzing Deep Learning (DL) APIs has received significant attention in the software engineering community. Many API-level DL fuzzers have been proposed to test individual DL APIs by generating malformed input. Although these fuzzers have been effective in detecting bugs and outperforming prior work, there remains a gap in benchmarking them against ground-truth, real-world bugs in DL libraries. Existing comparisons among these API-level DL fuzzers primarily focus on the bugs detected but do not offer a comprehensive, in-depth evaluation of the fuzzers' effectiveness. In this work, we perform the first in-depth evaluation of state-of-the-art API-level DL fuzzers that generate tests for single DL APIs, focusing on their effectiveness against real-world bugs. We manually created an extensive benchmark dataset, including 517 real-world DL bugs collected from PyTorch and TensorFlow libraries that can be triggered by malformed inputs. We then apply seven state-of-the-art DL fuzzers-FreeFuzz, DeepRel, NablaFuzz, DocTer, ACETest, TitanFuzz, and FuzzGPT-to our benchmark dataset, following their respective instructions. Our results show that these fuzzers detect only 6.5% (34 out of 517) of the unique real-world bugs in the dataset. Our analysis identifies two dominant factors that impact the effectiveness of these fuzzers in detecting real-world bugs. These findings suggest opportunities for improving the performance of fuzzers in future work. Overall, this study extends previous work on DL fuzzers by providing an extensive evaluation and benchmarking platform for fuzzing DL libraries.
Keyword:
fuzzing
deep learning
benchmarking
期刊
A
IF:
6.2
论文数:
1.2K
被引数:
3.4K
机构
引用论文
Image classification with deep learning in the presence of noisy labels: A survey在存在噪声标签的情况下使用深度学习进行图像分类: 一项调查

