arrow
返回

Test Input Prioritization for Graph Neural Networks

delete2024-06-01
delete0
delete
OA
AI
Y
Yinghua Li
X
Xueqi Dang *
W
Weiguo Pian
A
Andrew Habib
J
Jacques Klein
T
Tegawendé F. Bissyandé
DOI:10.1109/TSE.2024.3385538delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
GNNs have shown remarkable performance in a variety of classification tasks. The reliability of GNN models needs to be thoroughly validated before their deployment to ensure their accurate functioning. Therefore, effective testing is essential for identifying vulnerabilities in GNN models. However, given the complexity and size of graph-structured data, the cost of manual labelling of GNN test inputs can be prohibitively high for real-world use cases. Although several approaches have been proposed in the general domain of Deep Neural Network (DNN) testing to alleviate this labelling cost issue, these approaches are not suitable for GNNs because they do not account for the interdependence between GNN test inputs, which is crucial for GNN inference. In this paper, we propose NodeRank, a novel test prioritization approach specifically for GNNs, guided by ensemble learning-based mutation analysis. Inspired by traditional mutation testing, where specific operators are applied to mutate code statements to identify whether provided test cases reveal faults, NodeRank operates on a crucial premise: If a test input (node) can kill many mutated models and produce different prediction results with many mutated inputs, this input is considered more likely to be misclassified by the GNN model and should be prioritized higher. Through prioritization, these potentially misclassified inputs can be identified earlier with limited manual labeling cost. NodeRank introduces mutation operators suitable for GNNs, focusing on three key aspects: the graph structure, the features of the graph nodes, and the GNN model itself. NodeRank generates mutants and compares their predictions against that of the initial test inputs. Based on the comparison results, a mutation feature vector is generated for each test input and used as the input to ranking models for test prioritization. Leveraging ensemble learning techniques, NodeRank combines the prediction results of the base ranking models and produces a misclassification score for each test input, which can indicate the likelihood of this input being misclassified. NodeRank sorts all the test inputs based on their scores in descending order. To evaluate NodeRank, we build 124 GNN subjects (i.e., a pair of dataset and GNN model), incorporating both natural and adversarial contexts. Our results demonstrate that NodeRank outperforms all the compared test prioritization approaches in terms of both APFD and PFD, which are widely-adopted metrics in this field. Specifically, NodeRank achieves an average improvement of between 4.41% and 58.11% on original datasets and between 4.96% and 62.15% on adversarial datasets.
Keyword:
Test input prioritization
graph neural networks
mutation analysis
learning to rank
labelling

期刊

IEEE Transactions on Software Engineering 封面图
IEEE Transactions on Software Engineering
IF:
5.6
论文数:
2.8K
被引数:
1.1W

机构

U
university of luxembourg
学者数:
5.2K
论文数: 4.8K
被引数: 4
引用论文

引用论文

Geopolitics and discourse地缘政治与话语
err1992-03-01
err0
PREAI
errGearóid Ó Tuathail; John Agnew
err分享
err收藏
Test Input Prioritization for 3D Point Clouds3D点云的测试输入优先级
err2024-06-04
err0
errOAAI
errLi, Yinghua; Dang, Xueqi; Ma, Lei; Klein, Jacques; Le Traon, Yves; Bissyande, Tegawende F.
err分享
err收藏
Women Drug Users
err
IF0
err2023-10-31
err0
PREAI
errAvril Taylor
err分享
err收藏
学者 查看更多内容