返回
Statistically Significant Pattern Mining With Ordinal Utility
DOI:10.1109/TKDE.2022.3208626.png)
摘要
En 中文
Statistically significant pattern mining (SSPM), which evaluates each pattern via a hypothesis test, is an essential and challenging data mining task for knowledge discovery. We introduce a preference relation between patterns and aim to discover the most preferred patterns under the constraint of statistical significance, which has never been considered in existing SSPM problems. We propose an iterative multiple testing procedure that can alternately reject a hypothesis and safely ignore the less useful hypotheses than the rejected one. By filtering out patterns with low utility, we can avoid the significance budget consumption of rejecting useless (uninteresting) patterns and focus the significance budget on more useful patterns, leading to more useful discoveries. We show that the proposed method can control the familywise error rate (FWER) under certain assumptions, which can be satisfied by a realistic problem class in SSPM. We also show that the proposed method always discovers equally or more useful patterns than Tarone-Bonferroni and Subfamily-wise Multiple Testing (SMT). Finally, we conducted several experiments with both synthetic and real-world data to evaluate the performance of our method. The proposed method discovered many more useful patterns in the experiments with real-world datasets than the existing method for all five conducted tasks.
Keyword:
High-utility pattern
multiple testing
significant pattern mining
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Dynamic evolution behavior of cracks for single-track and multi-track clads in laser cladding激光熔覆单轨和多轨熔覆层裂纹的动态演化行为
Utility-based association rule mining: A marketing solution for cross-selling基于效用的关联规则挖掘: 交叉销售的营销解决方案

