返回
Sampling scheme-based classification rule mining method using decision tree in big data environment
DOI:10.1016/j.knosys.2022.108522.png)
摘要
En 中文
Obtaining comprehensible classification rules may be extremely important in many real applications such as data-driven decision-making and classification tasks. Decision-tree methods are powerful and popular tools for acquiring classification rules. However, they do not show good performance, and the base data processing methods lack strong theoretical support in big data scenarios. This study introduces a sampling scheme with and without the replacement of the implementations of decision tree methods. This method, called sampling-based classification rule mining (SCRM), is designed to improve the adaptation and generalization ability of classification rules in a big-data environment. Sampling without replacement is conducted to refine classification rules using the concept of conflict and coverage rules, while sampling with replacement is applied to determine rule reliability; the reliability approximation property of classification rules is proved by using the law of large numbers. The effectiveness of the SCRM was evaluated and verified using seven UCI datasets. Theoretical analysis and experimental results show that SCRM is generic with good classification ability, thereby improving the classification accuracy of the rules. SCRM has a significant advantage as it provides theoretical and methodological support for the classification rule mining of big data. Therefore, the SCRM can be used in many applications. (c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Classification rules
Decision tree
Sampling
Reliability
Big data
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
暂无机构信息
引用论文
Mining unexpected patterns using decision trees and interestingness measures: a case study of endometriosis
SOFT COMPUTING
IF2.5
A Cross-Domain Recommender System With Kernel-Induced Knowledge Transfer for Overlapping Entities面向重叠实体的基于内核诱导知识转移的跨域推荐系统

