Return
HUPSP-LAL: Efficiently mining utility-driven sequential patterns in uncertain sequences
DOI:10.1016/j.eswa.2025.126536.png)
Abstract
En 中文
Data mining encompasses various subfields, among which an important branch is high utility itemset mining. Within this domain, exploring high utility sequential patterns is an emerging field of interest, which is to identify high utility sequential patterns (HUSPs) within databases. In practice, there are many fields with application of high utility sequential pattern mining, including DNA sequence analysis and network intrusion detection, etc. However, most HUSPM assume that the data in the database is accurate, which is not consistent with the actual situation in the real world. Inevitably, data uncertainty arises due to the collection process, which involves sensors of varying degrees of precision. Although the methods of high utility probability sequential pattern mining (HUPSPM) in the context of uncertain sequences have been proposed, their performance is unsatisfactory when dealing with a low utility/probability threshold or largescale datasets. Therefore, we propose an efficient HUPSPM algorithm called HUPSP-LAL. We have proposed a new probability calculation framework to mathematically represent the collected uncertain data. We designed the compact structure, PUL - IA - EL , which HUPSP-LAL uses for projection to accelerate the calculation of the utility, probability, and upper bounds of the candidates. This paper introduces two probability-based pruning strategies, complemented by two additional utility-based pruning strategies, all aimed at diminishing the search space. The experimental findings from real datasets indicate that HUPSP-LAL outperforms the leading algorithms significantly regarding patterns, runtime, candidates, and memory consumption.
Keywords:
Data mining
Uncertain data
High utility sequential pattern
Database projection
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W
Organization
No organization information available

