arrow
返回

Frequent high minimum average utility sequence mining with constraints in dynamic databases using efficient pruning strategies

delete2021-09-01
delete8
PRE
AI
T
Tin Truong
H
Hai Duong
B
Bac Le
P
Philippe Fournier‐Viger *
U
Unil Yun
DOI:10.1007/s10489-021-02520-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
High utility sequence mining is a popular data mining task, which aims at finding sequences having a high utility (importance) in a quantitative sequence database. Though it has several applications, state-of-the-art algorithms have one or more of the following limitations: (1) they rely on a utility function that tends to be biased toward finding long patterns, (2) some algorithms do take pattern length into account using an average-utility function but they adopt an optimistic perspective that can be risky or misleading for some applications, (3) they do not let the user specify additional constraints on patterns to be found. To address these three limitations, this paper defines a novel task of mining frequent high minimum average-utility sequences (FHAUS) with constraints in a quantitative sequence database. This task has the following benefits. First, it uses the average-utility au function based on the minimum utility, which takes the length of a pattern into account to calculate its utility. This helps finding short patterns missed by traditional algorithms and it is based on more safe pessimistic utility calculations. Second, the user can specify a set of monotonic and anti-monotonic constraints C on patterns to filter irrelevant patterns and improve the performance of the mining process. To efficiently find all FHAUSs with constraints, this paper first proposes some novel upper bounds (UBs) and weak upper bounds (WUBs) on the average-utility, which satisfy downward-closure (DC) properties or DC-like properties. Then, to effectively reduce the search space, the paper designs novel width pruning, depth pruning, reducing, and tightening strategies based on the proposed bounds. These proposed novel theoretical results are integrated into an algorithm named C-FHAUSPM (Constrained Frequent High minimum Average-Utility Sequential Pattern Mining) for efficiently discovering all FHAUSs with constraints. Results from extensive experiments on both real-life and synthetic quantitative sequence databases show that C-FHAUSPM is highly efficient in terms of runtime and memory usage.
Keyword:
Utility mining
High average-utility sequence
Upper bound
Weak upper bound
Pruning strategy

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
V
vietnam national university ho chi minh city (vnuhcm) system
学者数:
7.2K
论文数: 4.2K
被引数: 8
D
Dalat University
学者数:
149
论文数: 95
被引数: 278
学者 查看更多机构
引用论文

引用论文

High throughput analysis of MHC‐I and MHC‐DR diversity of Brazilian cattle populations
errHLA
IF0
err2021-06-17
err0
errOAAI
errDeepali Vasoya; Priscila Silva Oliveira; Laura Agundez Muriel; Thomas Tzelos; Christina Vrettou; W. Ivan Morrison; Isabel Kinney Ferreira de Miranda Santos; Timothy Connelley
err分享
err收藏
Mining cost-effective patterns in event logs
err2020-03-01
err31
PREAI
errFournier-Viger, Philippe; Li, Jiaxuan; Lin, Jerry Chun-Wei; Tin Truong Chi; Kiran, R. Uday
err分享
err收藏
An efficient algorithm to mine high average-utility itemsets
err2016-04-01
err71
PREAI
errLin, Jerry Chun-Wei; Li, Ting; Fournier-Viger, Philippe; Hong, Tzung-Pei; Zhan, Justin; Voznak, Miroslav
err分享
err收藏
err分享
err收藏
A MOBILE LESION IN THE CAROTID ARTERY
err2008-01-21
err0
PREAI
errJ. Stewart; J. Gover; D. Tridgell; And J. Frawley
err分享
err收藏
CCSpan: Mining closed contiguous sequential patterns
err2015-11-01
err49
errOAAI
errZhang, Jingsong; Wang, Yinglin; Yang, Dingyu
err分享
err收藏
Applying the maximum utility measure in high utility sequential pattern mining
err2014-09-01
err92
PREAI
errLan, Guo-Cheng; Hong, Tzung-Pei; Tseng, Vincent S.; Wang, Shyue-Liang
err分享
err收藏
err分享
err收藏
学者 查看更多内容