arrow
Return

Hashing for Fast Pattern Set Selection

delete2026-01-01
delete0
PRE
AI
M
Maiju Karjalainen *
P
Pauli Miettinen
DOI:10.1007/978-3-032-06096-9_8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Pattern set mining, which is the task of finding a good set of patterns instead of all patterns, is a fundamental problem in data mining. Many different definitions of what constitutes a good set have been proposed in recent years. In this paper, we consider the reconstruction error as a proxy measure for the goodness of the set, and concentrate on the adjacent problem of how to find a good set efficiently. We propose a method based on bottom-k hashing for efficiently selecting the set and extend the method for the common case where the patterns might only appear in approximate form in the data. Our approach has applications in tiling databases, Boolean matrix factorization, and redescription mining, among others. We show that our hashing-based approach is significantly faster than the standard greedy algorithm while obtaining almost equally good results in both synthetic and real-world data sets.
Keywords:
Hashing
Pattern Set Mining
BMF
Redescription Mining

Journal

M
MACHINE LEARNING AND KNOWLEDGE DISCOVERY IN DATABASES. RESEARCH TRACK, ECML PKDD 2025, PT V
IF:
0
Papers:
26
Citations:
0

Organization

U
University of Eastern Finland
Scholars:
1.4W
Papers: 1.2W
Citations: 1.5W