返回
Redescription mining on data with background network information
DOI:10.1016/j.knosys.2022.110109.png)
摘要
En 中文
Redescription mining aims at finding subsets of instances that can be re-described, characterized in multiple ways, using one or more disjoint sets of attributes that describe some set of instances. Current redescription mining algorithms either work with tabular data or with relational data - where binary relations between objects are used which allow representing descriptions as graphs. In this work, we propose novel type of redescription mining methodology that allows using tabular data in combination with background network information, where nodes of a network are instances in the tabular data. Background information is used to locate subsets of instances with some desired network property whereas tabular data are used to re-describe such interesting subsets. Methodology can be classified as constraint-based redescription mining, where we allow for a large variety of complex network-based soft constraints. The proposed framework is extensible, thus any network-related measure can be used to localize subsets of instances of interest. In addition, different types of network such as undirected, directed graphs, graph sequences or multiplex can be used as a background information. We demonstrate the applicability of the proposed framework on three use -case datasets involving country trade networks, biological (gene spatial) networks and social networks. The experimental evaluation demonstrates that the proposed approach outperforms existing, general redescription mining approaches with respect to intensity of network properties of the re-described instances without loss of accuracy, mostly even improving redescription accuracy.(c) 2022 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keyword:
Redescription mining
Network information
Trade network
Biological network
Social network
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
引用论文
Scalable Triangle Discovery Algorithm for Large Scale-Free Network with Limited Internal Memory有限内存的大规模无标度网络可扩展三角形发现算法
Patterns of diverse gene functions in genomic neighborhoods predict gene function and phenotype
SCIENTIFIC REPORTS
IF3.9

