arrow
返回

Distribution-based aggregation for relational learning with identifier attributes

delete2006-01-27
delete74
delete
OA
AI
C
Claudia Perlich
F
Foster Provost
DOI:10.1007/s10994-006-6064-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Identifier attributes-very high-dimensional categorical attributes such as particular product ids or people's names-rarely are incorporated in statistical modeling. However, they can play an important role in relational modeling: it may be informative to have communicated with a particular set of people or to have purchased a particular set of products. A key limitation of existing relational modeling techniques is how they aggregate bags (multisets) of values from related entities. The aggregations used by existing methods are simple summaries of the distributions of features of related entities: e.g., MEAN, MODE, SUM, or COUNT. This paper's main contribution is the introduction of aggregation operators that capture more information about the value distributions, by storing meta-data about value distributions and referencing this meta-data when aggregating-for example by computing class-conditional distributional distances. Such aggregations are particularly important for aggregating values from high-dimensional categorical attributes, for which the simple aggregates provide little information. In the first half of the paper we provide general guidelines for designing aggregation operators, introduce the new aggregators in the context of the relational learning system ACORA (Automated Construction of Relational Attributes), and provide theoretical justification. We also conjecture special properties of identifier attributes, e.g., they proxy for unobserved attributes and for information deeper in the relationship network. In the second half of the paper we provide extensive empirical evidence that the distribution-based aggregators indeed do facilitate modeling with high-dimensional categorical attributes, and in support of the aforementioned conjectures.
Keyword:
identifiers
relational learning
aggregation
networks

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

暂无机构信息
引用论文

引用论文

Time-series genome-centric analysis unveils bacterial response to operational disturbance in activated sludge
err
IF0
err2019-03-06
err0
errOAAI
errMaría Victoria Pérez; Leandro D. Guerrero; Esteban Orellana; Eva L. Figuerola; Leonardo Erijman
err分享
err收藏
On the evolution of noise-dependent vocal plasticity in birds
err2012-09-12
err0
errOAAI
errSophie Schuster; Sue Anne Zollinger; John A. Lesku; Henrik Brumm
err分享
err收藏
err分享
err收藏
The Poynting–Robertson effect: A critical perspective
err2014-04-01
err0
PREAI
errJ. Klačka; J. Petržala; P. Pástor; L. Kómar
err分享
err收藏
err分享
err收藏
学者 查看更多内容