arrow
Return

Controlling the False Split Rate in Tree-Based Aggregation

delete2024-09-24
delete0
delete
OA
AI
S
Simeng Shao
J
Jacob Bien *
A
Adel Javanmard
DOI:10.1080/01621459.2024.2376285delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In many domains, data measurements can naturally be associated with the leaves of a tree, expressing the relationships among these measurements. For example, companies belong to industries, which in turn belong to ever coarser divisions such as sectors; microbes are commonly arranged in a taxonomic hierarchy from species to kingdoms; street blocks belong to neighborhoods, which in turn belong to larger-scale regions. The problem of tree-based aggregation that we consider in this article asks which of these tree-defined subgroups of leaves should really be treated as a single entity and which of these entities should be distinguished from each other. We introduce the false split rate, an error measure that describes the degree to which subgroups have been split when they should not have been. While expressible as the false discovery rate in a special case, we show that these measures can be quite different for the general tree structures common in our setting. We then propose a multiple hypothesis testing algorithm for tree-based aggregation, which we prove controls this error measure. We focus on two main examples of tree-based aggregation, one which involves aggregating means and the other hich involves aggregating regression coefficients. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Keywords:
False discovery rate
Hierarchy
Multiple testing
Rare features

Journal

J
Journal of the American Statistical Association
IF:
3
Papers:
5.1K
Citations:
4.8W

Organization

U
university of southern california
Scholars:
4.6W
Papers: 3.8W
Citations: 51
A
amazon.com
Scholars:
698
Papers: 505
Citations: 8