arrow
Return

DATA HARMONIZATION VIA REGULARIZED NONPARAMETRIC MIXING DISTRIBUTION ESTIMATION

delete2026-03-01
delete0
PRE
AI
S
Steven Wilkins‐Reeves *
Y
Yen‐Chi Chen
C
Chan, Kwun Chuen Gary
DOI:10.1214/25-AOAS2024delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data harmonization is the process of developing an equivalence between two measurements of a common domain. Our problem is motivated by dementia research in which multiple neuropsychological tests have been used in practice to measure the same underlying cognitive ability, such as memory or attention. We connect this statistical problem to mixing distribution estimation common in empirical Bayes approaches. We introduce and study a nonparametric latent trait model, develop a method that enforces the uniqueness of the regularized maximum likelihood estimator, show how a nonparametric EM algorithm will converge weakly to its maximizer, and illustrate its superior computational efficiency to off-the-shelf solvers. Furthermore, we develop methods for model selection and assessing the goodness-of-fit for the measurement model, an area neglected in most mixing distribution estimation problems. We develop methods for score conversion with uncertainty quantification in order to draw inferences on a whole population with multiple score scales. We apply our method to the National Alzheimer's Coordination Center Uniform Dataset and show that we can use our method to convert between score measurements and account for the measurement error. We show that this method outperforms standard techniques commonly used in dementia research.
Keywords:
Alzheimer's disease
nonparametric
expectation-maximization algorithm
latent trait model
measurement error model

Journal

A
Annals of Applied Statistics
IF:
1.4
Papers:
88
Citations:
5.1K

Organization

U
university of washington seattle
Scholars:
1.8K
Papers: 994
Citations: 0
U
university of washington
Scholars:
8.8K
Papers: 4.1K
Citations: 2