arrow
Return

Improved distance correlation estimation

delete2025-01-03
delete0
delete
OA
AI
B
Blanca E. Monroy-Castillo *
M
María Amalia Jácome
R
Ricardo Cao
DOI:10.1007/s10489-024-05940-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Distance correlation is a novel class of multivariate dependence measure, taking positive values between 0 and 1, and applicable to random vectors of not necessarily equal arbitrary dimensions. It offers several advantages over the well-known Pearson correlation coefficient, the most important being that distance correlation equals zero if-and-only if- the random vectors are independent. There are two different estimators of the distance correlation available in the literature. The first estimator, proposed by Sz & eacute;kely et al. (Ann Stat 35:2769-279 2007), is based on an asymptotically unbiased estimator of the distance covariance, which is a V-statistic. The second builds on an unbiased estimator of the distance covariance proposed in Sz & eacute;kely and Rizzo (Stat 42:2382-2412 2014), shown to be a U-statistic by Huo and Sz & eacute;kely (Technometrics 58:435-447 2016). This study evaluates their efficiency (mean squared error) and compares computational times for both methods under different dependence structures. Under conditions of independence or near-independence, the V-estimates are biased, while the U-estimator frequently cannot be computed due to negative values. To address this challenge, a convex linear combination of the former estimators is proposed and studied, yielding good results regardless of the level of dependence. Additionally, a medical database is studied and discussed.
Keywords:
Distance correlation
U-statistic
V-statistic
Simulation study

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.5K
Citations:
1.7W

Organization

U
Universidade da Coruna
Scholars:
6.6K
Papers: 5.7K
Citations: 11