arrow
Return

Why Overfitting Is Not (Usually) a Problem in Partial Correlation Networks

delete2022-10-01
delete7
delete
OA
AI
D
Donald R. Williams *
J
Josue E. Rodriguez
DOI:10.1037/met0000437delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Translational Abstract It is vital to clearly understand the benefits and limitations of regularized networks as inferences drawn from them may hold methodological and clinical implications. This article addresses a core rationale for the increasing adoption of regularized estimation. Namely, that it reduces overfitting. Accordingly, we elucidate important aspects of overfitting and the bias-variance tradeoff that are especially relevant for network research, where the number of variables is small relative to the number of observations (i.e., a low p/n ratio). We find that bias, and especially variance, are the most problematic aspects for inference in p/n ratios that are rare to psychometric settings. We then introduce a nonregularized method based on classical techniques that fulfill two desiderata: (1) reducing or controlling chance findings and (2) avoiding overfitting by providing accurate predictions. In several simulation studies, our nonregularized method provided more than competitive predictive performance, and in many cases, outperformed regularized networks. It appears to be nonregularized, as opposed to regularized estimation, that best satisfies these desiderata. We then provide insights into using our methodology. Here we discuss the multiple comparisons problem in relation to prediction: stringent alpha levels, resulting in a network with few associations, can deteriorate predictive accuracy. We end by emphasizing key advantages of our approach that make it ideal for both inference and prediction in network analysis. Network psychometrics is undergoing a time of methodological reflection. In part, this was spurred by the revelation that l(1)-regularization does not reduce spurious associations in partial correlation networks. In this work, we address another motivation for the widespread use of regularized estimation: the thought that it is needed to mitigate overfitting. We first clarify important aspects of overfitting and the bias-variance tradeoff that are especially relevant for the network literature, where the number of nodes or items in a psychometric scale are not large compared to the number of observations (i.e., a low p/n ratio). This revealed that bias and especially variance are most problematic in p/n ratios rarely encountered. We then introduce a nonregularized method, based on classical hypothesis testing, that fulfills two desiderata: (a) reducing or controlling the false positives rate and (b) quelling concerns of overfitting by providing accurate predictions. These were the primary motivations for initially adopting the graphical lasso (glasso). In several simulation studies, our nonregularized method provided more than competitive predictive performance, and, in many cases, outperformed glasso. It appears to be nonregularized, as opposed to regularized estimation, that best satisfies these desiderata. We then provide insights into using our methodology. Here we discuss the multiple comparisons problem in relation to prediction: stringent alpha levels, resulting in a sparse network, can deteriorate predictive accuracy. We end by emphasizing key advantages of our approach that make it ideal for both inference and prediction in network analysis.
Keywords:
partial correlation network
overfitting
prediction
frequentist inference
mean squared error

Journal

Psychological Methods cover
Psychological Methods
IF:
7.8
Papers:
1.3K
Citations:
2.1W

Organization

University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K