arrow
Return

Testing Significance Testing

delete2018-04-26
delete5
delete
OA
AI
J
Joachim I. Krueger *
P
Patrick R. Heck *
DOI:10.1525/collabra.108delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The practice of Significance Testing (ST) remains widespread in psychological science despite continual criticism of its flaws and abuses. Using simulation experiments, we address four concerns about ST and for two of these we compare ST's performance with prominent alternatives. We find the following: First, the p values delivered by ST predict the posterior probability of the tested hypothesis well under many research conditions. Second, low p values support inductive inferences because they are most likely to occur when the tested hypothesis is false. Third, p values track likelihood ratios without raising the uncertainties of relative inference. Fourth, p values predict the replicability of research findings better than confidence intervals do. Given these results, we conclude that p values may be used judiciously as a heuristic tool for inductive inference. Yet, p values cannot bear the full burden of inference. We encourage researchers to be flexible in their selection and use of statistical methods.
Keywords:
statistical significance testing
null hypotheses
Bayes' Theorem
NHST
p values

Journal

Collabra Psychology cover
Collabra Psychology
IF:
3.2
Papers:
616
Citations:
1.5K

Organization

B
Brown University
Scholars:
2.4W
Papers: 2.2W
Citations: 3.2W
G
geisinger health system
Scholars:
1.1K
Papers: 786
Citations: 2