Skip to content
Sunday, August 23, 2026
BLOGDAILYGADGETS · APPS · REVIEWS
Apps

What a p-value actually means (and what it doesn't)

A p-value is the probability of data at least as extreme as the observed data if the null hypothesis were true — not the probability that the hypothesis is true, and not the size of the effect.

NB
Naomi Bergman, · July 12, 2026 · 3 min read
Laboratory notebook and calculator beside scattered data cards

A p-value is the probability, computed under a statistical model, of obtaining data at least as extreme as what a study observed, if there were actually no effect — the null hypothesis. A p-value of 0.03 means data like this would arise 3 percent of the time by chance alone under that model. It is not the probability the finding is true, and it says nothing about how big the effect is — the two most common misreadings, flagged formally by the American Statistical Association's 2016 statement on p-values. Engevity News explains the statistic; the explanation is the whole story of why "statistically significant" means less than it sounds.

The p-value answers a narrow question precisely. The trouble begins when it is asked the questions readers actually care about.

What question does the p-value actually answer?

One question: how surprising is this data if nothing is going on? The framework is hypothesis testing. A researcher assumes the null hypothesis — no difference between groups, no association between variables — and computes how often chance alone would produce results at least as extreme as the observed ones. Small p-values mean the data would be unusual under the null; convention, dating to Ronald Fisher's 1925 writing and hardened by journal practice, sets the threshold at 0.05. The convention is arbitrary. Fisher himself described 0.05 as one informal guideline among others, and the American Statistical Association's statement urges treating 0.05 as a line of no special honor.

What doesn't a p-value tell you?

Three things, per the ASA's six principles. First, it is not the probability the null hypothesis is false — that probability, the thing readers want, requires Bayesian methods the p-value does not perform. Second, it carries no effect-size information: a study of two million people can find p < 0.001 for an effect so small it matters to no one, while a large, genuinely important effect in a small study can miss the threshold. Third, it does not measure the quality of the study design — a p-value can be computed from biased data as smoothly as from clean data. Statistical significance and practical significance are different claims, and only the first belongs to the p-value.

Why do so many false findings pass?

Arithmetic, mostly. If researchers test many hypotheses, five in a hundred will pass a 0.05 threshold by chance alone — and modern research tests many hypotheses, sometimes thousands per paper, and reports the passes. When a field's tested hypotheses are mostly false, most positive findings will be false too, a structural argument the epidemiologist John Ioannidis made in a widely cited 2005 essay and that later replication projects bore out: large replication attempts in psychology and other fields have reproduced substantially fewer than half of tested findings, per the projects' published results. Selective reporting compounds the arithmetic: analyses run many ways, one passing result reported, p < 0.05 attached — a practice the meta-science literature has documented from authors' own accounts.

How should a reader treat a p-value?

As one line in a larger ledger. The questions that matter alongside it:

  1. How big is the effect? Effect sizes and confidence intervals describe it; the p-value does not.
  2. How many comparisons were run? A p-value among thousands of tests means far less than the same p-value from a single pre-registered one.
  3. Was the analysis plan fixed in advance? Pre-registration removes the freedom to mine for significance.
  4. Has anyone repeated it? A replicated small effect outranks an unreplicated large one.

What the statistical literature establishes is a tool that answers its narrow question well and is routinely read as answering three broader ones it cannot touch. The correction fits in a sentence: ask how big, ask how often tested, ask who repeated it. Then the p-value earns its keep.