Lindley's Paradox — significant, but Bayes favors H₀

Dr. R. Düsing · Osnabrück University
📋 The Puzzle
A test against the point null hypothesis H₀: θ = 0 on n = 50,000 people yields p ≈ .05 — so significant, H₀ is rejected. Yet the Bayes factor says: the data are many times more probable under H₀ than under H₁. Frequentist and Bayesian reach opposite conclusions — from the same data. Who's right, and why?
Frequentist vs. Bayes
Test statistic z
δ·√n
p-value (2-sided)
BF₀₁ (for H₀)
data×more probable under H₀
BF₁₀ (for H₁)
Jeffreys scale
Why? The Distribution of the Test Statistic under H₀ vs. H₁
Marginal distribution of z — H₀ (fixed) vs. H₁ (widens with n)
The Divergence with Growing n
p-value (red) vs. Bayes factor BF₀₁ (purple) over sample size
Concepts
What exactly is the paradox?
At a fixed significance level, a just-significant result (p = .05) can become ever-stronger evidence for H₀ as n grows. The frequentist rejects H₀, the Bayesian supports it — both from the same data. Lindley (1957) proved this formally.
The effect becomes tiny
For p = .05 to persist, the observed effect δ = z/√n must keep shrinking as n grows large. A "significant" effect at n = 1,000,000 is practically zero. Significance here only measures that the effect is not exactly 0 — not that it's meaningful.
Why does Bayes favor H₀?
H₁ spreads its probability over all possible effect sizes (prior τ). The larger n, the more sharply a real effect would show up — a tiny z then speaks against large effects and thus for H₀. H₁ is automatically "penalized" for its spread (Occam's razor).
The role of the prior
The Bayes factor depends on the prior width τ — a well-known point of criticism. A wider prior → stronger support for H₀ (more "penalty" for H₁). The p-value ignores this question entirely, but pays for that with significance's dependence on n.
Why does the BF "flip" in free mode?
At a fixed true effect δ (Lindley mode off), BF₀₁ goes through two phases, because it's made of two opposing forces: BF₀₁ = √(1+nτ²) · exp(−½ z²·…).
• The Occam factor √(1+nτ²) — grows with n and favors H₀: it penalizes H₁ for spreading its probability over a wide prior.
• The fit term exp(−½z²·…) — collapses once z gets large (a real effect becomes visible), and favors H₁.
Small n: z = δ√n is tiny → the effect can't be resolved yet → Occam wins → BF favors H₀. Large n: the real effect is resolved → fit wins → BF favors H₁. The "flip" is thus Occam's razor at work: H₁ has to earn its advantage through enough data first.
Punchline: if a real effect exists, p and BF eventually agree (both → H₁) — then it's really just different sensitivity plus an Occam surcharge. The true paradox only occurs in Lindley mode: there the effect shrinks (δ = z/√n → 0), p stays at α forever, and BF₀₁ rises monotonically toward H₀ — they disagree permanently.
What about a credible interval / probability of direction?
With a flat prior, the posterior is N(x̄, SE²). The probability of direction is then pd = Φ(|z|) = 1 − p/2 — a monotone function of the p-value. In Lindley mode, pd would stay constant at 1 − α/2 (e.g. 0.975) and the flat-prior CrI would always exclude 0 narrowly → both behave like the p-value and do not show the paradox. Reason: pd and CrI are estimation measures ("where does θ lie?") without a point null and thus without an Occam factor. Only the Bayes factor (a point-null comparison) or a ROPE — where the CrI, shrinking with n, falls entirely inside the region of practical equivalence — resolve the paradox toward H₀. In short: the Lindley paradox is a phenomenon of point-null model comparison, not of Bayesian estimation.
Resolution & connection to the Bayes Thinking Lab
The paradox is not a calculation error — it shows that the p-value and the Bayes factor answer different questions. The p-value: "How surprising are the data under H₀?" The Bayes factor: "Which hypothesis predicts the data better?" At large n with a point null, the two diverge. For a deeper dive into priors, Bayes factors, and ROPE decisions, see the sister project:
→ Bayes Thinking Lab
Lindley's Paradox — Background
The setup

We test the mean θ of a normal distribution with known spread σ. We observe the sample mean x̄ from n observations, standardized as z = (x̄ − μ₀)·√n / σ. All effects are expressed in σ units (σ = 1); the observed effect is δ = (x̄ − μ₀)/σ, so z = δ·√n.

Frequentist: the p-value
p = 2·(1 − Φ(|z|)) (two-sided)

H₀ is rejected when p < α. For fixed δ > 0, z = δ√n grows with n — eventually any effect, no matter how tiny, becomes significant. Significance is thus strongly dependent on n.

Bayesian: the Bayes factor

H₀: θ = 0 (point null) against H₁: θ ~ N(0, τ²). The Bayes factor compares how well each hypothesis predicts the data (marginal likelihood). For this normal-normal model:

BF₀₁ = √(1 + n·τ²) · exp( −½ · z² · n·τ²/(1 + n·τ²) )

BF₀₁ > 1 means evidence for H₀, BF₁₀ = 1/BF₀₁ evidence for H₁. Jeffreys' rule of thumb: 1–3 anecdotal, 3–10 moderate, 10–30 strong, >30 very strong.

The paradox

If you hold the result just significant (z = zcrit, i.e. p = α constant) and let n grow, √(1+nτ²) → ∞ while the exponential term tends toward exp(−½ zcrit²). So BF₀₁ → ∞: the same "significant" result becomes ever-stronger evidence for H₀. Visually (Panel ②): the z distribution under H₀ stays fixed at N(0,1), while the one under H₁ grows ever wider and flatter with n. A fixed z = 1.96 then still lies in H₀'s rejection region, but under the wide H₁ it has an even smaller density — the data fit H₀ better.

The resolution

The p-value and the Bayes factor answer different questions. The p-value conditions only on H₀ ("how extreme are the data, if H₀ holds?"). The Bayes factor compares H₀ and H₁ directly and "penalizes" H₁ for its vagueness (Occam). Also: an effect that's just significant at large n is practically zero — significance only tells you that θ ≠ 0 exactly. Anyone who wants to test for a relevant minimum size needs an interval/ROPE instead of a point null anyway.

Limits of this tool

A normal prior is used under H₁ (analytically clean). Common default Bayes factors (JZS, Rouder et al.) use a Cauchy prior — qualitatively the same paradox, different numbers. The BF visibly depends on τ.

References

Lindley, D. V. (1957). A statistical paradox. Biometrika, 44, 187–192.
Jeffreys, H. (1939). Theory of Probability.
Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychon. Bull. Rev., 14, 779–804.
Rouder et al. (2009). Bayesian t tests. Psychon. Bull. Rev., 16, 225–237.