Range Restriction — Variance Restriction & Correlation Attenuation

Dr. R. Düsing · Osnabrück University

Help — Range Restriction

Does / does NOT

Does: shows how selecting on the predictor X systematically biases the correlation r observed in the selected group — and how Thorndike's formulas back-correct the true value ρ. Does NOT: no estimation from real data, no indirect case with a measured third variable (only a demonstration of the mechanism under bivariate normality).

What is this about?

If a sample covers only part of the value range of X (e.g. because only applicants above a test cutoff are hired), the spread of X shrinks. The correlation between X and Y is then underestimated — even though the true relationship in the overall population is much stronger. This is direct range restriction (Thorndike Case II).

The running example

A company validates an aptitude test X against later career success Y. But: only those who score above the cutoff are hired — only for them is a Y value available later. The validity r observed this way is smaller than the true validity ρ. Anyone who ignores range restriction wrongly writes off a good test as useless.

The mechanism (why?)

Correlation = standardized covariance. Selecting on X caps the spread of X (sx drops), while the residual spread of Y around the regression line stays the same. As a result, the systematic variation makes up a smaller share of Y's total variance → r drops. The key quantity is u = s'x / sx (selected spread ÷ total spread).

The most important point: the slope stays!

Under direct selection on X, the regression slope bY·X theoretically stays unchanged — only the correlation (and R²) drop. In the right plot, the red line has essentially the same slope as the blue population line; the points just scatter over a narrower X range. Correlation ≠ slope.

The visualizations

Top left: overall population. Gray points = not selected, red = selected. The orange line is the selection cutoff on X; the shaded area drops out. Top right: only the selected group — the narrower X range is immediately visible. Red line = regression in the selection, dashed blue = population line (same slope!). The spread bars below the plots show ±1 sx.

Bottom curve: the Thorndike relationship r(u). For the current true correlation ρ, it shows how the observed correlation depends on the spread ratio u. At u = 1 (no restriction), r = ρ. To the left (u < 1, restriction), r drops; to the right (u > 1, expansion via extreme groups), r rises. The point marks your current state.

Controls

Selection mode: Top (classic selection via cutoff) · Extreme groups (top + bottom — increases r!) · Middle (middle range). Selection fraction: what % remains. ρ: true population correlation. Thorndike correction: back-calculates the corrected value from rselected and u — it should hit ρ.

The symbols in detail

ρ (rho) — the true correlation in the overall population, i.e. the test's real validity.
r — the correlation observed in the selected group (what you actually measure).
sx — spread (standard deviation) of X in the overall population.
s′x — spread of X in the selected group (the prime ′ stands for "restricted").
u = s′x / sx — the spread ratio: how much of the original spread remains. u = 1 means no restriction, u < 1 restriction, u > 1 expansion (extreme groups).
U = 1/u = sx / s′x — simply the reciprocal of u; it only shows up because the correction formula computes "from restricted to overall" (≥ 1 under restriction).
b — the regression slope of Y on X; it stays unbiased under direct selection on X.
ρ (rho-hat) — the Thorndike-corrected estimate of ρ. The hat "^" always marks an estimated/computed value in statistics.

Formulas (Thorndike, Case II)

Attenuation: r = ρu / √(1 − ρ²(1 − u²))  ·  Correction: ρ = rU / √(1 + r²(U² − 1)) with U = 1/u = sx/s′x.

References

Thorndike, R. L. (1949). Personnel Selection: Test and Measurement Techniques. Wiley.
Sackett, P. R. & Yang, H. (2000). Correction for range restriction: An expanded typology. Journal of Applied Psychology, 85(1), 112–118.
Hunter, J. E., Schmidt, F. L. & Le, H. (2006). Implications of direct and indirect range restriction for meta-analysis methods and findings. Journal of Applied Psychology, 91(3), 594–612.

📋 Example — Aptitude Testing
A company checks whether the aptitude test X predicts later career success Y. The true validity in the applicant population is ρ = 0.60. But only 20% are hired (selection on X). In this restricted group, only r = is observed — the test looks weaker than it is. The Thorndike correction recovers ρ = .
What do the symbols mean?
ρtrue correlation in the overall population (here: the test's true validity)
robserved correlation in the selected group
sxspread (SD) of X in the overall population
s′xspread of X in the selected group (prime ′ = restricted)
uspread ratio s′x/sx: < 1 = restriction, = 1 = none, > 1 = expansion
Ureciprocal 1/u (= sx/s′x) — notational helper in the correction formula
bregression slope Y on X — stays unbiased under selection on X
ρThorndike-corrected estimate of ρ from r and u (should hit ρ)
Overall Population vs. Restricted Group
ρ population
true correlation
u = s′ / s
spread ratio
r selected
observed
ρ Thorndike
corrected back
Conclusion:
Thorndike Relationship — Observed r over the Spread Ratio u
Formula (Thorndike, Case II — direct selection on X)
Attenuation:   r = ρ·u / √( 1 − ρ²(1 − u²) ) with u = s′x/sx
Correction:   ρ = r·U / √( 1 + (U² − 1) ) with U = 1/u = sx/s′x
Current: u =  ·  robserved =  ·  rformula(ρ,u) =  ·  ρcorrected =
Concepts
When a sample is obtained such that only part of a variable's value range is represented, this is called range restriction. Because correlation measures shared standardized variation, it drops as soon as the predictor's spread is capped. The observed relationship then underestimates the true one — a common reason why tests, selection procedures, or predictors seem "not to work."
Direct vs. indirect (Case II vs. III)
Direct range restriction (Thorndike Case II): selection happens on the predictor X itself (e.g. a cutoff on an aptitude test). Indirect (Case III): selection happens on a third variable Z that correlates with X — X only gets "restricted along with it." The indirect case is more common in practice (e.g. selection on an earlier overall rating) and needs Lawley's extended formula / the correction by Hunter, Schmidt & Le (2006). This tool shows the direct case.
Correlation drops, slope stays
The most important lesson: under direct selection on X, only the correlation is biased, while the regression slope bY·X stays unbiased. Reason: the conditional distribution Y|X doesn't change under X-selection. In the right plot, the red (selection) and dashed blue (population) lines have the same slope. If you need predictions, you can still estimate the regression from restricted data — if you're comparing effect sizes, you must correct.
Extreme groups — the reversal
Keeping only the top and bottom values (extreme-groups design) increases the spread of X (u > 1) and the correlation is artificially inflated. This is the mirror-image bias: popular for making effects "more visible," but r and d are then no longer transferable to the population. Try Preset C: the same true ρ, but r shoots up.
Connection to Taylor-Russell
The Taylor-Russell tables answer the economic follow-up question: if a test has (corrected) validity ρ, a selection ratio is applied, and there's a base rate of "suitable" candidates — what fraction of those selected succeed? Range restriction supplies the correct ρ as input; without correction, you systematically underestimate the procedure's usefulness. → Taylor-Russell Tables
Practice & literature
Range restriction is a core topic in meta-analyses of validity studies (personnel selection, clinical prediction, university admissions). Rules of thumb: (1) always report u = s′/s; (2) use the Case-II correction for selection on X, the Case-III correction for selection on third variables; (3) label corrected values as such. References: Thorndike (1949); Sackett & Yang (2000); Hunter, Schmidt & Le (2006).
Distinguishing it from Berkson's paradox
In common: both are selection effects — a subgroup selected on some criterion biases the observed correlation, and you never see the whole population. Difference: in range restriction, you select directly on the predictor X → r is dampened (or inflated for extreme groups), but the regression slope bY·X stays unbiased. In Berkson, you select on a collider Z (a common effect X→Z←Y) → a spurious correlation arises out of nowhere, and the slope is biased too. In short: here a variance problem (correctable with Thorndike), there a structural problem (never control for a collider). The bridge between them is indirect range restriction (Case III, selection on a third variable). → Berkson's Paradox & Collider Bias