Running Example
XPredictor: selection-test score (cognitive aptitude measure, z-standardized)
YCriterion: job performance (supervisor rating, z-standardized)
Group AReference group (majority group, e.g. group without an immigrant background)
Group BFocal group (minority group, e.g. group with an immigrant background)
Is the selection test equally valid for both groups — or does it systematically produce wrong performance predictions for one group? And even if it's fair: can different selection rates still result?
① Scatterplot
Module 1 — Cleary Test: Differential Prediction
Cleary (1968) — Definition of Test Bias
A test is considered biased if the regression equation for the overall sample
produces systematically wrong predictions for one group. Cleary tests this with a
hierarchical regression model in three steps:M1: Ŷ = b₀ + b₁·X (shared regression line) M2: Ŷ = b₀ + b₁·X + b₂·G (same slope, different intercepts) M3: Ŷ = b₀ + b₁·X + b₂·G + b₃·(X×G) (different slopes) M2 vs. M1 tests intercept differences (intercept bias).
M3 vs. M2 tests slope differences (slope bias).
Module 2 — Meade & Fetzer (2009): What's Behind It?
Differential Prediction ≠ Test Bias
A significant intercept difference (M2 vs. M1) can have four different causes.
Only one of them is genuine test bias in Cleary's sense.Source 1 — Test bias: the test measures the construct differently for the two groups — group B scores lower on X even though their real Y performance is comparable. d_X large, d_Y small. Effect: qualified people from B are systematically screened out; the test underestimates their actual capability.
Source 2 — Criterion bias: the criterion Y (e.g. supervisor rating) rates one group better or worse, independent of their actual performance. d_X ≈ 0, d_Y large. Effect: the test itself is fair — but the validation target is biased. Selection based on this criterion cements the disadvantage.
Source 3 — Missing variables: additional predictors (e.g. length of training, socioeconomic status) would explain the intercept difference. Effect: the apparent bias disappears after controlling for them — action: extend the model.
Source 4 — Sampling error: random deviations, especially with small n. Effect: no systematic bias; replication settles the question.
Diagnosis: the standardized mean differences d_X and d_Y point to the most likely source (cf. Meade & Fetzer, 2009, Fig. 1–3).
Module 3 — Adverse Impact & the 4/5 Rule
Adverse Impact
Adverse impact is present when one group is selected considerably less often
than another in a selection decision — even if the test itself is fair.
The US Uniform Guidelines (1978) formulate the 4/5 rule:
the selection rate of the disadvantaged group should be at least 80% of the rate
of the favored group.AIR = selection rate_B / selection rate_A Adverse impact ≠ test bias. A fair test can produce adverse impact when real group differences exist on the predictor.
Concepts
Test Bias per Cleary (1968)
A test is biased if the shared regression of criterion (Y) on predictor (X) yields systematically wrong predictions for one group — differential prediction. It's tested whether the intercept and slope of the regression line are equal for the reference (A) and focal (B) groups. If they are, the test is fair in Cleary's sense.
Intercept vs. Slope Bias
Intercept difference: one group is consistently over- or under-predicted (parallel but offset lines). Slope difference: the test predicts the criterion less well for one group (differing validity). Module 1 tests both hierarchically via F-test (intercept: M2 vs. M1; slope: M3 vs. M2).
Diagnosing the Cause (Meade & Fetzer)
The standardized mean differences
d_X (predictor) and d_Y (criterion) point to the source: d_X large, d_Y small → test bias; both large & proportional → adverse impact without bias; only d_Y → criterion bias; both small → no problem. Four patterns, only one is genuine bias per Cleary.Adverse Impact & the 4/5 Rule
Adverse impact concerns the selection outcomes, not test quality:
AIR = selection rate_B / selection rate_A. If the AIR falls below 0.80 (four-fifths rule, US EEOC), that counts as a legal indicator of disadvantage. It's about hit rates between groups, not prediction error.Test Bias ≠ Adverse Impact
The central point: a fair test (identical regression for both groups) can still produce adverse impact — namely when the groups have different predictor means. Bias is a prediction problem, adverse impact a distribution/selection problem. The two must not be confused.
Context & Related Tools
Differential prediction at the test level has a counterpart at the item level: Differential Item Functioning. The underlying validity concept and the selection utility are explored in depth in the validity tools. → Differential Item Functioning · → Diagnostic Validity · → Taylor-Russell