An IQ test has M = 100. Every observed score X is the sum of the true scoreT
(the actual trait level) and a random measurement errorE: X = T + E.
Reliability is the proportion of variance attributable to true differences — not to noise.
With σT = 14 and σE = 5, this gives rtt ≈ .89 and an SEM of 5 points.
A single person takes the test once and gets the observed score
X (the "Observed score X" slider in the sidebar). Because of measurement error, X is not exactly the
true score τ. CTT offers two confidence intervals for this:
the equivalence hypothesis (CI around the observed score X) and the
regression hypothesis (CI around the regression-corrected estimate X′).
95% confidence interval: equivalence (red) vs. regression hypothesis (green)
⚠ A confidence interval is not a probability range for τ and not an
estimate of the true score itself — it is a coverage interval (under repeated measurement, it contains τ
in, say, 95% of cases). The point estimate of the true score is X or X′; the interval only expresses the
uncertainty. A direct probability statement about τ requires the Bayesian credible interval — CIs and CrIs are
covered in detail in the tool
→ Diagnostic Intervals
rtt tells you what proportion of the observed spread is due to true
differences between people. r = .90 means: 90% true variance, 10% measurement noise. Reliability is not a property
"of the test" but of test + population: the same scale is less reliable in a homogeneous
group.
SEM Depends on Reliability
The standard error of measurementSEM = SD·√(1−r) translates reliability into
the scale unit. Here, SEM even equals σE. It is the basis for confidence bands around individual scores —
and thus for change measurement (reliable change).
→ Jacobson-Truax
Longer Tests Are More Reliable
The Spearman-Brown formula shows: more (equivalent) items increase
reliability — but with diminishing returns. Doubling helps a lot; a tenfold increase, little. The tool computes
how many items are needed for a target reliability.
Measurement Error Distorts Associations
Unreliability attenuates correlations: the observed
association is smaller than the true one. The disattenuation formula corrects for this.
→ Measurement Error Attenuation
CTT Foundations — Background
What This Tool Shows — and What It Doesn't
Shows: the true-score model X = T + E, reliability as a proportion of variance, SEM,
confidence bands around individual scores, and the Spearman-Brown test-length formula. Not here: the
concrete estimation procedures (α, ω, split-half) — covered by the "Reliability: α vs. ω" tool; the measurement
models behind it by the "Measurement Models" tool.
The Classical Test Theory Model
X = T + E (observation = true score + error)
E[E] = 0, Cov(T,E) = 0
Var(X) = Var(T) + Var(E)
The error is random, zero on average, and uncorrelated with the true score. From this follows the
central definition of reliability:
Two parallel measurements X₁ and X₂ of the same people correlate exactly at rtt —
because their shared part is the true score T. This is the operational basis of test-retest and
parallel-test reliability. The empirically estimated value fluctuates around the theoretical one (sampling error —
visibly smaller with larger n).
A single person gets the observed score X on one test administration. Because of
measurement error, X deviates from the (unknown) true score τ. The standard error of measurement
describes how much individual measurements scatter around τ:
SEM = SD_X · √(1 − r_tt) (here it also equals σ_E)
From this, CTT constructs a confidence interval — in two readings, exactly as in the "Diagnostic
Intervals" tool:
Equivalence hypothesis (red): the interval is centered
around the observed score X. Assumption: X is an unbiased estimator of τ.
CI = X ± z · SEM
Regression hypothesis (green): the interval is centered
around the regressed estimate X′, which accounts for regression to the mean — extreme scores contain more
measurement error and get pulled toward the mean:
X′ = M + r_tt · (X − M)
SE_reg = SD_X · √(r_tt · (1 − r_tt))
CI = X′ ± z · SE_reg
Important for interpretation: a confidence interval is not
a probability statement about τ and is not itself "the estimate of the true score." The point estimate is X or
X′; the interval only expresses that, under repeated measurement, it covers the true score in, say, 95%
of cases. A direct statement "with 95% probability τ lies in …" is only possible with the
Bayesian credible interval. Confidence and credible intervals (as well as percentile-rank
placement) are covered in detail in the tool
Diagnostic Intervals.
Reliability can be increased by adding items — because the shared true-score part
of several items sums up more strongly than their independent measurement error. The Spearman-Brown prophecy
formula predicts what reliability a test lengthened by a factor of k would have:
r_k = (k · r) / (1 + (k − 1) · r)
k is the length factor: k = 2 means double the number of items, k = 0.5 a halved
test. The curve in the plot rises steeply at first and then flattens — diminishing returns: for an
already reliable test, doubling barely helps; for an unreliable one, it helps a lot. At k = 1 sits the current
test (red dot); the purple line marks the target reliability.
Rearranging the formula gives the required test length for a target reliability r*:
k = r* · (1 − r) / [ r · (1 − r*) ]
Example: a test with r = .70 should reach r* = .90 → k = .90·.30 / (.70·.10) ≈ 3.86, i.e. almost
four times the number of items. Requirement: the added items are equivalent to the existing ones
(same quality, parallel) — otherwise the formula overestimates the gain.
Reliability Is Context-Dependent
Since rtt = Var(T)/Var(X), it decreases when the true variance shrinks (homogeneous
group, range restriction) — with the same measurement error. Reliability thus belongs to test
and population, not to the test alone.