CTT — Foundations · X = T + E

Dr. R. Düsing · Osnabrück University
📋 Example — IQ Test (IQ Scale)
An IQ test has M = 100. Every observed score X is the sum of the true score T (the actual trait level) and a random measurement error E: X = T + E. Reliability is the proportion of variance attributable to true differences — not to noise. With σT = 14 and σE = 5, this gives rtt ≈ .89 and an SEM of 5 points.
A test score is usually a sum of several items — how do you get from there to reliability? → What measurement model formally underlies this? → Measurement Models
Reliability & Measurement Error
Reliability rtt
σ²T / (σ²T+σ²E)
SEM
SDX·√(1−rtt) = σE
SD observed
√(σ²T+σ²E)
Empirical r (retest)
cor(X₁, X₂) of the sample
true variance
error
True variance (reliable) = rtt Error variance = 1 − rtt
Reliability = Correlation of Two Measurements
Spearman-Brown — Reliability & Test Length
Confidence Interval Around a Single Test Score

A single person takes the test once and gets the observed score X (the "Observed score X" slider in the sidebar). Because of measurement error, X is not exactly the true score τ. CTT offers two confidence intervals for this: the equivalence hypothesis (CI around the observed score X) and the regression hypothesis (CI around the regression-corrected estimate X′).

95% confidence interval: equivalence (red) vs. regression hypothesis (green)
⚠ A confidence interval is not a probability range for τ and not an estimate of the true score itself — it is a coverage interval (under repeated measurement, it contains τ in, say, 95% of cases). The point estimate of the true score is X or X′; the interval only expresses the uncertainty. A direct probability statement about τ requires the Bayesian credible interval — CIs and CrIs are covered in detail in the tool → Diagnostic Intervals
Concepts
Reliability Is a Proportion of Variance
rtt tells you what proportion of the observed spread is due to true differences between people. r = .90 means: 90% true variance, 10% measurement noise. Reliability is not a property "of the test" but of test + population: the same scale is less reliable in a homogeneous group.
SEM Depends on Reliability
The standard error of measurement SEM = SD·√(1−r) translates reliability into the scale unit. Here, SEM even equals σE. It is the basis for confidence bands around individual scores — and thus for change measurement (reliable change).
→ Jacobson-Truax
Longer Tests Are More Reliable
The Spearman-Brown formula shows: more (equivalent) items increase reliability — but with diminishing returns. Doubling helps a lot; a tenfold increase, little. The tool computes how many items are needed for a target reliability.
Measurement Error Distorts Associations
Unreliability attenuates correlations: the observed association is smaller than the true one. The disattenuation formula corrects for this.
→ Measurement Error Attenuation
CTT Foundations — Background
What This Tool Shows — and What It Doesn't

Shows: the true-score model X = T + E, reliability as a proportion of variance, SEM, confidence bands around individual scores, and the Spearman-Brown test-length formula. Not here: the concrete estimation procedures (α, ω, split-half) — covered by the "Reliability: α vs. ω" tool; the measurement models behind it by the "Measurement Models" tool.

The Classical Test Theory Model
X = T + E (observation = true score + error) E[E] = 0, Cov(T,E) = 0 Var(X) = Var(T) + Var(E)

The error is random, zero on average, and uncorrelated with the true score. From this follows the central definition of reliability:

r_tt = Var(T) / Var(X) = Var(T) / [Var(T) + Var(E)]
Reliability as Correlation (Panel ②)

Two parallel measurements X₁ and X₂ of the same people correlate exactly at rtt — because their shared part is the true score T. This is the operational basis of test-retest and parallel-test reliability. The empirically estimated value fluctuates around the theoretical one (sampling error — visibly smaller with larger n).

Standard Error of Measurement & Confidence Intervals (Panel ④)

A single person gets the observed score X on one test administration. Because of measurement error, X deviates from the (unknown) true score τ. The standard error of measurement describes how much individual measurements scatter around τ:

SEM = SD_X · √(1 − r_tt) (here it also equals σ_E)

From this, CTT constructs a confidence interval — in two readings, exactly as in the "Diagnostic Intervals" tool:

Equivalence hypothesis (red): the interval is centered around the observed score X. Assumption: X is an unbiased estimator of τ.

CI = X ± z · SEM

Regression hypothesis (green): the interval is centered around the regressed estimate X′, which accounts for regression to the mean — extreme scores contain more measurement error and get pulled toward the mean:

X′ = M + r_tt · (X − M) SE_reg = SD_X · √(r_tt · (1 − r_tt)) CI = X′ ± z · SE_reg

Important for interpretation: a confidence interval is not a probability statement about τ and is not itself "the estimate of the true score." The point estimate is X or X′; the interval only expresses that, under repeated measurement, it covers the true score in, say, 95% of cases. A direct statement "with 95% probability τ lies in …" is only possible with the Bayesian credible interval. Confidence and credible intervals (as well as percentile-rank placement) are covered in detail in the tool Diagnostic Intervals.

Spearman-Brown: Test Length & Reliability (Panel ③)

Reliability can be increased by adding items — because the shared true-score part of several items sums up more strongly than their independent measurement error. The Spearman-Brown prophecy formula predicts what reliability a test lengthened by a factor of k would have:

r_k = (k · r) / (1 + (k − 1) · r)

k is the length factor: k = 2 means double the number of items, k = 0.5 a halved test. The curve in the plot rises steeply at first and then flattens — diminishing returns: for an already reliable test, doubling barely helps; for an unreliable one, it helps a lot. At k = 1 sits the current test (red dot); the purple line marks the target reliability.

Rearranging the formula gives the required test length for a target reliability r*:

k = r* · (1 − r) / [ r · (1 − r*) ]

Example: a test with r = .70 should reach r* = .90 → k = .90·.30 / (.70·.10) ≈ 3.86, i.e. almost four times the number of items. Requirement: the added items are equivalent to the existing ones (same quality, parallel) — otherwise the formula overestimates the gain.

Reliability Is Context-Dependent

Since rtt = Var(T)/Var(X), it decreases when the true variance shrinks (homogeneous group, range restriction) — with the same measurement error. Reliability thus belongs to test and population, not to the test alone.

Literature

Lord, F. M. & Novick, M. R. (1968). Statistical Theories of Mental Test Scores. Addison-Wesley.
Spearman, C. (1910) & Brown, W. (1910), prophecy formula.
Gulliksen, H. (1950). Theory of Mental Tests. Wiley.