Diagnostic Validity — Test & Criterion

Dr. R. Düsing · Osnabrück University
Running Example
XTest score — e.g. depression scale (PHQ-9, z-standardized) YCriterion — clinical diagnosis by structured interview (z-standardized) rValidity = correlation between test and criterion (r=0.50) Cut-Off XPHQ threshold for a positive screening result Cut-Off YCriterion threshold: above this, someone is considered "clinically affected"
What does a positive PHQ-9 finding actually get you? Validity r describes test quality — but PPV and NPV determine how usable the result is in this specific situation. Base rate, cut-off, and validity together determine the practical utility.
Confusion Matrix
Test +Test −Σ
Crit. + TPtrue positive FNfalse negative
Crit. − FPfalse positive TNtrue negative
Σ
Metrics
MetricFormulaCalculationValue
Concepts
Three Kinds of Validity
Validity asks: does the test measure what it's supposed to measure? A distinction is drawn between construct validity (does the test capture the theoretical construct?), criterion validity (does it predict an external criterion — concurrently or predictively?), and content validity (do the items representatively cover the trait?). This tool shows criterion validity: the relationship between test score X and a criterion Y.
Validity r & Variance Explained
Validity is operationalized here as the correlation r between test and criterion. is the proportion of shared variance — at r = 0.50, that's 25%. The 95% confidence ellipse shows how tight the relationship is: the narrower it is, the higher r. Good diagnostic tests typically fall around r = 0.40–0.70; r = 1 (perfect prediction) does not occur in practice.
Test Property vs. Situational Utility
Sensitivity (TP/(TP+FN)) and specificity (TN/(TN+FP)) are properties of the test — prevalence-independent and transferable to other populations. PPV and NPV, by contrast, answer the clinical question for this specific person: how likely is the condition really present given a positive (PPV) or negative (NPV) finding? They depend on the base rate.
The Base Rate Decides
PPV = TP/(TP+FP) drops drastically when the condition is rare: at low prevalence, the denominator is full of unaffected people, a small percentage of whom are false-positive — and the few true cases get lost among them. A highly valid test can still have a disappointingly low PPV for a rare condition. Move the criterion cut-off to change the base rate. → Sensitivity & Specificity
Likelihood Ratios
LR+ = Sens/(1−Spec) and LR− = (1−Sens)/Spec are prevalence-independent measures of result strength. Via the Bayes formula, they translate pre-test into post-test probability: post-odds = pre-odds · LR. Rule of thumb: LR+ > 10 shifts the diagnosis strongly toward "affected", LR− < 0.1 largely rules it out.
Validity is necessary but not sufficient — only together with base rate and cut-off does utility (a secondary quality criterion) emerge. In selection contexts (personnel, therapy slots), PPV is called the success rate: the proportion of those selected who meet the criterion — mathematically identical (TP/(TP+FP)), just a different field of application. The Taylor-Russell tables formalize this from validity, selection rate, and base rate. → Taylor-Russell Tables
Diagnostic Validity — Help
Example

A depression questionnaire (PHQ-9, z-standardized: X) is compared with the result of a clinical interview (criterion Y). The correlation r=0.50 describes the validity. Both cut-off values split the scatterplot into four quadrants — TP, TN, FP, FN. The central utility metric is PPV = TP/(TP+FP): of everyone who tests positive — how many are actually affected?

Validity r & Scatterplot

The correlation r determines how tight the relationship is between test (X) and criterion (Y). At r=0: the point cloud is circular — no diagnostic utility. At r=1: all points on a line — perfect prediction.

In the example: r=0.50 means the PHQ-9 explains about 25% of the variance in the criterion (r²=0.25). Good diagnostic tests typically have r=0.40–0.70.

Cut-Off Lines & Confusion Matrix

The vertical line (X = test cut-off) and horizontal line (Y = criterion cut-off) define four groups:

TP: affected (Y > Cut_Y) + tested positive (X > Cut_X) FP: unaffected (Y ≤ Cut_Y) + tested positive → false alarm FN: affected + tested negative → missed diagnosis TN: unaffected + tested negative → correctly excluded

In the example: moving the test cut-off to the left raises sensitivity (more TP), but also FP. The cut-off lines can be dragged directly on the canvas.

Diagnostic Metrics — PPV & NPV in Focus

Sensitivity and specificity describe how well the test itself works — prevalence-independent, a property of the test. In everyday clinical and diagnostic practice, though, a different question is central: what does this specific result tell me about this specific person? PPV and NPV answer that question.

PPV (positive predictive value) = TP/(TP+FP): if someone tests positive — how likely is it that a condition is actually present? NPV (negative predictive value) = TN/(TN+FN): if someone tests negative — how certain can one be that no condition is present?

PPV = (Sens · prevalence) / (Sens · prev + (1−Spec) · (1−prev))

Both values depend strongly on the base rate (prevalence). For rare conditions, PPV stays low even with a good test — the denominator is full of unaffected people, a small proportion of whom are still false-positive. In the example: if depression prevalence is 10% instead of 50%, PPV drops drastically at the same sens and spec.

Sensitivity = TP/(TP+FN) and specificity = TN/(TN+FP) are the foundation: prevalence-independent and transferable across populations. Via LR, they determine how strongly a test result changes the pre-test probability.

LR+ = Sens/(1−Spec), LR− = (1−Sens)/Spec: prevalence-independent measures of result strength; directly usable to convert pre- into post-test probability (Bayes formula).

Utility as a Secondary Quality Criterion

In psychological assessment, alongside the primary quality criteria (objectivity, reliability, validity), secondary quality criteria are also required — among them utility. It asks: is it worth using this test in this situation — does it improve decisions relative to the status quo?

PPV and NPV are the direct operationalizations of utility in the diagnostic context: a test is useful if a positive finding substantially raises the probability of a condition (high PPV) and a negative finding substantially lowers it (high NPV). Validity r is a necessary but not sufficient condition — base rate and cut-off determine how much utility r delivers under real conditions.

Relationship r → Metrics

As r rises, the point clouds move apart — sensitivity and specificity improve simultaneously, and so do PPV and NPV (at the same base rate). As r falls, sensitivity and specificity become a zero-sum game: any improvement in one comes at the cost of the other — and PPV/NPV stagnate or worsen depending on the cut-off choice.

Taylor-Russell Tables — a Preview

For selection contexts (personnel, therapy slots, support programs), the Taylor-Russell tables (1939) formalize the relationship between utility, validity, and boundary conditions. They connect three quantities:

From these three quantities the success rate emerges: the proportion of those selected who are actually qualified. This shows when a test with moderate validity (r=0.30) has substantial utility under favorable conditions (low SR, medium BR) — and when a test with high validity delivers little (e.g. when almost everyone would be selected anyway). The interactive Taylor-Russell tool displays this success rate as a table and a nomogram with a target crosshair.

Success rate = PPV — mathematically the same thing, different language: in clinical diagnostics, the proportion of true positives among everyone who tests positive is called PPV; in aptitude assessment and personnel psychology, the proportion of qualified people among everyone selected is called the success rate. The formula is identical: TP/(TP+FP). The term changes with the field of application — the concept stays the same.

References

Taylor, H. C. & Russell, J. T. (1939). The relationship of validity coefficients to the practical effectiveness of tests in selection. Journal of Applied Psychology, 23(5), 565–578.