Inter-Rater Agreement

© Dr. Rainer Düsing · Diagnostics Course · Osnabrück University · Interactive Tools by Claude
Running Example
Stimuli (n)80 patient videos from an outpatient clinic — each video shows an intake session Raters (k)3 clinical psychologists — rate each video independently Dimension 1depression severity (PHQ category: none / mild / moderate / severe) Dimension 2risk classification (0 = no risk … 4 = acute) Format3 CSV files (80 rows × 2 columns each) — one per rater
How well do the three psychologists agree — and does it matter whether we correct for chance agreement (κ, α), partial out rater bias (ICC), or use robust corrections (Gwet AC)?
Upload rater files, choose scale level and method,
then click Compute.
Global Agreement
Agreement by Dimension
Per-Stimulus Agreement
Overall Measure — All Raters
Box-and-whisker plot: distribution of the agreement measure across all stimuli. Diamond = mean, line = median.
Histogram of the stimulus-specific agreement values with an overlaid normal distribution (blue) and mean line.
Trace plot: agreement value per stimulus in dataset order. LOESS trend line shows local fluctuations; reference lines (green = good ≥ .80, orange = moderate ≥ .60, red = weak < .40).
Pairwise Comparisons
Box-and-whisker plots of the pairwise similarity (Hamming) or mean absolute difference (MAD) per stimulus for each rater pair.
KDE density plots of the pairwise distributions (colored lines) and all raters combined (bold, accent color).
Kendall's W — Rank Concordance
Measures the concordance of all raters simultaneously based on ranks. W = 0: no agreement; W = 1: all raters rank identically.
Computed per dimension; χ² test checks whether W > 0 (df = n_stimuli − 1).
ICC — Intraclass Correlation
ICC(x,1): reliability of a single measurement (1 rater per person).
ICC(x,k): reliability of the mean of all k raters (Spearman-Brown corrected).

Reference values (Koo & Mae 2016):
< .50 poor · .50–.74 moderate · .75–.89 good · ≥ .90 excellent
Heatmap — Dimension × Rater
Heatmap
Shows the pairwise agreement value (chosen method) for each rater pair (row) and each dimension (column).
Red = low agreement; green = high agreement.
Red cells show on which dimensions two raters agree especially poorly; green cells signal good pairwise agreement.
Bland-Altman Plot
Shows for each rater pair: x-axis = mean of both values, y-axis = difference.
Bias = mean difference. LoA = bias ± 1.96 SD.
Points outside the LoA = systematic discrepancy between the rater pair.