Running Example
Stimuli (n)80 patient videos from an outpatient clinic — each video shows an intake session
Raters (k)3 clinical psychologists — rate each video independently
Dimension 1depression severity (PHQ category: none / mild / moderate / severe)
Dimension 2risk classification (0 = no risk … 4 = acute)
Format3 CSV files (80 rows × 2 columns each) — one per rater
Upload rater files, choose scale level and method,
then click Compute.
then click Compute.
① Global Agreement
② Agreement by Dimension
③ Per-Stimulus Agreement
Overall Measure — All Raters
Box-and-whisker plot: distribution of the agreement measure across all stimuli. Diamond = mean, line = median.
Histogram of the stimulus-specific agreement values with an overlaid normal distribution (blue) and mean line.
Trace plot: agreement value per stimulus in dataset order. LOESS trend line shows local fluctuations; reference lines (green = good ≥ .80, orange = moderate ≥ .60, red = weak < .40).
Pairwise Comparisons
Box-and-whisker plots of the pairwise similarity (Hamming) or mean absolute difference (MAD) per stimulus for each rater pair.
KDE density plots of the pairwise distributions (colored lines) and all raters combined (bold, accent color).
Kendall's W — Rank Concordance
Measures the concordance of all raters simultaneously based on ranks.
W = 0: no agreement; W = 1: all raters rank identically.
Computed per dimension; χ² test checks whether W > 0 (df = n_stimuli − 1).
Computed per dimension; χ² test checks whether W > 0 (df = n_stimuli − 1).
ICC — Intraclass Correlation
ICC(x,1): reliability of a single measurement (1 rater per person).
ICC(x,k): reliability of the mean of all k raters (Spearman-Brown corrected).
Reference values (Koo & Mae 2016):
< .50 poor · .50–.74 moderate · .75–.89 good · ≥ .90 excellent
ICC(x,k): reliability of the mean of all k raters (Spearman-Brown corrected).
Reference values (Koo & Mae 2016):
< .50 poor · .50–.74 moderate · .75–.89 good · ≥ .90 excellent
Heatmap — Dimension × Rater
Heatmap
Shows the pairwise agreement value (chosen method) for each rater pair (row) and each dimension (column).Red = low agreement; green = high agreement.
Red cells show on which dimensions two raters agree especially poorly; green cells signal good pairwise agreement.
Bland-Altman Plot
Shows for each rater pair: x-axis = mean of both values, y-axis = difference.
Bias = mean difference. LoA = bias ± 1.96 SD.
Points outside the LoA = systematic discrepancy between the rater pair.
Bias = mean difference. LoA = bias ± 1.96 SD.
Points outside the LoA = systematic discrepancy between the rater pair.