Methods Lab beta 0.7

Dr. Rainer Düsing · University of Osnabrück DE
Interactive Statistics · Open Source · University of Osnabrück

Methods Lab

Interactive visualizations for statistical methods — from simple regression through causal inference to psychometric models. No code required, every concept is directly experienceable.

Dr. Rainer Düsing · University of Osnabrück · Department of Research Methods, Diagnostics & Evaluation · ResearchGate ↗
Version beta 0.7 48 tools available 2 tools planned HTML / JS only No server
📖 Glossary: Key terms used across this lab are explained in the shared Bayes Thinking Lab Glossary ↗ — a curated reference for both labs.
📚 References: The literature cited across the tools is collected and categorized in the shared Bayes Thinking Lab Reference List ↗.
01

Regression & Association

The foundation for every other section: how do you estimate linear relationships, what does "controlling for" actually mean, and how do you decompose effects into direct and indirect paths? From simple OLS regression through mediation and moderation to hierarchically nested data — these tools form the basis for understanding the statistical paradoxes in Section 3 and the causal methods in Section 4.

02

Inference & Planning

What does a p-value actually tell you? How large does a sample need to be? What counts as a meaningful effect? And what happens when data are missing? These questions determine the quality of every empirical study — before data collection and after it.

03

Statistical Paradoxes

Counter-intuitive effects that trip up even experienced researchers. All the paradoxes here rest on regression logic — which is why this section comes after the fundamentals. Once you understand Section 1, you'll see why these results are not surprising after all.

↩️
Regression to the Mean
Blood-pressure example, r slider, bidirectional selection. Why "bad" extreme scores look better the second time around.
📡
Measurement Error & Attenuation
Intelligence & school achievement — disattenuation, flashcards on N-instability and reliability choice. How measurement imprecision distorts correlations.
🌀
Berkson's Paradox
Talent + effort → success — DAG, diagonal selection boundary, collider bias. Negative correlation in selected samples.
🔀
Lord's Paradox
ANCOVA vs. difference score — when do the two analyses lead to opposite conclusions? DAG, key formula, best-practice decision.
📊
Simpson's Paradox
Aggregated trend reverses on disaggregation — within- vs. pooled regression, confounding slider, amplification & reversal. Cross-link to the Multilevel tool.
🔭
Range Restriction
Variance restriction interactively — how selection on X attenuates the correlation. Scatter full vs. restricted, spread bars, Thorndike relation r(u), Case II correction. Slope stays, r drops.
⚔️
Lindley's Paradox
p < .05, yet the Bayes factor supports H₀ — the conflict between frequentist and Bayesian conclusions. Lindley mode, z-distribution H₀ vs. H₁, divergence over n. Synergy with the Bayes Thinking Lab.
🎭
Will Rogers Phenomenon
When the Okies left Oklahoma for California, the average IQ rose in both states. Reshuffling raises both group means while the overall mean stays the same — animated. Also: stage migration.
🔦
Data Dredging & the Streetlight Paradox
Search a large dataset for significance and you will inevitably find chance hits. Multiple comparisons, P(≥1)=1−(1−α)ᵐ, Bonferroni, HARKing headlines + stepwise regression (Freedman).
🚌
Inspection Paradox
You tend to land in above-average-length intervals by chance (buses, clinical cohorts, survival data). Length-biased sampling: experienced gap = E[L²]/E[L] ≥ E[L].
✈️
Survivorship Bias
Wald's bombers: where to add armor? Click zones, then reveal — the missing (shot-down) planes show where it actually matters. Selecting on the outcome biases every conclusion.
🗺️
Law of Small Numbers
The "record-holding" counties (highest and lowest rates) are systematically the smallest ones — pure sampling variability. Funnel plot, ranking, SE ∝ 1/√n. Related to regression to the mean.
🕸️
Friendship Paradox
On average, your friends have more friends than you do — size-biased sampling over edges. Click people, make a prediction, reveal. Related to the inspection paradox.
04

Causal Inference

Under what conditions is it justified to infer causation from association? This section covers the potential-outcomes framework, natural experiments and matching methods — the methodological toolkit of modern causal analysis.

05

Test Theory & Measurement

How do you measure psychological constructs, and how well does a test do it? From classical reliability to modern IRT models — this section covers the fundamentals of psychometrics.

📋
Classical Test Theory — Basics
True-score model X = T + E, reliability as a proportion of variance, SEM, confidence band around test scores, Spearman-Brown test length. The anchor for the reliability topic.
🪞
Measurement Models (CFA Intro)
Parallel, tau-equivalent, congeneric — path diagram, implied covariance matrix, automatic model classification, ω vs. α live. Bridge from EFA to reliability.
🧩
Factor Analysis
PAF (genuine EFA), correct oblique rotation (pattern matrix Λ), reliability panel (α / ωt / ωh Schmid-Leiman), 3 scenarios, biplot.
🔬
Confirmatory Factor Analysis
2-factor CFA with genuine ML fitting, misspecification switches (force orthogonal, ignore cross-loading), χ²/CFI/TLI/RMSEA/SRMR live, lavaan syntax box. Theory fixes the structure in advance.
🧭
Structural Equation Model
Directed structural paths between latent factors, optionally as a mediation model (X→M→Y). Live comparison of latent vs. manifest (sum-score) path directly shows attenuation from measurement error.
🕸️
Multitrait-Multimethod Matrix
Self-, peer-, and parent-report × three traits — the MTMM matrix decomposed by cell type (reliability, MTHM, HTMM, HTHM), automatic Campbell-Fiske criteria check, the same approach as a CT-UM SEM (toggle: CT-CM), plus an overview of all four CFA-MTMM model variants.
🎯
Reliability: α vs. ω
Cronbach's α vs. McDonald's ω_t/ω_h — when α underestimates (congenericity) and when it fakes unidimensionality (multidimensionality). Variance decomposition, split-half distribution.
📐
IRT — Dichotomous Models
1PL / 2PL / 3PL / 4PL + Rasch — ICC, TIF, Wright map, MLE person estimator. Rasch vs. 1PL difference made explicit.
🎚️
IRT — Ordinal Models
PCM, GPCM, GRM — CRF, ESC, item information for all items. Disordered-threshold warning, factor-analysis connection in help.
🔍
Differential Item Functioning
2PL model, 4 items, Δb/Δa sliders, ICC comparison, difference curve, group distributions, Raju SA/UA, ETS A/B/C.
🔄
Measurement Invariance
Configural / metric / scalar — the CFA twin of the DIF tool. Set a true difference and non-invariance and see how much of a group difference is real vs. a measurement artifact.
06

Diagnostics & Test Quality

How well does an instrument detect what it is supposed to detect? Clinical and psychological diagnostics need precise indicators — from sensitivity and specificity through inter-rater agreement to differential validity.

🩺
Sensitivity & Specificity
ROC curve, AUC, PPV/NPV as a function of prevalence — interactive cutoff, live 2×2 table.
Diagnostic Validity
Construct, criterion and content validity — validity coefficients, Taylor-Russell tables, utility analysis.
🏆
Taylor-Russell Tables
Selection utility of a test — success rate (PPV) from validity, selection ratio and base rate. Interactive table + nomogram with crosshair. Bivariate normal distribution, cross-link to Range Restriction.
⚖️
Test Bias
Cleary model, Meade & Fetzer, adverse impact — 3 modules, 4 scenarios. Differential prediction vs. fairness in practice.
📶
Diagnostic Intervals
Single-case diagnostics — the true score τ from a test score: confidence interval via equivalence vs. regression hypothesis and a Bayesian credible interval. SEM, regression to the mean, direct probability statements P(τ>T).
🩹
Jacobson-Truax Analysis
Reliable Change Index & clinical significance — RCI band, cutoff a/b/c, 4-cell classification in the pre-post plot. BDI-II as an example.
📉
Single Case Designs
Baseline as prediction, AB/ABA/ABAB reversal designs, multiple-baseline design with a confound check, visual criteria plus live-computed PND/NAP/Tau-U.
🌡️
Norming Bias
T-scores under right-skewed distributions (SCL-90 analogy) — gamma, log-normal, exponential, ex-Gaussian. Empirical PR vs. T-norm PR, discrepancy table, interactive parameters.
👤
Profile Analysis
Profile comparison with Cattell's rₚ, McCrae's Iₚₐ / rₚₐ, ICC_de — elevation, scatter and shape of a profile assessed separately.
👥
ICC Lab
Intraclass correlation — all 6 Shrout-&-Fleiss forms, number of raters, absolute vs. consistency agreement.
Tool only
🤝
Inter-Rater Agreement
Cohen's κ, weighted κ, the kappa paradox, prevalence & bias — when κ misleads and what to report instead.
🎲
Sister project · Bayesian statistics
Bayes Thinking Lab

Frequentist methods are covered here — but for priors, posterior distributions, ROPE decisions, Bayes factors and brms models, there's the Bayes Thinking Lab: interactive tools that explain Bayesian thinking from the ground up. The two labs complement each other: Lindley's Paradox, for example, needs both perspectives.

Disclaimer

Methods Lab is a free, open-source teaching and learning project. All content is provided for educational purposes only. Despite careful preparation of the content and implementations, no guarantee is given for the correctness, completeness or currency of the calculations, visualizations or other content shown. Use is at your own risk.

Liability for damages arising from the use of, or reliance on, the information provided — in particular from incorrect calculations due to software bugs — is expressly excluded to the extent permitted by law.

For academic, clinical or other professional decisions, independent verification by qualified professionals is always required.

About the project

Methods Lab is developed and maintained by Dr. Rainer Düsing, University of Osnabrück, Department of Research Methods, Diagnostics & Evaluation. It is the sister project of the Bayes Thinking Lab.