What you see
10 baseline sessions (Phase A, blue) — then two continuations of the same time span: the projection (dashed), which simply extends the baseline trend, and the actual data (red), which would result if the set intervention effect also acts on top of it.
Two functions of the baseline (Kazdin, 2017)
Descriptive: the baseline describes the current level of the target behavior. Predictive: it forecasts how the behavior would continue without intervention — that's the dashed line. Only the comparison "actual vs. predicted" allows a statement about the intervention's effect.
Trend and stability are decisive
Drag Baseline trend: a trend already pointing in the desired direction makes it harder to detect an additional intervention effect — the prediction is already heading there. Drag Baseline variability up: with strong noise, a small effect disappears into the "static" of baseline fluctuation.
Detectability
Bottom right expresses the difference between the actual and predicted level relative to baseline variability — the same logic as Cohen's d (→ Effect Sizes): the same intervention effect is easy to spot against a stable baseline, but barely distinguishable from random fluctuation against an unstable one.
What you see
Choose a design: AB (just one baseline, then intervention — not recommended, since any coincidentally time-matched third variable could explain the "effect"), ABA (return to baseline tests the prediction), ABAB (an additional second intervention phase replicates the effect), or BABA (starting with intervention).
The logic of reversal
Every new phase gets a dashed prediction line from the last same-type phase (e.g. A₂ from A₁, B₂ from B₁). If a new phase matches this prediction, that argues against an intervention effect; if it deviates systematically, that's evidence that the intervention controls the behavior.
The reversal switch
Turn ↺ Reversal off: the second (and any further) A phase does not fully return to the original baseline. This matches Kazdin's (2017) two scenarios: if the reversal occurs, the causal evidence is strong — but the client's behavior deliberately worsens again, which can be ethically delicate. If the reversal does not occur, the causal conclusion is weaker (alternative explanations aren't ruled out) — but that isn't necessarily a contradiction either: some behavior simply persists after successful training.
What you see
The same intervention is introduced for three students at different points in time (controlled via offset). Each person stays in baseline until it's "their turn" — there is never a need to return to the original baseline, unlike ABAB.
Why is this causally informative?
If the behavior changes only after each individual's own intervention start — staggered in time, equally strongly for all three people — then no single third variable acting at one fixed point in time can be the cause. The staggered replication is the control condition.
Test the counter-case
Activate ⚡ External event: a confound (e.g. a teacher change, school holidays) acts on all three people at the same time, regardless of each individual's intervention onset. The result is immediately visible: a jump that occurs at the exact same session for all three trajectories instead of staggered — a clear warning sign that the intervention isn't the (sole) cause, but a shared external factor is.
Four visual criteria (Kazdin, 2017)
① Mean: Do the phase means differ? ② Trend: Does the rise/fall of the data change from A to B? ③ Level: Is there a "break" right at the phase transition (last A value vs. first B value), independent of the mean? ④ Latency: How quickly after the intervention starts does the change appear? Click the four buttons above to mark each criterion individually on the graph.
Why not just the mean?
An example: drag Trend in Phase B to a high value with intervention effect = 0. The means of A and B can stay nearly identical — yet the behavior visibly changes within B. Pure mean comparisons (as in the classic group comparison) can miss that; that's why a single index isn't enough for SCD.
Beyond visual inspection: PND & Tau-U
PND (Percentage of Non-overlapping Data): the share of B values that exceed the most extreme A value in the desired direction — simple, but wasteful with information, using only a single baseline point.
Tau-U (Parker, Vannest, Davis & Sauber, 2011): compares all A–B pairs (related to Kendall's Tau and Non-overlap of All Pairs, NAP) and thus uses the entire baseline, not just its extreme. Values near ±1 mean (almost) complete non-overlap in the respective direction, values near 0 mean strong overlap.
PND compares B only to the most extreme A value — simple, but wasteful with
information. NAP/Tau-U (Parker et al., 2011) use all A–B pairs and are
today's more common indices for supplementing visual inspection — never meant to replace it.
Randomized group experiments are considered the gold standard but need many cases. Single Case (Experimental Research) Designs (SCD) allow causal conclusions about interventions from a single case (an individual, but also a class, school, company, or city). They are especially appropriate when an intervention is new, resources are limited, or the question is specifically about the change in this one unit.
AB — one baseline, one intervention. Not recommended: no control for third variables.
ABA / ABAB (reversal design) — baseline, intervention, baseline again (, intervention again). Each repetition tests the previous phase's prediction; a reversal of the effect when the intervention is withdrawn is strong causal evidence.
Multiple-baseline design — the same intervention is introduced in a staggered fashion across several people, behaviors, or settings. If the behavior changes only after each individual's own intervention onset, a shared external cause is unlikely — with no return to baseline at all.
Primarily through visual inspection: change in mean, in trend, in level (break at the phase transition), and the latency of the change. Quantitative indices have become established as a supplement:
These indices don't replace visual inspection — they supplement it with a more objective, reproducible number.
The most frequently cited concern is external validity: results from one or a few cases don't straightforwardly generalize to other people. SCD are precisely suited to situations where the problem is rare enough that recruiting a larger, similar sample isn't practical anyway.