The Multitrait-Multimethod Matrix (MTMM, Campbell & Fiske, 1959) measures several traits (constructs) simultaneously with several methods (informants/procedures) and sorts the correlations between all resulting measurements into four cell types. Tab ① shows the matrix itself with color-coded cell types and an automatic criteria check; tab ② shows that the same logic can be written as an SEM with trait and method factors.
The example
Three traits — depression, anxiety, competence — are assessed with three methods: self-report, peer report (e.g. classmates), parent report. Every trait-method combination yields one measurement, giving 9 variables and a 9×9 correlation matrix.
The four cell types
Reliability diagonal (gray): the correlation of a measurement with itself across repeated measurement — shown here as a reliability coefficient h², not as a trivial 1.0.
Validity diagonal / Monotrait-Heteromethod (MTHM) (green): same trait, different method — e.g. self-reported depression × peer-reported depression. These are the convergent validity coefficients: they should be clearly different from 0 and as large as possible.
Heterotrait-Monomethod (HTMM) (orange): different traits, same method — e.g. depression × anxiety, both self-reported. These correlations are distorted by shared method variance (e.g. response tendencies, halo effects) and should be lower than the validity diagonal.
Heterotrait-Heteromethod (HTHM) (blue): different traits, different methods — the methodologically "cleanest" comparison, undistorted by shared method variance.
Campbell & Fiske criteria for convergent & discriminant validity
① The validity diagonal (MTHM) should be significantly different from 0 and substantial — convergent validity. ② An MTHM value should be higher than the HTHM values in the same row/column. ③ An MTHM value should be higher than the HTMM values involving the same trait — this is the strictest criterion, because method variance systematically inflates correlations. ② and ③ together are discriminant validity. ④ The correlation pattern between traits should be similar across all methods (here already satisfied by the model's fixed trait correlations).
The model behind tab ②
X_tm = λ_T · Trait_t + λ_M · Method_m + e_tm
Every measurement X is explained by its trait factor, its method factor, and measurement error. This yields closed-form formulas for all four cell types (ρ = trait correlation):
This structurally implies always HTMM ≥ HTHM (method variance can only add to a correlation) — exactly Campbell & Fiske's observation that monomethod correlations are usually inflated. Whether MTHM stays larger than HTMM depends on the ratio of the two loadings: λT (trait loading) and λM (method loading) — readable directly off the paths in tab ② for each scenario.
Limits of this CT-UM model
This tool assumes uncorrelated method factors — that is not the "simplest variant" of the Correlated-Trait-Correlated-Method model (CT-CM), but a distinctly named model in its own right: the Correlated-Trait-Uncorrelated-Methods model (CT-UM), a special case of the CT-CM model (Eid et al., 2008). Allowing correlations between the method factors instead (true CT-CM) frequently leads, in practice, to improper solutions (e.g. negative error variances) and identification problems. A common way out is the CT-C(M−1) model (one method factor is dropped as a reference) — details, diagrams, and which model fits when: tab ③.
Literature
Campbell, D. T. & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105. Geiser, C., Eid, M., Nussbeck, F. W., Lischetzke, T. & Cole, D. A. (2010). Multitrait-Multimethod-Analyse. In H. Holling & B. Schmitz (Eds.), Handbuch Statistik, Methoden und Evaluation (pp. 679–685). Hogrefe. Widaman, K. F. (1985). Hierarchically nested covariance structure models for multitrait-multimethod data. Applied Psychological Measurement, 9(1), 1–26. Kenny, D. A. (1976). An empirical application of confirmatory factor analysis to the multitrait-multimethod matrix. Journal of Experimental Social Psychology, 12(3), 247–252. Jöreskog, K. G. (1971). Statistical analysis of sets of congeneric tests. Psychometrika, 36(2), 109–133. Eid, M. (2000). A multitrait-multimethod model with minimal assumptions. Psychometrika, 65(2), 241–261. Eid, M., Nussbeck, F. W., Geiser, C., Cole, D. A., Gollwitzer, M. & Lischetzke, T. (2008). Structural equation modeling of multitrait-multimethod data: Different models for different types of methods. Psychological Methods, 13(3), 230–253.
Scenarios
Cell types
Reliability (diagonal)
MTHM — convergent
HTMM — method effect
HTHM — clean
📋 Example
Depression, anxiety, and competence are assessed in the same individuals via self-report, peer report, and parent report — 3 traits × 3 methods = 9 measurements. The 9×9 correlation matrix answers two questions at once: Do different methods agree on the same trait (convergent validity)? And can the traits be told apart from each other, even when measured by the same method (discriminant validity)?
② Convergent & discriminant validity — criteria check
—
Trait
min. |MTHM|
max. |HTMM|
max. |HTHM|
Convergent
Discriminant
Convergent: the trait's smallest |MTHM| value is clearly > 0. Discriminant: that same value exceeds both the strongest HTMM and the strongest HTHM correlation involving this trait.
Flashcards
Why HTMM ≥ HTHM holds structurally
In the model, HTMM = HTHM + λM². Method variance can only add to a correlation (when λM > 0), never lower it. That's why correlations between different traits measured with the same method are almost never smaller than the same traits measured across different methods — one of the most robust empirical MTMM findings.
The strictest criterion
Criterion ③ (MTHM > HTMM) fails most often in real data — including in Geiser et al.'s (2010) example: the self-peer and self-parent correlations for depression (.20–.30) are clearly smaller than the within-self-report correlation between depression and anxiety (.67). Discriminant validity is thus not cleanly established there. Yet the CT-UM model in tab ② still fits this very data reasonably well overall — no contradiction, just two different questions (details there).
Where does method variance come from?
Shared response tendencies (e.g. socially desirable responding), halo effects in raters, a shared observation context (parents see their child in the same situations), or item-wording style — all causes that shift all traits of the same method up or down together, regardless of their actual trait content.
What's next?
Tab ② shows that the same logic is just an ordinary structural equation model — just with two kinds of latent factors (trait and method) instead of one. → Switch to tab ②
📋 From matrix to model
The same 9 measurements, the same example — drawn by default as a Correlated-Trait-Uncorrelated-Methods model (CT-UM): 3 trait factors, 3 method factors, each measurement loads on exactly one trait and exactly one method; the method factors do not correlate with each other by default (via the toggle below, also as CT-CM with correlated method factors — details and further model variants in tab ③). The loadings here are not freely adjustable, but read directly from the matrix in tab ① for each scenario — the scenario selection on the left thus drives both at once. For the synthetic scenarios B–D that's a single λ_T/λ_M for all 9 items (the matrix was generated that way in the first place); for scenario A (real data) every item gets its own trait and method loading, estimated via maximum likelihood.
③ Path diagram — trait and method factors
CT-UM ↔ CT-CM (scenario A only)
—
Method factors are drawn uncorrelated here by default (CT-UM model) — the button above lets you (for scenario A only) draw them correlated instead, with real ML values (CT-CM). Details, limits, and further model variants like CT-C(M−1) are in tab ③ and the help. Trait factors correlate according to the values shown above in the arcs.
Tab ① (table) and tab ② (path diagram) are two views of the same model. For the synthetic scenarios B–D, the path diagram's implied matrix is exactly the matrix from tab ① — both were generated from exactly these loadings, after all. For scenario A (real data), the implied matrix instead deviates from the real example — the model fit above shows by how much.
Why two factors per item?
In an ordinary CFA, each item loads on only one factor. In the CT-UM model (as here) and the more general CT-CM model, each item loads on two: its trait (what it's actually supposed to measure) and its method (how it was measured). That's exactly the separation of "true" construct variance from method artifacts — algebraically the same split that the four MTMM cell types in tab ① make visible.
Model fit ≠ validity
A good model fit (SRMR above) only says: the form of the model — correlated traits, uncorrelated methods, a simple loading structure — matches the observed correlation structure. It says nothing about whether the estimated values themselves are good. In the Geiser example (scenario A), the model fits well — and that's exactly why we can trust the estimated parameters. And those reveal precisely the problem that tab ① flags as a violation of criterion ③: a trait correlation of ρ=0.90 (depression–anxiety), plus a Heywood case at x4. A good fit makes a poor validity diagnosis more credible, not less — a poorly fitting model couldn't have earned trust in these numbers in the first place. Fit and validity are independent questions: you can have a model that matches the data's structure exactly and still exposes poorly measured constructs — or, conversely, a poorly fitting model whose (unreliable) parameters happen to look good.
📋 One model family, four variants
Tab ② shows only one way to formalize trait and method effects as an SEM (CT-UM). That's not an arbitrary choice, but one of four classic CFA-MTMM models, which differ only in which correlations between error/method factors are allowed — and how "difficult" method factors are handled. No model here is populated with values from the Geiser example (that's tabs ①/②) — this is purely about structure.
Kenny, 1976
A — Correlated-Trait-Correlated-Uniqueness (CT-CU)
Each item loads only on one trait factor. Method effects are not modeled as a separate factor, but via correlated error variables ("uniquenesses") between items of the same method. Drawback: method effects are confounded with measurement error (reliability is underestimated), and with many variables, many error correlations must be estimated — not parsimonious. No decomposition into trait, method, and error variance is possible.
Jöreskog, 1971
B — Correlated-Trait-Correlated-Method (CT-CM)
Each item loads on one trait factor and one method factor. All trait factors may correlate, and all method factors may also correlate (only trait↔method is disallowed). Advantage over CT-CU: method effects are modeled explicitly, and a decomposition into trait/method/error variance becomes possible. Drawback: in practice, this frequently leads to improper solutions (e.g. negative error variances) and identification problems — especially when the method factors correlate strongly. Available via the toggle in tab ② for the Geiser example — see it live there: the Heywood case at x4 becomes even more extreme under CT-CM than under CT-UM.
Special case of CT-CM
C — Correlated-Trait-Uncorrelated-Methods (CT-UM)
The default model in tab ② of this tool. Like CT-CM, but the method factors are not allowed to correlate with each other. Avoids most of CT-CM's estimation problems. Especially suited to interchangeable methods — e.g. raters randomly drawn from a larger population (not which raters, but that they form a representative sample). For structurally different methods (self-, parent-report, a physiological measure — as in the Geiser example), the uncorrelatedness assumption is a simplification, not an exact picture of reality.
Eid, 2000; Eid et al., 2008
D — Correlated-Trait-Correlated-(Methods-Minus-One) [CT-C(M−1)]
One method is chosen as the reference (here M1) — its items load only on their trait and get no method factor of their own. For every other method, a method factor is specified representing the deviation from the reference method; these method factors may correlate with each other. Avoids CT-CM's problems without needing CT-UM's (often unrealistic) uncorrelatedness assumption — the current standard in the MTMM-SEM literature whenever a sensible reference method exists (e.g. the most established or most objective method).
What if I have multiple items per trait-method combination?
All four models above assume exactly one indicator per trait-method unit (as in the Geiser example: 3 traits × 3 methods = 9 single items) — this implicitly assumes that the method effect is equally strong for every trait. With multiple indicators per combination (e.g. several items per trait and method), trait-specific method effects can instead be modeled — the method effect is then allowed to vary in strength from trait to trait. Such extensions are described in detail in Eid, Nussbeck, Geiser, Cole, Gollwitzer & Lischetzke (2008), but not implemented here.