Multitrait-Multimethod Matrix — MTMM

Dr. R. Düsing · University of Osnabrück
Multitrait-Multimethod Matrix — Background
What does this tool show?

The Multitrait-Multimethod Matrix (MTMM, Campbell & Fiske, 1959) measures several traits (constructs) simultaneously with several methods (informants/procedures) and sorts the correlations between all resulting measurements into four cell types. Tab shows the matrix itself with color-coded cell types and an automatic criteria check; tab shows that the same logic can be written as an SEM with trait and method factors.

The example

Three traits — depression, anxiety, competence — are assessed with three methods: self-report, peer report (e.g. classmates), parent report. Every trait-method combination yields one measurement, giving 9 variables and a 9×9 correlation matrix.

The four cell types

Reliability diagonal (gray): the correlation of a measurement with itself across repeated measurement — shown here as a reliability coefficient h², not as a trivial 1.0.

Validity diagonal / Monotrait-Heteromethod (MTHM) (green): same trait, different method — e.g. self-reported depression × peer-reported depression. These are the convergent validity coefficients: they should be clearly different from 0 and as large as possible.

Heterotrait-Monomethod (HTMM) (orange): different traits, same method — e.g. depression × anxiety, both self-reported. These correlations are distorted by shared method variance (e.g. response tendencies, halo effects) and should be lower than the validity diagonal.

Heterotrait-Heteromethod (HTHM) (blue): different traits, different methods — the methodologically "cleanest" comparison, undistorted by shared method variance.

Campbell & Fiske criteria for convergent & discriminant validity

The validity diagonal (MTHM) should be significantly different from 0 and substantial — convergent validity. An MTHM value should be higher than the HTHM values in the same row/column. An MTHM value should be higher than the HTMM values involving the same trait — this is the strictest criterion, because method variance systematically inflates correlations. and together are discriminant validity. The correlation pattern between traits should be similar across all methods (here already satisfied by the model's fixed trait correlations).

The model behind tab
X_tm = λ_T · Trait_t + λ_M · Method_m + e_tm

Every measurement X is explained by its trait factor, its method factor, and measurement error. This yields closed-form formulas for all four cell types (ρ = trait correlation):

Reliability h² = λ_T² + λ_M² MTHM (validity) = λ_T² HTHM = λ_T,t · λ_T,t' · ρ_tt' HTMM = λ_T,t · λ_T,t' · ρ_tt' + λ_M²

This structurally implies always HTMM ≥ HTHM (method variance can only add to a correlation) — exactly Campbell & Fiske's observation that monomethod correlations are usually inflated. Whether MTHM stays larger than HTMM depends on the ratio of the two loadings: λT (trait loading) and λM (method loading) — readable directly off the paths in tab for each scenario.

Limits of this CT-UM model

This tool assumes uncorrelated method factors — that is not the "simplest variant" of the Correlated-Trait-Correlated-Method model (CT-CM), but a distinctly named model in its own right: the Correlated-Trait-Uncorrelated-Methods model (CT-UM), a special case of the CT-CM model (Eid et al., 2008). Allowing correlations between the method factors instead (true CT-CM) frequently leads, in practice, to improper solutions (e.g. negative error variances) and identification problems. A common way out is the CT-C(M−1) model (one method factor is dropped as a reference) — details, diagrams, and which model fits when: tab .

Literature

Campbell, D. T. & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.
Geiser, C., Eid, M., Nussbeck, F. W., Lischetzke, T. & Cole, D. A. (2010). Multitrait-Multimethod-Analyse. In H. Holling & B. Schmitz (Eds.), Handbuch Statistik, Methoden und Evaluation (pp. 679–685). Hogrefe.
Widaman, K. F. (1985). Hierarchically nested covariance structure models for multitrait-multimethod data. Applied Psychological Measurement, 9(1), 1–26.
Kenny, D. A. (1976). An empirical application of confirmatory factor analysis to the multitrait-multimethod matrix. Journal of Experimental Social Psychology, 12(3), 247–252.
Jöreskog, K. G. (1971). Statistical analysis of sets of congeneric tests. Psychometrika, 36(2), 109–133.
Eid, M. (2000). A multitrait-multimethod model with minimal assumptions. Psychometrika, 65(2), 241–261.
Eid, M., Nussbeck, F. W., Geiser, C., Cole, D. A., Gollwitzer, M. & Lischetzke, T. (2008). Structural equation modeling of multitrait-multimethod data: Different models for different types of methods. Psychological Methods, 13(3), 230–253.

📋 Example
Depression, anxiety, and competence are assessed in the same individuals via self-report, peer report, and parent report — 3 traits × 3 methods = 9 measurements. The 9×9 correlation matrix answers two questions at once: Do different methods agree on the same trait (convergent validity)? And can the traits be told apart from each other, even when measured by the same method (discriminant validity)?
← Structural Equation Model Compare groups instead of methods? → Measurement Invariance
The 9×9 matrix
Reliability h² (diagonal)
MTHM — same trait, different method (convergent)
HTMM — different traits, same method
HTHM — different traits, different method
Criterion violated
Convergent & discriminant validity — criteria check
Traitmin. |MTHM|max. |HTMM|max. |HTHM|ConvergentDiscriminant
Convergent: the trait's smallest |MTHM| value is clearly > 0. Discriminant: that same value exceeds both the strongest HTMM and the strongest HTHM correlation involving this trait.
Flashcards
Why HTMM ≥ HTHM holds structurally
In the model, HTMM = HTHM + λM². Method variance can only add to a correlation (when λM > 0), never lower it. That's why correlations between different traits measured with the same method are almost never smaller than the same traits measured across different methods — one of the most robust empirical MTMM findings.
The strictest criterion
Criterion (MTHM > HTMM) fails most often in real data — including in Geiser et al.'s (2010) example: the self-peer and self-parent correlations for depression (.20–.30) are clearly smaller than the within-self-report correlation between depression and anxiety (.67). Discriminant validity is thus not cleanly established there. Yet the CT-UM model in tab still fits this very data reasonably well overall — no contradiction, just two different questions (details there).
Where does method variance come from?
Shared response tendencies (e.g. socially desirable responding), halo effects in raters, a shared observation context (parents see their child in the same situations), or item-wording style — all causes that shift all traits of the same method up or down together, regardless of their actual trait content.
What's next?
Tab shows that the same logic is just an ordinary structural equation model — just with two kinds of latent factors (trait and method) instead of one. → Switch to tab
📋 From matrix to model
The same 9 measurements, the same example — drawn by default as a Correlated-Trait-Uncorrelated-Methods model (CT-UM): 3 trait factors, 3 method factors, each measurement loads on exactly one trait and exactly one method; the method factors do not correlate with each other by default (via the toggle below, also as CT-CM with correlated method factors — details and further model variants in tab ). The loadings here are not freely adjustable, but read directly from the matrix in tab for each scenario — the scenario selection on the left thus drives both at once. For the synthetic scenarios B–D that's a single λ_T/λ_M for all 9 items (the matrix was generated that way in the first place); for scenario A (real data) every item gets its own trait and method loading, estimated via maximum likelihood.
Path diagram — trait and method factors
CT-UM ↔ CT-CM (scenario A only)
Method factors are drawn uncorrelated here by default (CT-UM model) — the button above lets you (for scenario A only) draw them correlated instead, with real ML values (CT-CM). Details, limits, and further model variants like CT-C(M−1) are in tab and the help. Trait factors correlate according to the values shown above in the arcs.
x1=Self×Depr. · x2=Self×Anx. · x3=Self×Comp. · x4=Peer×Depr. · x5=Peer×Anx. · x6=Peer×Comp. · x7=Parent×Depr. · x8=Parent×Anx. · x9=Parent×Comp.
What does this scenario show?
Flashcards
Two representations, one model
Tab (table) and tab (path diagram) are two views of the same model. For the synthetic scenarios B–D, the path diagram's implied matrix is exactly the matrix from tab — both were generated from exactly these loadings, after all. For scenario A (real data), the implied matrix instead deviates from the real example — the model fit above shows by how much.
Why two factors per item?
In an ordinary CFA, each item loads on only one factor. In the CT-UM model (as here) and the more general CT-CM model, each item loads on two: its trait (what it's actually supposed to measure) and its method (how it was measured). That's exactly the separation of "true" construct variance from method artifacts — algebraically the same split that the four MTMM cell types in tab make visible.
Model fit ≠ validity
A good model fit (SRMR above) only says: the form of the model — correlated traits, uncorrelated methods, a simple loading structure — matches the observed correlation structure. It says nothing about whether the estimated values themselves are good. In the Geiser example (scenario A), the model fits well — and that's exactly why we can trust the estimated parameters. And those reveal precisely the problem that tab flags as a violation of criterion : a trait correlation of ρ=0.90 (depression–anxiety), plus a Heywood case at x4. A good fit makes a poor validity diagnosis more credible, not less — a poorly fitting model couldn't have earned trust in these numbers in the first place. Fit and validity are independent questions: you can have a model that matches the data's structure exactly and still exposes poorly measured constructs — or, conversely, a poorly fitting model whose (unreliable) parameters happen to look good.
📋 One model family, four variants
Tab shows only one way to formalize trait and method effects as an SEM (CT-UM). That's not an arbitrary choice, but one of four classic CFA-MTMM models, which differ only in which correlations between error/method factors are allowed — and how "difficult" method factors are handled. No model here is populated with values from the Geiser example (that's tabs /) — this is purely about structure.
Kenny, 1976
A — Correlated-Trait-Correlated-Uniqueness (CT-CU)
Each item loads only on one trait factor. Method effects are not modeled as a separate factor, but via correlated error variables ("uniquenesses") between items of the same method. Drawback: method effects are confounded with measurement error (reliability is underestimated), and with many variables, many error correlations must be estimated — not parsimonious. No decomposition into trait, method, and error variance is possible.
Jöreskog, 1971
B — Correlated-Trait-Correlated-Method (CT-CM)
Each item loads on one trait factor and one method factor. All trait factors may correlate, and all method factors may also correlate (only trait↔method is disallowed). Advantage over CT-CU: method effects are modeled explicitly, and a decomposition into trait/method/error variance becomes possible. Drawback: in practice, this frequently leads to improper solutions (e.g. negative error variances) and identification problems — especially when the method factors correlate strongly. Available via the toggle in tab for the Geiser example — see it live there: the Heywood case at x4 becomes even more extreme under CT-CM than under CT-UM.
Special case of CT-CM
C — Correlated-Trait-Uncorrelated-Methods (CT-UM)
The default model in tab of this tool. Like CT-CM, but the method factors are not allowed to correlate with each other. Avoids most of CT-CM's estimation problems. Especially suited to interchangeable methods — e.g. raters randomly drawn from a larger population (not which raters, but that they form a representative sample). For structurally different methods (self-, parent-report, a physiological measure — as in the Geiser example), the uncorrelatedness assumption is a simplification, not an exact picture of reality.
Eid, 2000; Eid et al., 2008
D — Correlated-Trait-Correlated-(Methods-Minus-One) [CT-C(M−1)]
One method is chosen as the reference (here M1) — its items load only on their trait and get no method factor of their own. For every other method, a method factor is specified representing the deviation from the reference method; these method factors may correlate with each other. Avoids CT-CM's problems without needing CT-UM's (often unrealistic) uncorrelatedness assumption — the current standard in the MTMM-SEM literature whenever a sensible reference method exists (e.g. the most established or most objective method).
What if I have multiple items per trait-method combination?
All four models above assume exactly one indicator per trait-method unit (as in the Geiser example: 3 traits × 3 methods = 9 single items) — this implicitly assumes that the method effect is equally strong for every trait. With multiple indicators per combination (e.g. several items per trait and method), trait-specific method effects can instead be modeled — the method effect is then allowed to vary in strength from trait to trait. Such extensions are described in detail in Eid, Nussbeck, Geiser, Cole, Gollwitzer & Lischetzke (2008), but not implemented here.