Classical IRT models dichotomous responses (correct/incorrect). Ordinal IRT extends this framework to polytomous items with ordered response categories (e.g. Likert scales, partial credit). As with binary IRT, person ability θ and item parameters lie on the same scale — but now with several threshold parameters per item.
The threshold parameter δⱼ is the θ value at which the adjacent categories j−1 and j are equally probable. As with the binary Rasch model: if an item doesn't fit, it is revised — not the model.
The boundary parameters bₖ must be strictly monotonic (b₁ < b₂ < ... < bₖ₋₁), otherwise negative response probabilities result. This is a structural difference from PCM/GPCM, where disordered thresholds can occur.
PCM/GPCM (δⱼ): at δⱼ, categories j−1 and j are equally probable. The δⱼ need not be ordered. Disordered thresholds (δⱼ₊₁ < δⱼ) mean that category j is the most likely response at no θ point — it is empirically redundant. Cause: adjacent categories worded too similarly.
GRM (bₖ): at bₖ, P*(X≥k|θ=bₖ) = 0.5 — the cumulative curve crosses the 50% line. Disordered thresholds are impossible in the GRM by model structure.
The GRM and ordinal factor analysis with polychoric correlations are mathematically equivalent: the GRM parameter a corresponds to a factor loading (a = λ/√(1−λ²)). The boundary parameters bₖ correspond to the threshold parameters τₖ in polychoric FA. The person parameter θ corresponds to the factor score.
The GPCM corresponds to a graded FA model with varying factor loadings. The PCM corresponds to the Rasch ideal for ordinal data — equal loadings (a=1), measurement-theoretic rigor.
The difference lies in the estimation approach: the GRM estimates parameters directly via ML; polychoric FA first estimates pairwise correlations, then factor loadings. For large samples, both give similar results.
For the GPCM, I(θ) = a² · Var(X|θ) — higher discrimination and more response variance produce more information. Compared to dichotomous items, polytomous items can carry more information — every intermediate category contributes to measurement. But only if the categories discriminate well: disordered thresholds and weak a reduce information despite many categories.
Expected Score Curve (ESC)
E[X|θ] = Σₖ k · P(X=k|θ) shows the expected response score as a function of person ability. It runs monotonically from 0 (θ → −∞) to K−1 (θ → +∞). Where the ESC is steepest, I(θ) is highest — because that's where the item discriminates best between similar ability levels.
Which Model to Choose When?
The choice depends on the item format, the research question, and the substantive understanding of the response categories:
PCM: suitable when the categories represent graded partial performance — e.g. 0 = no solution, 1 = partial approach recognizable, 2 = complete solution. The model enforces measurement-theoretic rigor (a = 1): items that don't fit are revised or removed. Useful for item banking, computerized adaptive testing, and achievement surveys where item parameters must be stable across samples.
GPCM: suitable when items are allowed different discrimination — e.g. in heterogeneous test batteries with mixed item formats. The GPCM is the standard model in large-scale educational assessments (e.g. PISA, TIMSS) for open-ended tasks with partial credit. Unlike the PCM, no measurement ideal is postulated: a is a free parameter estimated from the data.
GRM: suitable when the categories represent continuous intensity levels or degrees of agreement — e.g. Likert scales in personality or attitude questionnaires. The model is mathematically closely related to confirmatory FA via polychoric correlations. The boundary response curves are substantively easy to interpret: P*(X≥k|θ) is the probability of choosing at least category k.
Rating Scale Model (RSM), Andrich 1978: a special case of the PCM in which all items share the same threshold spacing. The threshold parameter for item i and transition j is τⱼ + βᵢ — where τⱼ is the shared category threshold and βᵢ the item-specific difficulty. Useful when all items use exactly the same response scale and one can assume that the psychological distances between categories are constant across items. The most parsimonious of the four approaches; often fails when items differ substantively.
Model Comparison and Fit
PCM and GPCM are nested (PCM ⊂ GPCM): a likelihood-ratio test (LRT) checks whether the additional a parameter significantly improves model fit. GPCM and GRM are not nested — here AIC and BIC are recommended as comparison criteria (lower values = better). For the PCM, Rasch-specific fit statistics are available: infit (information-weighted fit, sensitive to mid-range θ) and outfit (unweighted fit, sensitive to outliers at the θ extremes). Common rule of thumb: infit/outfit between 0.7 and 1.3 are considered acceptable.
Sample Size and Number of Categories
Sample size: for the PCM, N ≥ 200–250 is considered a minimum for stable parameter estimation. GPCM and GRM have an additional a parameter per item and need N ≥ 300–500. A pragmatic rule of thumb: at least 10 observations per parameter to be estimated. With small samples, Bayesian estimation (e.g. via Stan/brms or the TAM/mirt package in R) can help, since it uses priors to stabilize parameters.
Number of categories: more categories don't automatically mean more information. What matters is whether each category is chosen by a meaningful proportion of people. Edge categories with very low frequency (< 5% of responses) lead to unstable parameter estimates and should be merged with adjacent categories. In practice, 4–7 categories are usually optimal for Likert scales. Disordered thresholds are often a warning sign that a category is effectively not being used.
In R: the package mirt is available for all three models (Chalmers, 2012). The package TAM specializes in the Rasch family (PCM, RSM). ltm implements the GRM. For an FA perspective on ordinal data: lavaan with polychoric correlations or psych::fa() with cor="poly".
Connection to DIF — Test Fairness and Item Bias
Ordinal IRT models are the methodological basis for Differential Item Functioning (DIF) analyses. DIF is present when an item is differentially difficult for two groups (e.g. by gender, background, or language) — even after latent ability θ has been statistically controlled for. That is the decisive difference from mere impact (group differences in mean ability): DIF is a problem of the item, impact is a characteristic of the groups.
PCM, GPCM, and GRM enable group-specific item parameter estimation. With uniform DIF, difficulty b is systematically higher for one group — the CRF curves are shifted horizontally without crossing. With non-uniform DIF, discrimination a additionally differs — the curves cross, so the item is differentially fair across ability levels. IRT-based DIF analyses are explored in depth in the separate DIF tool in MethodsLab.
References
Masters, G. N. (1982). A Rasch model for partial credit scoring. Psychometrika, 47(2), 149–174. Muraki, E. (1992). A generalized partial credit model. Applied Psychological Measurement, 16(2), 159–176. Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika Monograph, 17. Andrich, D. (1978). A rating formulation for ordered response categories. Psychometrika, 43(4), 561–573. de Ayala, R. J. (2022). The Theory and Practice of Item Response Theory (2nd ed.). Guilford Press. Tutz, G. (2025). A Short Guide to Item Response Theory Models. Springer.
Choose Model
Select Item
Parameters — Item 1
a — discrimination1.20
a = 1 by axiom (PCM — not estimated)
Person Ability θ
θ0.0
📋 Example — Burnout Screening
4-item workplace well-being questionnaire, 5-point Likert scale (0 = never · 4 = always). θ represents general well-being.
✓ Why Separate Models for Likert/Partial-Credit Items?
Assumes the basic concepts θ, ICC, and discrimination from IRT — Dichotomous Models, now for items with more than two ordered response categories. The payoff: every intermediate category contributes to measurement — a 5-point item can carry substantially more information than a binary-coded one. But that only holds if the categories actually discriminate well; weak or disordered thresholds (step ③) wipe out this advantage immediately.
Approach — From Binary to Polytomous Item
①From binary to polytomous. Instead of one difficulty b, an item with K categories has K−1 threshold parameters — transitions between adjacent (PCM/GPCM) or cumulated (GRM) categories. → View category response curves
②Model choice: PCM, GPCM, or GRM? Partial-credit levels (0/partial approach/complete) → PCM. Heterogeneous discrimination in mixed item formats → GPCM. Degrees of agreement on Likert scales → GRM (closely related to ordinal factor analysis via polychoric correlations).
③Check for disordered thresholds. If the thresholds are not ordered in sequence, a category is empirically redundant — a warning sign for response levels worded too similarly. → View the warning note
④Expected score & item information. Where along the θ scale does the item contribute most to measurement? → Expected score · → Item information
⚠ Disordered thresholds: at least two threshold parameters are swapped (δⱼ > δⱼ₊₁).
An intermediate category then has the highest response probability at no θ point — it is empirically redundant.
Not a calculation error — a sign of poorly worded categories.
⚠ GRM: threshold order automatically adjusted. In the GRM the boundary parameters must be strictly increasing (b₁ < b₂ < b₃ < b₄) — otherwise negative category probabilities result. The slider values were sorted in increasing order for all calculations. The labels above the plot show the order actually used.