Lawrence T. DeCarlo’s recent article introduces a psychological framework for mixed-format exams, combining signal detection theory (SDT) for multiple-choice items and item response theory (IRT) for open-ended items. This fusion allows for a unified model that captures the nuances of each item type while providing insights into the underlying cognitive processes of examinees.
Background
Mixed-format exams, commonly used in large-scale assessments, present a challenge for researchers seeking to model responses across different item types. Historically, multiple-choice items have been analyzed using frameworks like signal detection theory, while open-ended items are typically modeled using item response theory. DeCarlo’s work builds on these approaches, introducing a method to unify them through the probability of knowing, a concept that bridges both models.
Figure 1. DeCarlo's fused model routes multiple-choice items through SDT and open-ended items through IRT, unifying them via the probability of knowing.
Key Insights
Figure 2. Fitting the SDT and IRT models together lets a single analysis compare item-type processes and examine shared covariates.
Unified Framework: The article demonstrates how the SDT choice model and IRT sequential logit model can be integrated into a single framework. This approach captures latent states such as “know” and “do not know” to analyze responses across item types.
Psychological Processes: By modeling both item types simultaneously, the approach highlights differences in the cognitive processes involved in multiple-choice and open-ended responses. This sheds light on how examinees interact with each type of item.
Estimation Benefits: Fitting the SDT and IRT models together offers potential computational advantages and allows for the examination of shared covariates, improving the overall utility of the framework.
Significance
This fusion of SDT and IRT models represents a significant step forward in psychometric analysis. By addressing the differences and connections between item types, the framework provides a deeper understanding of examinee behavior. This has implications for designing fairer and more reliable assessments, particularly in international exams where mixed-format tests are prevalent.
Future Directions
Future research could focus on expanding the application of this model to other testing contexts, including formative assessments or specialized exams. Additionally, exploring how this framework performs with diverse populations and item designs could further validate its effectiveness and versatility.
Conclusion
DeCarlo’s work offers a robust framework for analyzing mixed-format exams by integrating SDT and IRT models. This unified approach not only enhances our understanding of psychological processes in test-taking but also opens the door to more comprehensive and equitable assessments.
DeCarlo, L. T. (2024). Fused SDT/IRT Models for Mixed-Format Exams. Educational and Psychological Measurement, 84(6), 1076–1106. https://doi.org/10.1177/00131644241235333
Related questions
What is item response theory?
Every time you take a standardized test — an IQ assessment, a college entrance exam, a professional certification — the questions have been calibrated using sophisticated statistical models that most test-takers never learn about. Item Response Theory (IRT) is the mathematical framework behind virtually all modern psychological and educational testing, and understanding its basics illuminates why tests work the way they do. Read more →
What is attenuation-corrected reliability?
Most psychometrics textbooks teach the classical "correction for attenuation" — Spearman's century-old technique for estimating what the correlation between two psychological constructs would be if the tests measuring them were perfectly reliable. The technique is simple: divide the observed correlation by the square root of the product of the two reliabilities. The technique is also limited: it adjusts the relationship between two scales, but assumes the reliability values plugged into the denominator are themselves accurate. A 2022 paper by Jari Metsämuuronen in Applied Psychological Measurement argues that this assumption is broken in practice. Reliability estimates produced by Cronbach's alpha and similar formulas are themselves attenuated by the same mechanical errors that attenuate correlations — and in some datasets, alpha may be deflated by 0.40–0.60 units of reliability. Metsämuuronen's contribution is a class of deflation-corrected reliability estimators that apply the classical attenuation logic inside the reliability formula rather than only to correlations between scales. Read more →
How are JCCES General Knowledge items structured?
The General Knowledge (GK) subtest of the Jouve Cerebrals Crystallized Educational Scale (JCCES) measures factual breadth — the accumulated stock of information about the world that crystallized intelligence theory treats as a core component of acquired cognitive ability. A subtest of this kind has to satisfy two structural requirements: items should span a meaningful range of difficulty, and they should order along a single underlying continuum of factual breadth rather than tapping multiple unrelated dimensions. This study examined the item structure of the JCCES GK subtest using multidimensional scaling (MDS) on response data from 588 respondents, and recovered the empirical signature of a well-ordered unidimensional construct: a horseshoe-shaped scaling pattern that is the canonical evidence of a Guttman simplex — items ordered cleanly along a single difficulty continuum. Read more →
How well does coefficient alpha perform with non-normal data?
Cronbach's coefficient alpha is the most-reported reliability statistic in psychology and educational measurement. It is also one of the most-misunderstood. The classical formula assumes that test items measure a single construct with equal factor loadings (tau-equivalence), uncorrelated errors, and continuously distributed scores. Real psychological measurement rarely meets all three assumptions: most scales use Likert responses (discrete), have items with unequal contributions to the construct (congeneric), and produce score distributions that depart from normality. The natural question is how badly alpha breaks under these violations and which alternatives perform better. A 2023 simulation study by Xiao and Hau in Educational and Psychological Measurement provides a systematic answer, with implications for the routine reliability reporting that fills psychometric methods sections. Read more →
Why do mixed-format exams pose a modeling challenge?
Mixed-format exams, commonly used in large-scale assessments, present a challenge for researchers seeking to model responses across different item types. Historically, multiple-choice items have been analyzed using frameworks like signal detection theory, while open-ended items are typically modeled using item response theory. DeCarlo’s work builds on these approaches, introducing a method to unify them through the probability of knowing, a concept that bridges both models.
What are the key insights of the unified SDT-IRT framework?
Unified Framework: The article demonstrates how the SDT choice model and IRT sequential logit model can be integrated into a single framework. This approach captures latent states such as "know" and "don’t know" to analyze responses across item types.
Psychological Processes: By modeling both item types simultaneously, the approach highlights differences