Four properties, four different questions
The COSMIN taxonomy separates reliability, measurement error, validity, and responsiveness because evidence for one property cannot stand in for another. [1]
- Reliability: are differences between people distinguishable from random measurement error?
- Measurement error: how much uncertainty in a score is not attributed to real change in the construct?
- Validity: does the available evidence support the proposed interpretation for the intended use?
- Responsiveness: can the measure detect change over time in the construct it intends to measure?
Reliability is not a permanent badge
Reliability estimates depend on the population, score distribution, administration conditions, and source of variation considered, so a coefficient should be read with its study context and confidence interval. [2]
- Internal consistency concerns relationships among items under a reflective measurement model.
- Test-retest reliability concerns stability when the construct is expected not to have changed.
- Inter-rater or intra-rater reliability matters when judgement or observation contributes to the score.
The four-lens evidence check
Four parallel questions that keep distinct measurement properties from being used as proxies for one another.
- Reliability
Can relevant score differences be distinguished despite random error?
- Error
How much score uncertainty is not attributed to real change?
- Validity
What evidence supports the intended interpretation in this context?
- Responsiveness
Can the measure detect relevant change over the intended interval?
Measurement error sets limits on score-level claims
Standard error of measurement, limits of agreement, and smallest detectable change are ways of expressing uncertainty or detectable difference, but they are not the same as a change that people or clinicians consider important. [1][2]
- Keep observed scores and uncertainty conceptually separate.
- Do not treat every numerical difference as evidence of real change.
- Distinguish detectable change from meaningful or important change.
- Check that error estimates apply to the relevant population and administration mode.
Validity is an evidence-based argument
Content validity asks whether the measure's content is relevant, comprehensive, and comprehensible for its construct, population, and context of use. [3]
Other validity evidence may address structure, expected relationships with other measures, known-group differences, or performance against an appropriate criterion where one genuinely exists. [1][4]
Responsiveness matters when the purpose is change
A measure selected for repeated use needs evidence that it can detect relevant change over the intended interval, not merely evidence that it distinguishes people at one time point. [1]
- Define the expected direction, magnitude, and time frame of change before choosing an analysis.
- Look for prespecified hypotheses and an appropriate comparison in responsiveness studies.
- Consider floor and ceiling effects that may restrict detectable movement.
- Keep statistical change, detectable change, and meaningful change as distinct claims.
A compact appraisal routine
- Name the exact property and intended interpretation being evaluated.
- Check the population, setting, version, language, administration, and score used.
- Inspect study design, sample size, missing data, estimates, and uncertainty.
- Compare the evidence with a prespecified adequacy criterion rather than a post hoc impression.
- Record limitations and decide whether they matter for the proposed use.
Sources and further reading
- COSMIN Taxonomy of Measurement Properties (opens in a new tab)COSMIN. Accessed 2026-07-13. Official taxonomy and definitions for measurement properties of health instruments.
- COSMIN Risk of Bias checklist for Patient-Reported Outcome Measures (opens in a new tab)COSMIN. Published 2018. Accessed 2026-07-13. Official checklist for evaluating study methods across measurement properties.
- COSMIN methodology for assessing the content validity of PROMs (opens in a new tab)COSMIN. Published 2018. Accessed 2026-07-13. Official content-validity method covering relevance, comprehensiveness, and comprehensibility.
- Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims (opens in a new tab)U.S. Food and Drug Administration. Published 2009. Accessed 2026-07-13. Regulatory guidance on instrument development and evidence for score interpretation in a defined context of use.