StatClinicStatClinicClinical Statistics Suite
Home/Tools/Reliability & Agreement Hub

Reliability & Agreement Hub

Quantifying the consistency of a measurement -- across items, raters, or methods.

Related Calculators

Related Articles

Medical Research Examples

Cronbach's Alpha: A 10-item anxiety questionnaire administered to 150 patients: Cronbach's alpha=0.86 -- good internal consistency.
Cohen's Kappa: Two radiologists independently classify 100 chest X-rays as normal/abnormal: observed agreement=88%, Kappa=0.74 -- substantial agreement.
ICC: Three physiotherapists measure knee range-of-motion in 30 patients: ICC(2,1)=0.91 (95% CI 0.85-0.95) -- excellent inter-rater reliability.
Bland-Altman: Two methods for measuring cardiac output (thermodilution vs. Doppler) in 50 ICU patients: mean bias=0.2 L/min, 95% limits of agreement −0.9 to 1.3 L/min -- acceptable agreement for clinical use.

Common Mistakes

Reporting Recommendations

Journals expect more than a p-value. For each test, report the effect estimate and its 95% confidence interval alongside the p-value:

Cronbach's AlphaCronbach's alpha with benchmark interpretation (e.g. > 0.7 acceptable, > 0.9 excellent)
Cohen's KappaKappa, 95% CI, Landis & Koch qualitative benchmark
ICC (Intraclass Correlation)ICC estimate with 95% CI, and the specific ICC model/type used (e.g. two-way random, absolute agreement)
Bland-Altman AnalysisMean bias, 95% limits of agreement, and a Bland-Altman plot

Recommended Learning Order

  1. Cronbach's Alpha — start with the most common reliability question: is my questionnaire internally consistent?
  2. Cohen's Kappa — move to categorical agreement between two raters
  3. ICC — extend to continuous measurements and more than two raters
  4. Bland-Altman Analysis — most specialized: comparing two measurement methods directly

Frequently Compared Tests

Cronbach's Alpha vs. Cohen's Kappa vs. ICC vs. Bland-Altman

Alpha measures multi-item scale consistency; Kappa measures categorical rater agreement; ICC measures continuous rater or test-retest agreement; Bland-Altman compares two measurement methods.

Frequently Asked Questions

What's the difference between Cronbach's Alpha, Kappa, ICC, and Bland-Altman?

Alpha measures internal consistency across multiple questionnaire items. Kappa measures agreement between raters on a categorical outcome. ICC measures agreement between raters or repeated measurements on a continuous variable. Bland-Altman compares two measurement methods for the same continuous variable.

What is a good Cronbach's Alpha value?

0.7-0.8 is acceptable, 0.8-0.9 is good, above 0.9 is excellent -- though values above 0.95 can indicate item redundancy rather than true reliability.

Why does Kappa sometimes seem low even when raters mostly agree?

This is the 'kappa paradox' -- Kappa is sensitive to how the categories are distributed (prevalence). Always report raw percent agreement alongside Kappa.

Which ICC model should I use?

It depends on your design: whether raters are a random or fixed sample, whether you want the reliability of a single measurement or the average of several, and whether you need absolute agreement or just consistency. Report exactly which of the six ICC forms you used.

Why use Bland-Altman instead of correlation to compare two methods?

Two methods can correlate very highly while still disagreeing systematically (e.g. one always reads 10% higher). Bland-Altman directly shows the bias and the range of disagreement (limits of agreement).

How many raters or subjects do I need for a reliability study?

More raters or repeated measurements narrow the confidence interval around your reliability estimate -- as a rough guideline, aim for at least 30 subjects rated by all raters.