Start Here
Not sure which test to use? The Reliability & Agreement Wizard asks a few questions and gives you a direct recommendation.
Open the Reliability & Agreement Wizard →
StatClinicClinical Statistics SuiteQuantifying the consistency of a measurement -- across items, raters, or methods.
Not sure which test to use? The Reliability & Agreement Wizard asks a few questions and gives you a direct recommendation.
Open the Reliability & Agreement Wizard →Journals expect more than a p-value. For each test, report the effect estimate and its 95% confidence interval alongside the p-value:
Alpha measures multi-item scale consistency; Kappa measures categorical rater agreement; ICC measures continuous rater or test-retest agreement; Bland-Altman compares two measurement methods.
Alpha measures internal consistency across multiple questionnaire items. Kappa measures agreement between raters on a categorical outcome. ICC measures agreement between raters or repeated measurements on a continuous variable. Bland-Altman compares two measurement methods for the same continuous variable.
0.7-0.8 is acceptable, 0.8-0.9 is good, above 0.9 is excellent -- though values above 0.95 can indicate item redundancy rather than true reliability.
This is the 'kappa paradox' -- Kappa is sensitive to how the categories are distributed (prevalence). Always report raw percent agreement alongside Kappa.
It depends on your design: whether raters are a random or fixed sample, whether you want the reliability of a single measurement or the average of several, and whether you need absolute agreement or just consistency. Report exactly which of the six ICC forms you used.
Two methods can correlate very highly while still disagreeing systematically (e.g. one always reads 10% higher). Bland-Altman directly shows the bias and the range of disagreement (limits of agreement).
More raters or repeated measurements narrow the confidence interval around your reliability estimate -- as a rough guideline, aim for at least 30 subjects rated by all raters.