Calculator — ICC(1,1) One-Way Random Effects
Each row/value in Rater 1 must correspond to the same subject in Rater 2. Minimum 3 pairs required.
StatClinicClinical Statistics SuiteCompute the Intraclass Correlation Coefficient (ICC) for continuous measurement reliability. Enter paired measurements from two raters or two time points per subject.
Each row/value in Rater 1 must correspond to the same subject in Rater 2. Minimum 3 pairs required.
The Intraclass Correlation Coefficient (ICC) measures the reliability of continuous measurements made by two or more raters on the same subjects. It expresses the proportion of total variance attributable to true between-subject differences rather than measurement error. ICC ranges from 0 (no reliability) to 1 (perfect reliability).
Model used — ICC(1,1) One-Way Random: Assumes each subject is rated by a different, randomly selected set of raters. MSB = Mean Square Between subjects; MSW = Mean Square Within (error). Formula:
ICC = (MSB − MSW) / (MSB + (k−1) × MSW)
Where k = number of raters (2). The 95% CI uses the F-distribution: Lower = (F/FU−1)/(F/FU+k−1); Upper = (F·FL−1)/(F·FL+k−1).
Reporting format: ICC(1,1) = 0.97 (95% CI: 0.91–0.99, F(9,10) = 260.8, p < 0.001)
The ICC is a reliability statistic for continuous measurements. It expresses the proportion of total variance due to true subject differences versus measurement error. ICC values range from 0 (no reliability) to 1 (perfect reliability). Unlike Pearson r, ICC is sensitive to both correlation pattern and mean shifts (systematic bias) between raters.
Three main models: ICC(1,1) one-way random — raters are randomly selected, different subjects may be rated by different raters; ICC(2,1) two-way random — all raters rate all subjects, raters are a random sample; ICC(3,1) two-way mixed — all raters rate all subjects, raters are the specific fixed raters of interest. This calculator uses ICC(1,1). For studies where the same two raters assess all subjects consistently, ICC(2,1) is often preferred but gives nearly identical values.
Koo & Mae (2016): ICC < 0.50 = Poor; 0.50–0.74 = Moderate; 0.75–0.89 = Good; ≥ 0.90 = Excellent. For clinical instruments used for individual patient decisions (e.g., diagnostic cutoffs), ICC ≥ 0.90 is recommended. For group-level research comparisons, ICC ≥ 0.70 may be sufficient.
Pearson r measures linear association but is blind to systematic mean differences. If Rater 2 consistently scores 10 points higher than Rater 1, Pearson r = 1.0 but ICC would be low. ICC assesses absolute agreement, making it the correct measure for measurement reliability studies.
Report the model, value, 95% CI, F-statistic, degrees of freedom, p-value, and N. Example: "Test-retest reliability was excellent (ICC(1,1) = 0.97, 95% CI: 0.91–0.99, F(9,10) = 260.8, p < 0.001, N = 10; Koo & Mae, 2016)."