Two radiologists independently classify 200 chest CT scans as showing pulmonary nodules (yes/no). They agree 91% of the time — a figure that looks reassuring at first glance. But when you account for the fact that 85% of the scans are nodule-free, two raters who always say "no" would agree 72% of the time purely by chance. The true agreement above chance is far lower than 91% suggests. This is exactly the problem that Cohen's Kappa was designed to solve. First published by Jacob Cohen in 1960, kappa is the gold standard measure of inter-rater agreement for categorical data in medical research. Understanding its formula, interpretation scale, and limitations is essential for any researcher conducting reliability studies, diagnostic test evaluations, or clinical scoring validation.
What Is Cohen's Kappa?
Cohen's Kappa (κ) is a chance-corrected measure of agreement between two raters classifying the same set of subjects into mutually exclusive categorical groups. Unlike simple percentage agreement (proportion of cases where both raters choose the same category), kappa accounts for the agreement that would be expected by chance given each rater's marginal distribution of responses.
The logic is straightforward: if Rater A classifies 90% of slides as "benign" and Rater B also classifies 90% as "benign," the two raters will agree on the "benign" classification by chance 81% of the time (0.9 × 0.9) — before even looking at the same slides. Kappa subtracts this chance agreement from the observed agreement and scales the remainder relative to the maximum possible improvement beyond chance.
Kappa = 1 indicates perfect agreement beyond chance. Kappa = 0 indicates agreement exactly at the chance level. Kappa < 0 indicates systematic disagreement (raters disagree more than chance would predict). Values between 0 and 1 represent the degree of agreement corrected for chance.
Why Percentage Agreement Is Insufficient
Before kappa became standard, researchers reported simple percentage agreement: the proportion of cases where both raters made the same classification. This measure has a fundamental flaw: it is inflated by the marginal frequencies of the categories.
The prevalence paradox: When one category is very common (e.g., 95% of biopsies are benign), two raters can achieve 90%+ agreement simply by saying "benign" to everything — without any genuine discriminatory ability. High percentage agreement in highly skewed distributions is meaningless without chance correction.
Consider a diagnostic scenario where two pathologists examine 100 lymph node biopsies. Ninety are reactive hyperplasia and ten are lymphoma. Pathologist A identifies 8 lymphomas; Pathologist B identifies 9. They agree on exactly the same lymphoma cases in 7 instances:
| Path B: Lymphoma | Path B: Reactive | Row Total |
| Path A: Lymphoma | 7 | 1 | 8 |
| Path A: Reactive | 2 | 90 | 92 |
| Column Total | 9 | 91 | 100 |
Percentage agreement = (7 + 90) / 100 = 97% — impressive. But the expected chance agreement is (8×9/100) + (92×91/100) = 0.72 + 83.72 = 84.44%. Cohen's Kappa = (97% − 84.44%) / (100% − 84.44%) = 12.56% / 15.56% = 0.81 — almost perfect, but much more honest than 97%.
Now consider a more challenging scenario where the raters agree on only 3 of the 10 lymphoma cases. The percentage agreement would drop to around 93% (still seeming high), while kappa might fall to 0.40 (moderate) — a far more alarming assessment of diagnostic reliability.
For any square contingency table (two raters classifying subjects into the same k categories), kappa is defined as:
Step-by-Step Calculation (Binary Case)
Using the lymph node biopsy data above (n = 100, table: 7, 1, 2, 90):
- Po (observed agreement) = (7 + 90) / 100 = 0.97
- Pe (expected by chance) = (8×9/100 + 92×91/100) / 100 = (0.72 + 83.72) / 100 = 0.8444
- κ = (0.97 − 0.8444) / (1 − 0.8444) = 0.1256 / 0.1556 = 0.807
This kappa of 0.807 falls in the "almost perfect" range — substantially different from the deceptively high 97% agreement and giving a much more accurate picture of diagnostic concordance.
Landis & Koch Interpretation Scale
The most widely cited interpretation scale for kappa values in medical research was published by Landis and Koch in 1977:
| Kappa Value (κ) | Strength of Agreement | Clinical Interpretation |
| < 0.20 | Slight | Agreement barely exceeds chance; major rater training issues |
| 0.21 – 0.40 | Fair | Low reliability; requires protocol revision before clinical use |
| 0.41 – 0.60 | Moderate | Acceptable for exploratory research; borderline for clinical tools |
| 0.61 – 0.80 | Substantial | Good reliability; acceptable for most clinical decision support |
| 0.81 – 1.00 | Almost Perfect | Excellent reliability; suitable for guideline-level recommendations |
| 1.00 | Perfect | Complete agreement beyond chance (rare in practice) |
Important limitation: The Landis and Koch scale is empirical, not mathematically derived. Other published scales (e.g., Fleiss 1981, Altman 1991) use different thresholds. Always state which scale you are using when reporting kappa in your thesis or paper. The Landis and Koch (1977) benchmarks are the most widely adopted in medical research.
Clinical Examples from Radiology, Pathology, and Diagnosis
1
Radiology: Inter-Rater Agreement for Pulmonary Nodule Classification
Two radiologists independently review 150 CT chest scans and classify each nodule (when present) as benign, indeterminate, or suspicious. A third radiologist arbitrates discordant cases. The study aims to establish inter-rater reliability before implementing the classification in a multicentre lung cancer screening trial.
For simplicity, the data for the binary classification (suspicious vs non-suspicious) are:
| Rad B: Suspicious | Rad B: Not Suspicious |
| Rad A: Suspicious | 28 | 6 |
| Rad A: Not Suspicious | 8 | 108 |
Observed agreement
Po = (28 + 108) / 150 = 136/150 = 0.907
Expected agreement by chance
Row A1 total = 34, Row A2 = 116; Col B1 = 36, Col B2 = 114
Pe = [(34×36)/150 + (116×114)/150] / 150
= [8.16 + 88.08] / 150 = 96.24/150 = 0.6416
Cohen's Kappa
κ = (0.907 − 0.6416) / (1 − 0.6416) = 0.2654 / 0.3584 = 0.741
95% CI (from SE formula)
SE ≈ 0.051 → 95% CI: 0.64 – 0.84
Result: κ = 0.741 (95% CI: 0.64–0.84) — Substantial agreement (Landis & Koch 1977)
Interpretation: The two radiologists achieved substantial agreement (κ = 0.741) in classifying pulmonary nodules as suspicious or non-suspicious. Percentage agreement was 90.7%, but this is partly explained by the high prevalence of non-suspicious scans (82% of cases). Kappa reveals that the true agreement beyond chance is substantial but not near-perfect — the six cases where Radiologist A was suspicious but B was not, and the eight reverse cases, represent meaningful discordance that warrants arbitration protocols. For a clinical screening programme, a pre-specified target of κ ≥ 0.75 should be set in the study protocol.
2
Pathology: Gleason Score Agreement for Prostate Biopsy (Weighted Kappa)
Two pathologists grade 80 prostate biopsy cores using the modified Gleason system, simplified to three groups: low grade (Gleason ≤ 6), intermediate grade (Gleason 7), and high grade (Gleason ≥ 8). Weighted kappa is used because the categories are ordered — being off by one grade is less serious than being off by two.
The 3×3 agreement table:
| Path A \ Path B | Low (≤6) | Intermediate (7) | High (≥8) |
| Low (≤6) | 28 | 5 | 1 |
| Intermediate (7) | 4 | 18 | 3 |
| High (≥8) | 0 | 4 | 17 |
Observed agreement (diagonal)
Po = (28 + 18 + 17) / 80 = 63/80 = 0.788
Expected chance agreement
Row totals: 34, 25, 21 | Col totals: 32, 27, 21
Pe = [(34×32 + 25×27 + 21×21) / 80] / 80
= [1088 + 675 + 441] / 6400 = 2204/6400 = 0.344
Unweighted Kappa
κ = (0.788 − 0.344) / (1 − 0.344) = 0.444 / 0.656 = 0.677
Weighted Kappa (linear weights)
κ_w = 0.743 (penalises adjacent disagreements less; computed by SPSS)
Result: Unweighted κ = 0.677, Weighted κ = 0.743 (95% CI: 0.63–0.85) — Substantial agreement
Interpretation: For treatment decisions in prostate cancer (active surveillance vs radical therapy), Gleason grade concordance is critical. An unweighted kappa of 0.677 and weighted kappa of 0.743 both indicate substantial agreement — but the 12 discordant cases (15%) that fall off the diagonal represent patients whose treatment pathway might differ depending on which pathologist reads the biopsy. The weighted kappa is preferred here because the three grades are ordered — a Low/High discordance (2 cases) is clinically far more serious than a Low/Intermediate discordance (5 cases), and linear weighting captures this appropriately.
3
Clinical Diagnosis: Depression Screening Tool Validation
A psychiatric research team validates a new 9-item depression screening questionnaire (cutoff score ≥ 5 = probable depression) against a structured clinical interview (SCID-5) as the gold standard, in 200 consecutive outpatients. Kappa measures how well the screening tool classification agrees with the clinical diagnosis — a form of diagnostic agreement analysis.
| SCID: Depressed | SCID: Not Depressed |
| Screen: Positive (≥5) | 62 | 19 |
| Screen: Negative (<5) | 11 | 108 |
Observed agreement
Po = (62 + 108) / 200 = 170/200 = 0.850
Expected agreement
Row 1 = 81, Row 2 = 119 | Col 1 = 73, Col 2 = 127
Pe = [(81×73 + 119×127) / 200] / 200
= [5913 + 15113] / 40000 = 21026/40000 = 0.5257
Cohen's Kappa
κ = (0.850 − 0.5257) / (1 − 0.5257) = 0.3244 / 0.4743 = 0.684
Result: κ = 0.684 (95% CI: 0.60–0.77) — Substantial agreement
Interpretation: A kappa of 0.684 indicates substantial agreement between the screening tool and the gold-standard clinical interview. The simple percentage agreement of 85% would give an overly optimistic impression given the 36.5% prevalence of depression in this outpatient sample. The 30 discordant classifications (11 false negatives + 19 false positives) highlight the tool's limitations: the 11 missed cases (false negatives) are the more clinically concerning category, and the authors should report sensitivity (85%) and specificity (85%) alongside kappa to fully characterise diagnostic performance.
Weighted Kappa for Ordinal Data
When the categories are ordered (e.g., disease severity: none/mild/moderate/severe), not all disagreements are equally serious. Weighted kappa assigns partial credit based on how far apart two classifications are:
Linear Weights
Penalty proportional to the distance between categories:
wij = 1 − |i−j| / (k−1)
An adjacent disagreement (mild vs moderate) receives partial credit; a two-step disagreement (mild vs severe) receives less.
Quadratic Weights
Penalty proportional to the squared distance:
wij = 1 − (i−j)² / (k−1)²
Larger disagreements are penalised much more heavily. Quadratic-weighted kappa is numerically equivalent to ICC(2,1) for ordinal data — the two statistics are interchangeable.
When to use weighted kappa: Use weighted kappa (linear or quadratic) when your outcome is an ordered categorical scale with ≥ 3 levels and when different magnitudes of disagreement carry different clinical significance. For dichotomous (binary) outcomes, weighted and unweighted kappa are identical. Report which weighting scheme was used and cite the justification.
Kappa vs Other Agreement Measures
Cohen's Kappa vs Intraclass Correlation Coefficient (ICC)
The key distinction is the measurement level of the outcome:
- Cohen's Kappa: Categorical data (nominal or ordinal). Two raters classify subjects into discrete groups.
- ICC: Continuous or ratio-level data. Two or more raters provide a numerical measurement (blood pressure, tumour size, pain score on a continuous scale).
- For ordinal data with 3+ categories: quadratic-weighted kappa ≈ ICC(2,1) — both are acceptable, and you should report whichever is more common in your field.
Cohen's Kappa vs Bland-Altman Analysis
Bland-Altman analysis is for agreement between two quantitative measurement methods (both measuring the same continuous variable). It plots the difference between measurements against the mean, producing limits of agreement. This is used for method comparison studies (e.g., comparing two blood pressure devices), not for categorical classification agreement.
Multi-Rater Kappa: Fleiss's Kappa
Cohen's kappa applies to exactly two raters. When three or more raters independently classify the same subjects, use Fleiss's kappa — a generalisation that computes overall agreement across all rater combinations. Fleiss's kappa is available in R (irr package: kappam.fleiss()) and SPSS (Reliability analysis → Kappa with multiple raters).
Thesis and Research Reporting Recommendations
Reliability studies using Cohen's Kappa have specific reporting requirements that differ from standard inferential analyses:
Describe the Raters and Classification System
State the number of raters, their qualifications and experience level, whether they were blinded to each other's ratings, whether they received training before the study, and the exact classification system and categories used. The reproducibility of your kappa depends on who classified and under what conditions.
Report the Full Agreement Table
Always include the complete cross-tabulation table, not just the kappa value. Reviewers need to assess whether one category dominates the margins (which can cause the kappa paradox — high kappa with apparently low percentage agreement, or vice versa).
Include Both Kappa and Percentage Agreement
Report both the kappa value (with 95% CI) and the raw percentage agreement. Each provides different information, and readers of different backgrounds will expect both. The percentage agreement situates the result clinically; kappa provides the statistically corrected figure.
Model Reporting Sentence (Inter-Rater)
"Inter-rater agreement between the two radiologists for pulmonary nodule malignancy classification was assessed using Cohen's Kappa. Overall agreement was 90.7% (136/150 cases), with a kappa of 0.741 (95% CI: 0.64–0.84, p < 0.001), indicating substantial agreement according to the Landis and Koch (1977) benchmark scale. Discordant classifications occurred in 14 cases (9.3%), of which six represented over-classification by Radiologist A and eight by Radiologist B."
Model Reporting Sentence (Intra-Rater)
"Intra-rater reliability for the same radiologist re-reading 50 randomly selected scans after a 4-week washout was excellent (κ = 0.88, 95% CI: 0.77–0.99, p < 0.001), demonstrating consistent individual classification performance independent of memory effects."
State the Interpretation Scale
Different kappa interpretation scales exist. Always cite which one you used: Landis and Koch (1977) is most common in medical research; Fleiss (1981) uses different labels (poor/fair/good/excellent); Altman (1991) has a different five-tier system. The threshold for "acceptable" agreement should be pre-specified in your protocol, not selected after seeing results.
Common Mistakes Researchers Make
Mistake 1: Reporting Percentage Agreement Instead of Kappa
Stating that "raters agreed 92% of the time" without reporting kappa is increasingly unacceptable in peer-reviewed medical journals. Percentage agreement does not correct for chance and is systematically inflated when one outcome category is dominant.
Fix: Always report kappa with its 95% CI as the primary agreement statistic. Include percentage agreement as supplementary context, not the primary reliability measure.
Mistake 2: Interpreting Kappa Without Its Confidence Interval
A kappa of 0.62 from a study of 30 subjects has a very wide 95% CI (perhaps 0.35–0.89), spanning fair to almost-perfect agreement. The point estimate alone is misleading and cannot support strong clinical conclusions about reliability.
Fix: Always compute and report the 95% CI for kappa. In SPSS: Analyze → Scale → Reliability Analysis (check "Intraclass correlation coefficient" and select kappa). In R: use CohenKappa() from DescTools with conf.level = 0.95.
Mistake 3: Using Unweighted Kappa for Ordered Categorical Data
Applying unweighted kappa to a 5-point severity scale treats a one-step disagreement (mild vs moderate) as equivalent to a four-step disagreement (mild vs very severe). This distorts the reliability estimate and underestimates agreement when most discordances are minor.
Fix: Use linear-weighted kappa for ordinal scales where adjacent disagreements are less concerning. Use quadratic-weighted kappa when large disagreements are disproportionately serious. Report the weighting scheme used.
Mistake 4: Ignoring the Kappa Paradox
The "kappa paradox" occurs when highly imbalanced marginal distributions produce a low kappa despite high percentage agreement, or vice versa. This happens because kappa's expected chance agreement (Pe) is sensitive to marginal prevalence. A disease with very low prevalence can yield a deceptively low kappa even when raters genuinely agree on the rare cases.
Fix: Report the marginal totals and prevalence in your study. If the kappa paradox is suspected, supplement kappa with positive and negative agreement rates (PA and NA) or Gwet's AC1, which is less sensitive to prevalence imbalance.
Mistake 5: Treating Kappa as a Correlation Coefficient
Kappa measures agreement, not correlation. A kappa of 0.80 does not mean the two raters' classifications correlate 0.80 or that 80% of classifications are correct. It means the agreement is 80% above the chance level — a fundamentally different interpretation.
Fix: Use precise language. Say "κ = 0.80 indicates substantial agreement beyond chance" rather than "raters were 80% consistent" or "ratings correlated at r = 0.80."
Mistake 6: Using Cohen's Kappa for More Than Two Raters
Cohen's original formula is derived for exactly two raters. Attempting to apply it to three or more raters by averaging pairwise kappas produces a biased, non-standard statistic with undefined sampling properties.
Fix: For three or more raters, use Fleiss's kappa (generalised for multiple raters). In SPSS: Analyze → Scale → Reliability Analysis → Statistics → Intraclass Correlation (for continuous) or use the DescTools R package for Fleiss's kappa.
Sample Size and Statistical Significance
Cohen's Kappa is both a descriptive reliability measure and an inferential statistic. Its null hypothesis is H₀: κ = 0 (no agreement beyond chance). Most studies with ≥ 30 subjects will produce a statistically significant kappa — so the p-value for kappa is rarely the primary concern. What matters more is:
- The kappa point estimate and its 95% CI: Does the lower bound of the CI exceed your pre-specified minimum acceptable kappa (e.g., 0.60)?
- The number of subjects needed for a narrow CI: For a binary outcome with true κ = 0.70, you typically need 80–100 subjects to achieve a 95% CI width of ±0.10.
- Prevalence of the outcome: Studies with very low or very high prevalence need more subjects to achieve stable kappa estimates because few discordant cases accumulate for the rare category.
Rule of thumb for planning: Aim for at least 50 subjects per rater for binary outcomes, and at least 50 subjects per category level for ordinal outcomes with 3+ levels. Always perform a formal sample size calculation using anticipated prevalence and expected kappa before starting a reliability study.
Practical Tips for Medical Researchers
Blind Raters to Each Other's Classifications
Raters must classify subjects independently, with no knowledge of the other rater's decisions. Any communication or consensus during the rating process invalidates the kappa as a measure of independent agreement. Document blinding procedures in your methods section.
Pre-Specify Your Minimum Acceptable Kappa
Define in your study protocol the minimum kappa threshold for "acceptable" reliability before analysing data. For clinical decision support tools, κ ≥ 0.70 is a reasonable target. For screening tools in lower-stakes settings, κ ≥ 0.60 may suffice. Post-hoc threshold selection is a form of outcome reporting bias.
Conduct a Calibration Exercise Before the Study
Train raters using a reference set of cases (not included in the main study) and conduct a calibration exercise to align classifications. This improves kappa and reduces the burden on statistical correction. Document the calibration process and any training provided.
Report Subcategory Agreement
For multi-category scales, report category-specific agreement rates (sensitivity and specificity for each category) alongside the overall kappa. A high overall kappa can mask poor agreement in clinically important minority categories — exactly where reliability matters most.
Use SPSS or R for Confidence Intervals
Manual kappa calculation is feasible for 2×2 tables, but 95% confidence intervals require software. SPSS reports kappa under Analyze → Scale → Reliability Analysis or Analyze → Nonparametric → Cross-tabs with kappa option. R: cohen.kappa(matrix) from psych or CohenKappa() from DescTools.
Consider Gwet's AC1 for Skewed Prevalence
When prevalence is extreme (<10% or >90%), the kappa paradox can produce misleadingly low kappa values despite genuine high agreement. Gwet's AC1 is a more stable alternative that is less sensitive to marginal distributions. Report both kappa and AC1 if prevalence is highly skewed, noting the reason for the discrepancy.
Frequently Asked Questions
What is Cohen's Kappa and why is it used in medical research? +
Cohen's Kappa (κ) is a chance-corrected measure of agreement between two raters classifying the same subjects into categorical groups. It is used in medical research to quantify inter-rater or intra-rater reliability for categorical outcomes. Kappa is preferred over simple percentage agreement because it corrects for the agreement that would occur by chance given the marginal distributions of each rater — making it a more honest and informative measure of true agreement, particularly when one category is more common than others.
What does a kappa value of 0.60 mean? +
According to the Landis and Koch (1977) benchmarks, κ = 0.60 falls at the upper boundary of "Moderate" agreement (κ = 0.41–0.60) — just below the "Substantial" range (κ = 0.61–0.80). In clinical terms, this represents meaningful agreement beyond chance but may be insufficient for high-stakes diagnostic tools. Whether it is acceptable depends on context: a κ of 0.60 might be adequate for a low-stakes screening instrument but inadequate for a pathology grading system that determines treatment choice between active surveillance and radical surgery.
What is the difference between weighted and unweighted kappa? +
Unweighted kappa treats all disagreements as equally serious, regardless of how far apart the two classifications are. Weighted kappa assigns smaller penalties to disagreements between adjacent categories on an ordered scale. For ordinal outcomes (e.g., none/mild/moderate/severe), weighted kappa is more appropriate because being off by one grade (mild vs moderate) is clinically less serious than being off by three grades (none vs severe). Linear and quadratic weighting schemes are the two most common options; quadratic-weighted kappa is equivalent to the ICC for ordinal data.
When should I use Cohen's Kappa vs ICC? +
Use Cohen's Kappa for categorical data (nominal or ordinal categories). Use ICC for continuous or ratio-level numeric data (e.g., blood pressure readings, tumour diameters). For ordinal data with 4+ ordered levels, quadratic-weighted kappa and ICC(2,1) give nearly identical values and either can be used. For binary data, ICC and kappa produce the same result. Never use ICC for nominal (unordered) categorical data, as it assumes interval-level measurement.
How many subjects do I need to reliably estimate Cohen's Kappa? +
A minimum of 30–50 subjects is commonly cited, but this depends on the expected kappa, the number of categories, and the prevalence of each category. For binary outcomes with balanced prevalence (~50%), 50 subjects typically provides a 95% CI of approximately ±0.14. For rare outcomes (prevalence <10%) or multi-category scales, 80–120 subjects per category are recommended. Always compute a formal sample size estimate using published formulas (Donner & Eliasziw 1992) or R's kappaSize package.
Can Cohen's Kappa be negative? +
Yes. Kappa ranges from −1 to +1. A negative kappa means raters agree less than expected by chance — they are systematically disagreeing. This almost always indicates a measurement or protocol problem: perhaps the two raters are interpreting the scale in opposite directions (one uses 1 = most severe, the other uses 1 = least severe), or there is a systematic training difference. Negative kappas require investigation before any further data collection or analysis.
How do I report Cohen's Kappa in a thesis or journal paper? +
State the number of raters, subjects, and categories. Report: kappa (to 2 decimal places), 95% CI, p-value, percentage agreement, and a verbal interpretation citing a specific scale. Example: "Inter-rater agreement was substantial (κ = 0.72, 95% CI 0.61–0.83, p < 0.001; observed agreement 90.7%), using the Landis and Koch (1977) benchmark scale." Always include the full cross-tabulation table in your results section.
What is the kappa paradox? +
The kappa paradox occurs when highly skewed marginal distributions produce counter-intuitive kappa values — either a very low kappa despite high percentage agreement, or (less commonly) a moderate kappa with low percentage agreement. This happens because kappa's expected chance agreement is inflated by marginal imbalance. For example, if 95% of samples are negative, two raters who both always say "negative" achieve 95% agreement but κ = 0. When prevalence is extreme, consider supplementing kappa with Gwet's AC1, which is more robust to marginal skewness.
Can I use the McNemar test with Cohen's Kappa in the same study? +
Yes, and they answer different questions. Cohen's Kappa measures the degree of agreement between two raters or methods classifying the same subjects. The McNemar test asks whether the two methods have significantly different positivity rates in the same subjects (marginal homogeneity). A study validating a new diagnostic test might report kappa to show agreement, and McNemar to test whether the new test has a systematically higher or lower positive rate than the reference standard. Both can be reported from the same 2×2 table.
What is the difference between inter-rater and intra-rater reliability? +
Inter-rater (inter-observer) reliability measures agreement between two or more different observers classifying the same subjects. Intra-rater reliability measures how consistently a single observer classifies the same subjects when re-rated after a time interval (test-retest reliability). Cohen's Kappa measures both: for intra-rater, the two columns represent the same observer at two time points. Studies in clinical measurement should ideally report both to distinguish between observer-to-observer variability and within-observer temporal variability.
Calculate Cohen's Kappa Online
Use StatClinic's built-in kappa calculator to quantify inter-rater agreement from your 2×2 or k×k contingency table. Get kappa, weighted kappa, 95% confidence intervals, and p-values formatted for thesis and journal submission.
Open Kappa Calculator →