What Is a Cohort Study?
A cohort study is an observational epidemiological design in which a defined group of people — the cohort — is assembled based on their exposure status and followed over time to observe the development of outcomes. The word "cohort" comes from the Latin for a unit of Roman soldiers: a group with a shared characteristic who move through time together.
The defining feature of cohort studies is that exposure is measured before the outcome occurs. This temporal sequence — exposure → follow-up → outcome — is what gives cohort studies their strength in establishing the direction of association. Unlike case-control studies, cohort studies allow direct calculation of incidence rates, relative risk, attributable risk, and absolute risk differences in the exposed and unexposed groups.
Prospective vs Retrospective Cohort Studies
Cohort studies divide into two broad types based on their temporal relationship to the researcher at the time of study initiation.
| Feature | Prospective Cohort | Retrospective Cohort |
|---|---|---|
| Timeline at enrolment | Outcome has NOT yet occurred; participants are followed forward | Both exposure and outcome have already occurred; researcher looks back |
| Data source | Primary data collected during follow-up (questionnaires, clinical tests) | Existing records (hospital databases, employment records, disease registries) |
| Exposure measurement | Measured carefully at baseline and during follow-up; minimal recall bias | Depends on historical record quality; potential for incomplete data |
| Outcome ascertainment | Standardised prospectively by study protocol | Based on existing diagnostic codes or records — may vary by era |
| Time required | Years to decades for chronic disease outcomes | Months to years — data already exist |
| Cost | High — participant follow-up, repeated measurements | Lower — no active follow-up required |
| Loss to follow-up | Real risk — must plan correction in sample size | Often lower — records are fixed (though records may be incomplete) |
| Classic example | Framingham Heart Study (1948–present) | Nurses' Health Study reanalysis of historical questionnaire data |
| Sample size formula | Same formula — parameters differ | Same formula — parameters differ |
Why Sample Size Matters in Cohort Studies
Sample size is not a bureaucratic hurdle. It directly controls two statistical quantities that determine whether your study will produce actionable, publishable, and scientifically credible results: Type I error (α) and Type II error (β).
- Type I error (α; false positive): Concluding there is an association when there is none. Controlled by the significance threshold (typically α = 0.05).
- Type II error (β; false negative): Failing to detect a real association that truly exists. Controlled by study power = 1 − β. A study with 80% power accepts a 20% chance of missing a true effect.
An underpowered cohort study is not merely less informative — it is potentially actively misleading. A null result from an underpowered study (p > 0.05) cannot be interpreted as evidence of no association. It means only that the study was too small to see the association even if it exists. Systematic reviews routinely exclude underpowered studies because their confidence intervals are so wide they add no information about effect size.
The Sample Size Formula for Cohort Studies
The standard formula for comparing two proportions in a cohort study — the event rate in the exposed group (p₁) versus the event rate in the unexposed group (p₂) — is derived from the two-sample z-test for proportions. The formula gives the sample size required per group, assuming equal group sizes:
The formula has an intuitive structure: the numerator grows when the variance is large (events near 50% frequency) or when a higher z-score is demanded (stricter α, higher power). The denominator shrinks when the expected difference between groups is small, driving n upward — detecting a subtle effect requires a larger study than detecting a large one. Total sample size is simply 2n for equal allocation between exposed and unexposed groups.
Loss-to-Follow-Up Adjustment
Prospective cohort studies must account for participants who withdraw, are lost to follow-up, die of unrelated causes, or otherwise fail to contribute outcome data. The adjustment is straightforward:
All Variables Explained in Depth
Z-Score Reference Table
| Significance Level (α) | zα/2 (Two-Tailed) | Power (1−β) | zβ | Combined (zα/2 + zβ) |
|---|---|---|---|---|
| 0.10 | 1.645 | 80% | 0.842 | 2.487 |
| 0.05 | 1.960 | 80% | 0.842 | 2.802 ← most common |
| 0.05 | 1.960 | 85% | 1.036 | 2.996 |
| 0.05 | 1.960 | 90% | 1.282 | 3.242 |
| 0.05 | 1.960 | 95% | 1.645 | 3.605 |
| 0.01 | 2.576 | 80% | 0.842 | 3.418 |
| 0.01 | 2.576 | 90% | 1.282 | 3.858 |
Step-by-Step Calculation Guide
Define the primary outcome
State a single, clearly measured binary outcome (e.g., incident type 2 diabetes, first myocardial infarction, 30-day mortality). Sample size is calculated for one primary endpoint. Secondary endpoints use the same sample size but are acknowledged to be underpowered for individual analysis.
Determine p₂ from literature
Search existing cohort studies or disease registries for the cumulative incidence of your outcome over your planned follow-up period in a population similar to your unexposed group. Document the source. If multiple estimates exist, use the most conservative (closest to 0.50) or conduct sensitivity analyses.
Determine p₁ from expected RR or minimum detectable difference
Either obtain expected RR from prior evidence (p₁ = p₂ × RR) or define the minimum clinically important absolute risk difference you want to detect (p₁ = p₂ + ARD). The latter approach frames the study by clinical relevance rather than statistical convention.
Set α and power
Use α = 0.05 (two-tailed) and power ≥ 80% for most medical cohort studies. Use α = 0.01 or power = 90% for studies with regulatory implications, rare high-consequence outcomes, or where a false positive or negative would have major policy impact. Set these before calculation — never adjust them after seeing a sample size you dislike.
Apply the formula to get n per group
Substitute all values into the formula. Round up to the nearest whole number (never round down — this would reduce power below the target). Multiply by 2 for total sample size with 1:1 exposed:unexposed allocation.
Adjust for expected loss to follow-up
For prospective studies, inflate each group size: nₙᵈᵍ = n / (1 − L). Plan a conservative L based on study duration, population mobility, and disease severity. A 2-year study in a stable outpatient population might use L = 10–15%; a 10-year community study might use L = 20–30%.
Consider unequal group sizes if indicated
If exposed individuals are rare or costly to recruit, enrol more unexposed controls. The formula adjusts to: n₁ = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)/r] / (p₁−p₂)² where r = n₂/n₁. Maximum efficiency gain occurs at r = 2 (1:2 ratio). Beyond r = 4, gains are negligible.
Document and justify all assumptions
Write out every assumption in your research protocol: the values of p₁, p₂, α, power, L, and the source for each estimate. Ethics committees and journal reviewers will ask for the basis of each assumption. An unjustified assumption is a protocol weakness that invites revision requests.
Worked Epidemiology Examples
The following four examples cover a range of disease areas, follow-up durations, and parameter choices. Each shows the complete calculation with all working visible.
Smoking and Chronic Obstructive Pulmonary Disease — Classic Prospective Cohort
10-year follow-up • Binary outcome: incident COPD diagnosis • 1:1 allocation • α = 0.05, Power = 80%
A pulmonologist plans a 10-year prospective cohort study comparing current smokers (exposed) with lifelong non-smokers (unexposed) for incident COPD. Literature estimates the 10-year COPD incidence in non-smokers at 6% and the expected relative risk for smokers at 3.0, giving an expected incidence of 18% in smokers.
p₁ (exposed, smokers) = p₂ × RR = 0.06 × 3.0 = 0.18 (18%)
α = 0.05 (two-tailed) → zα/2 = 1.960
Power = 80% → zβ = 0.842
Expected attrition over 10 years: L = 20%
Variance Term p₁(1−p₁) = 0.18 × 0.82 = 0.1476
p₂(1−p₂) = 0.06 × 0.94 = 0.0564
Sum = 0.1476 + 0.0564 = 0.2040
Risk Difference Squared (p₁−p₂)² = (0.18−0.06)² = (0.12)² = 0.0144
Sample Size Per Group n = (1.960 + 0.842)² × 0.2040 / 0.0144
n = (2.802)² × 0.2040 / 0.0144
n = 7.851 × 0.2040 / 0.0144
n = 1.602 / 0.0144 = 111.3 → 112 per group
Attrition Adjustment (L = 20%) nₙᵈᵍ = 112 / (1−0.20) = 112 / 0.80 = 140 per group
Total Sample Total = 140 × 2 = 280 participants
Physical Activity and Type 2 Diabetes — Small Risk Difference, High Power Target
5-year follow-up • Outcome: new T2DM diagnosis • Power = 90% • Shows impact of small absolute difference
Researchers plan a 5-year prospective cohort study examining whether regular vigorous exercise (at least 150 minutes/week) reduces type 2 diabetes incidence in adults aged 40–65. The 5-year incidence in sedentary adults (unexposed) is estimated at 12% from national registry data. The research team considers a 4 percentage-point absolute reduction (from 12% to 8%) clinically meaningful. The study targets 90% power.
p₁ (active, exposed) = 0.08 (8%) → RR = 0.67 (33% risk reduction)
α = 0.05 (two-tailed) → zα/2 = 1.960
Power = 90% → zβ = 1.282
Attrition over 5 years: L = 15%
Variance Term p₁(1−p₁) = 0.08 × 0.92 = 0.0736
p₂(1−p₂) = 0.12 × 0.88 = 0.1056
Sum = 0.0736 + 0.1056 = 0.1792
Risk Difference Squared (p₁−p₂)² = (0.08−0.12)² = (−0.04)² = 0.0016
Sample Size Per Group n = (1.960 + 1.282)² × 0.1792 / 0.0016
n = (3.242)² × 0.1792 / 0.0016
n = 10.511 × 0.1792 / 0.0016
n = 1.884 / 0.0016 = 1,177.0 → 1,177 per group
Attrition Adjustment (L = 15%) nₙᵈᵍ = 1,177 / (1−0.15) = 1,177 / 0.85 = 1,385 per group
Total Sample Total = 1,385 × 2 = 2,770 participants
Occupational Solvent Exposure and Hepatocellular Carcinoma — Unequal Groups (1:3 Ratio)
Retrospective cohort • Rare exposed group (factory workers) • 1:3 allocation • Rare outcome
An occupational health researcher uses a company employment database to assemble a retrospective cohort. Factory workers with documented solvent exposure (exposed group) are rare — only about 200 workers meet the exposure criteria from 20 years of records. The researcher plans to match them with unexposed administrative staff in a 1:3 ratio (exposed:unexposed). The 20-year cumulative incidence of hepatocellular carcinoma in occupationally unexposed adults is approximately 0.8%. Expected RR with chronic solvent exposure is 4.0. α = 0.05, power = 80%.
p₁ (exposed workers) = 0.008 × 4.0 = 0.032 (3.2%)
r = 3 (three unexposed per one exposed)
α = 0.05 → zα/2 = 1.960 Power = 80% → zβ = 0.842
Retrospective design: L = 5% (incomplete records)
Unequal Ratio Formula for n₁ (exposed group) n₁ = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)/r] / (p₁−p₂)²
p₁(1−p₁) = 0.032 × 0.968 = 0.030976
p₂(1−p₂)/r = (0.008 × 0.992)/3 = 0.007936/3 = 0.002645
Sum = 0.030976 + 0.002645 = 0.033621
(p₁−p₂)² = (0.032−0.008)² = (0.024)² = 0.000576
n₁ = (2.802)² × 0.033621 / 0.000576
n₁ = 7.851 × 0.033621 / 0.000576 = 0.2640 / 0.000576 = 458.3 → 459 exposed
Unexposed Group n₂ = r × n₁ = 3 × 459 = 1,377 unexposed
Attrition Adjustment (5%) Exposed: 459 / 0.95 = 484 (but only ~200 available!)
Unexposed: 1,377 / 0.95 = 1,449 unexposed
Breastfeeding Duration and Childhood Asthma — Sensitivity Analysis on Power
Prospective cohort • 5-year child follow-up • Shows how power choice changes total N
A paediatrician studies whether exclusive breastfeeding for at least 6 months (exposed) versus formula feeding (unexposed) reduces childhood asthma diagnosis by age 5. The 5-year asthma incidence in formula-fed children is estimated at 22%. An absolute 6 percentage-point reduction to 16% in breastfed children is considered clinically important. The research team wants to compare sample sizes for 80%, 85%, and 90% power at α = 0.05 with L = 15% attrition.
Variance: p₁(1−p₁) + p₂(1−p₂) = (0.16×0.84) + (0.22×0.78) = 0.1344 + 0.1716 = 0.3060
(p₁−p₂)² = (0.16−0.22)² = (−0.06)² = 0.0036
Ratio: 0.3060 / 0.0036 = 85.0
80% Power (zβ = 0.842) n = (1.960 + 0.842)² × 85.0 = (2.802)² × 85.0 = 7.851 × 85.0 = 667.3 → 668
Adjusted: 668 / 0.85 = 787 per group → 1,574 total
85% Power (zβ = 1.036) n = (1.960 + 1.036)² × 85.0 = (2.996)² × 85.0 = 8.976 × 85.0 = 763.0 → 763
Adjusted: 763 / 0.85 = 898 per group → 1,796 total
90% Power (zβ = 1.282) n = (1.960 + 1.282)² × 85.0 = (3.242)² × 85.0 = 10.511 × 85.0 = 893.4 → 894
Adjusted: 894 / 0.85 = 1,052 per group → 2,104 total
The Attrition Correction: How Loss to Follow-Up Inflates Sample Size
Loss to follow-up is one of the most underappreciated threats to cohort study validity and the most common reason prospective studies arrive at analysis with insufficient events to answer their question. The table below shows how different attrition levels inflate an example base sample size of 300 per group:
Impact of Confidence Level, Power, and Risk Difference
Understanding how each parameter drives sample size helps you make informed trade-offs when full-scale recruitment is constrained by budget, time, or available population.
Effect of α Level (Confidence Level)
Tightening the significance threshold from α = 0.05 to α = 0.01 increases the zα/2 term from 1.96 to 2.576. Since sample size scales with the square of (zα/2 + zβ), this increases sample size by roughly 49% at 80% power. A α = 0.01 threshold is warranted when a false positive would have serious downstream consequences — for example, incorrectly declaring a substance carcinogenic and triggering costly regulatory action.
Effect of Power Level
Every 5-percentage-point increase in power approximately adds 14–20% more participants (the exact amount varies with the other parameters). The minimum acceptable power is 80% for most peer-reviewed medical research. Ethics committees will question studies powered below this threshold. A power of 70% or below is generally considered inadequately powered and unlikely to be published if the primary endpoint is not statistically significant.
Effect of Risk Difference
Because risk difference appears squared in the denominator, it has a disproportionately large effect on sample size. Halving the expected risk difference quadruples the required sample size. This is why large, expensive cohort studies investigating small exposures — such as the effect of low-dose environmental pollutants on cardiovascular outcomes — often require tens of thousands of participants and multi-decade follow-up periods.
Common Sample Size Calculation Mistakes
Mistake 1: Using prevalence instead of cumulative incidence for p₂
Cohort studies track new cases over time, so the correct measure is the cumulative incidence (risk) over the follow-up period — the proportion of disease-free participants at baseline who develop the outcome by end of follow-up. Plugging in the overall disease prevalence (which includes existing cases) inflates p₂ and artificially deflates the required sample size.
Mistake 2: Ignoring loss to follow-up in prospective studies
Researchers frequently calculate n from the formula and directly report that number as their recruitment target, without inflating for expected dropout. A prospective study enrolling exactly the formula-calculated n will almost certainly finish underpowered because some fraction of participants will not complete follow-up.
Mistake 3: Over-optimistic expected risk difference
Researchers sometimes choose p₁ and p₂ that produce a conveniently small and feasible n, rather than values grounded in realistic epidemiological expectations. If the true effect is smaller than assumed, the study will be underpowered for the actual difference and may miss it entirely.
Mistake 4: Forgetting that total sample size = 2n, not n
The formula gives n per group. A surprisingly common protocol error states the total sample as n rather than 2n, effectively cutting the study to half the required power without realising it. This halves the number of participants in each arm and substantially reduces power.
Mistake 5: Not conducting sensitivity analyses
A single point estimate of sample size based on one set of assumed parameters creates a brittle calculation. If any assumption is wrong — and at least some usually are — the planned sample size may be inappropriate. Ethics committees and grant reviewers increasingly expect authors to show how sample size changes across a plausible range of assumptions.
Mistake 6: Using one-tailed instead of two-tailed tests without justification
One-tailed tests require fewer participants than two-tailed tests for the same power, because they only look for an effect in one direction. This is tempting when recruits are scarce, but using a one-tailed test is only valid when there is a strong prior reason to believe the exposure can only increase (or only decrease) the outcome. In most epidemiological situations this cannot be justified, and using one-tailed p-values is likely to trigger reviewer rejection.
Research Reporting Examples
Journal methods sections should include all the information a reader needs to independently verify the sample size calculation. The following examples show compliant and non-compliant reporting formats.
Frequently Asked Questions
Calculate your cohort study sample size instantly
StatClinic's free sample size calculator handles two-proportion cohort studies with attrition correction, unequal group ratios, and sensitivity analysis — with plain-English output ready for your thesis or grant application.
Open Sample Size Calculator →