Analyze My Study
Sample Size Calculation

How to Calculate Sample Size for Case-Control Study in Medical Research

- 17 min read ... June 2025 Updated June 2025
S
StatClinic Editorial Team Statistical content for medical researchers and clinicians
Sample size is not a bureaucratic formality in research design it is the difference between a study that can detect a real effect and one that cannot. In case-control studies, the consequences of miscalculation are especially severe: cases are often rare, recruitment is expensive, and an underpowered study generates a non-significant result that tells you nothing useful while consuming months of effort. Yet a large proportion of medical theses and published case-control studies either skip formal sample size calculation entirely or apply formulas incorrectly, feeding in parameters that are unrealistic or poorly sourced. This guide walks through every component of case-control sample size calculation the formula, its inputs, the reasoning behind each parameter, and the most common errors that lead researchers astray with three fully worked epidemiological examples.

What Is a Case-Control Study?

A case-control study is an observational, retrospective epidemiological design that begins with the outcome and works backward to examine past exposures. Participants are selected based on whether or not they have the disease or condition of interest, rather than based on whether they were exposed to a risk factor.

Cases

Individuals With the Disease

  • Confirmed diagnosis of the condition under study
  • Incident cases (newly diagnosed) preferred over prevalent cases
  • Selected from hospital records, disease registries, or referral centers
  • Exposure history collected retrospectively (interviews, medical records)
Controls

Individuals Without the Disease

  • Free of the disease at the time of selection
  • Must come from the same source population that gave rise to the cases
  • Matched to cases (on age, sex, etc.) or selected independently (unmatched)
  • Same exposure history collection method as cases critical for validity

The fundamental measure of association in a case-control study is the odds ratio (OR) not relative risk. The OR approximates the relative risk when the outcome is rare in the population (prevalence < 10%), which holds for most diseases studied using case-control designs.

Case-control studies occupy a specific niche in the epidemiological hierarchy. They are particularly appropriate for:

The Fundamental Advantage Because you begin with known cases, case-control studies are far more efficient than cohort studies for rare outcomes. A cohort study of gastric cancer would need to follow tens of thousands of people for decades. A case-control study can assemble 80100 confirmed gastric cancer cases from a hospital registry and recruit matched controls from the same institution in a fraction of the time and cost.

Why Sample Size Calculation Matters Before Data Collection

The sample size for your study must be determined before data collection begins. Calculating it retrospectively after results are known is a form of statistical manipulation that inflates or deflates the apparent power of the study and invalidates the reasoning behind the chosen sample.

There are four concrete consequences of inadequate sample size planning:

Ethics and Sample Size Recruiting too few subjects violates the principle of scientific validity the study cannot answer its own question. Recruiting unnecessarily many subjects violates the principle of economy more people are burdened by the study than required. The formally calculated sample size defines the ethically and statistically appropriate scope of your study.

The Four Key Parameters for Case-Control Sample Size

Before applying any formula, you must define four parameters. Each represents a deliberate research decision grounded in prior literature, clinical judgment, or convention.

1. Confidence Level (+/-) Type I Error Rate

The significance level +/- defines the probability you are willing to accept of rejecting the null hypothesis when it is actually true a false positive. By convention in medical research, +/- = 0.05 is the universal standard, corresponding to a 95% confidence interval and a critical Z-value of 1.96 for a two-sided test. Some regulatory or safety studies use +/- = 0.01 (Z = 2.576), which requires a larger sample to compensate for the more stringent threshold.

2. Statistical Power (1 2) Type II Error Rate

Power is the probability of detecting a real effect if one truly exists, set conventionally at 80% (2 = 0.20) for most medical research. A power of 80% means your study has a 20% chance of producing a false negative of missing a real association. Higher power (90%, 95%) is preferred for high-stakes research but requires larger samples.

3. Exposure Prevalence Among Controls (p)

The proportion of people in the control group who have been exposed to the risk factor of interest. This must come from population data, published literature, or prior pilot studies it represents background exposure in the disease-free population. This is the parameter researchers most frequently misestimate, and errors here have a large impact on the final sample size.

4. Expected Odds Ratio (OR)

The anticipated strength of association between the exposure and the disease. This should come from prior case-control studies, cohort studies, systematic reviews, or meta-analyses on the same or similar exposure-outcome relationship. The OR is the most powerful driver of sample size: increasing the expected OR from 1.5 to 3.0 can reduce the required sample by 80%.

ParameterCommon ValueZ Critical ValueEffect on n
+/- = 0.05, two-sided Most common Z+/-/2 = 1.96 Baseline
+/- = 0.01, two-sided High rigor studies Z+/-/2 = 2.576 Larger n
Power = 80% (2 = 0.20) Most common Z2 = 0.842 Baseline
Power = 90% (2 = 0.10) Preferred Z2 = 1.282 ~30% larger n
Power = 95% (2 = 0.05) High-stakes Z2 = 1.645 ~60% larger n

The Kelsey Formula for Case-Control Sample Size

For unmatched (or individually unmatched) case-control studies, the most widely referenced sample size formula is the Kelsey formula (Kelsey, Whittemore, Evans & Thompson, 1996), used in EpiInfo, OpenEpi, and most biostatistics texts:

Kelsey Formula Unmatched Case-Control Study
n = [ Z+/-/2 ((1+1/r) p(1p)) + Z2 (p(1p) + p(1p)/r) ]2

(p p)2
n = Number of cases required
r = Ratio of controls to cases (1:1, 1:2, etc.)
p = Exposure prevalence among controls
p = Exposure prevalence among cases*
p = Weighted average exposure = (p + r-p)/(1+r)
Z+/-/2 = 1.96 (for +/- = 0.05, two-sided)
Z2 = 0.842 (for 80% power)
Total n = Cases (n) + Controls (n - r)
Deriving p from the Odds Ratio You never input p directly from the literature it is derived from p and the expected OR using this formula: p = (OR - p) / (1 p + OR - p). This formula converts the expected OR and the background exposure prevalence into the expected exposure prevalence among cases. This step is the most frequently confused part of the calculation.

The Control-to-Case Ratio (r) and Sample Efficiency

When recruiting cases is difficult (rare disease, specialized registry), increasing the number of controls per case can compensate for insufficient cases. The marginal gain in power from adding more controls diminishes rapidly beyond a 1:4 ratio. Common choices in medical research:

Step-by-Step Worked Calculation: Smoking and COPD

The following worked example demonstrates a complete sample size calculation for an unmatched case-control study investigating smoking as a risk factor for COPD in a tertiary care respiratory clinic. You can verify this calculation using the StatClinic Sample Size Calculator (select the OR/RR option).

Worked Example Smoking and COPD (Case-Control Study)
1
State the Study Parameters
From the literature (prior case-control studies on smoking and COPD in similar populations):
+/- = 0.05 (two-sided) Z+/-/2 = 1.96 Power = 80% Z2 = 0.842 p = 0.30 (30% smoking prevalence in general adult population) Expected OR = 2.5 (smokers have 2.5- the odds of COPD vs non-smokers) r = 1 (equal number of cases and controls)
2
Calculate p (Exposure Prevalence Among Cases)
p = (OR - p) / (1 p + OR - p) p = (2.5 - 0.30) / (1 0.30 + 2.5 - 0.30) p = 0.75 / (0.70 + 0.75) p = 0.75 / 1.45
p = 0.517 (51.7% of COPD cases expected to be smokers)
3
Calculate p (Weighted Average Exposure)
p = (p + r - p) / (1 + r) p = (0.517 + 1 - 0.30) / (1 + 1) p = 0.817 / 2
p = 0.409
4
Calculate the Numerator Term 1 (Z+/-/2 component)
Term = Z+/-/2 - ((1 + 1/r) - p(1 p)) Term = 1.96 - ((1 + 1) - 0.409 - 0.591) Term = 1.96 - (2 - 0.2417) Term = 1.96 - 0.4834 Term = 1.96 - 0.6953
Term = 1.363
5
Calculate the Numerator Term 2 (Z2 component)
Term = Z2 - (p(1 p) + p(1 p)/r) Term = 0.842 - (0.517 - 0.483 + 0.30 - 0.70 / 1) Term = 0.842 - (0.2497 + 0.2100) Term = 0.842 - 0.4597 Term = 0.842 - 0.6780
Term = 0.571
6
Calculate the Denominator
Denominator = (p p)2 Denominator = (0.517 0.30)2 Denominator = (0.217)2
Denominator = 0.04709
7
Calculate Final Sample Size
n = (Term + Term)2 / Denominator n = (1.363 + 0.571)2 / 0.04709 n = (1.934)2 / 0.04709 n = 3.740 / 0.04709
n = 79.4 round up to 80 cases
Controls: 80 - 1 = 80 controls. Add 15% for non-response: 80 / 0.85 = 95 cases, 95 controls.
95 Cases + 95 Controls = 190 Total
Minimum sample size: 80 cases + 80 controls (160 total) before non-response correction
Recommended recruitment target: 95 cases + 95 controls = 190 total (including 15% buffer for exclusions and non-response)

How the Expected Odds Ratio Drives Sample Size

The expected OR is the single most consequential parameter in your calculation. Even small changes in the anticipated OR produce dramatic changes in the required sample size. The following table shows required case numbers for varying ORs, holding all other parameters constant (p = 0.25, +/- = 0.05, power = 80%, r = 1):

Expected ORp (cases)Cases RequiredControls RequiredTotal N
1.30.302 7867861,572
1.50.333 310310620
2.00.400 9595190
2.50.455 5151102
3.00.500 323264
4.00.571 181836
5.00.625 121224

This table illustrates a critical point: overestimating the expected OR leads to a dangerously underpowered study. A researcher who anticipates OR = 3.0 but the true OR is 1.5 will recruit 32 cases when 310 are needed and will produce a non-significant result with 90% probability, regardless of whether the exposure truly is a risk factor.

Examples from Epidemiology Research

Example 1 Diabetes Mellitus and Chronic Kidney Disease

Context: A nephrology researcher at a tertiary hospital designs a case-control study to quantify the association between Type 2 diabetes and early-stage CKD (eGFR 3059 mL/min/1.73 m2). Cases are incident CKD patients from the nephrology outpatient clinic; controls are age- and sex-matched patients from general internal medicine without kidney disease.

Parameters (from regional epidemiology literature): Diabetes prevalence in the general adult population = 12.5% (p = 0.125). Prior studies suggest OR 2.8. +/- = 0.05, power = 80%, r = 2 (two controls per case, as CKD cases are limited).

Derived p: (2.8 - 0.125) / (1 0.125 + 2.8 - 0.125) = 0.35 / (0.875 + 0.35) = 0.35 / 1.225 = 0.286

Result: Kelsey formula yields n 63 cases, 126 controls = 189 total. Adding 15% buffer 74 cases + 148 controls = 222 total participants. The 2:1 ratio compensates for limited case availability while maintaining 80% power.
Example 2 HPV Infection and Cervical Cancer

Context: A gynecology oncology department investigates the association between high-risk HPV (types 16/18) and cervical cancer (squamous cell carcinoma) in a population with limited prior vaccination. Cases are confirmed incident cervical cancer patients; hospital-based controls are women attending the same hospital for non-gynecological complaints.

Parameters: HPV prevalence among control women = 18% (p = 0.18, from prior cervical screening data). Published case-control studies in similar populations report OR 4.06.0; the researcher conservatively uses OR = 4.0. +/- = 0.05, power = 90%, r = 1.

Derived p: (4.0 - 0.18) / (1 0.18 + 4.0 - 0.18) = 0.72 / (0.82 + 0.72) = 0.72 / 1.54 = 0.468

Result: With 90% power (Z2 = 1.282), the formula yields n 38 cases + 38 controls = 76 total. Adding 20% buffer (higher non-response expected for cancer cases) 46 cases + 46 controls = 92 total. The strong OR (4.0) drives a remarkably small required sample but only if the OR assumption is correct. The researcher notes this in the study limitations.
Example 3 Antibiotic Use and Clostridioides difficile Infection

Context: An infectious disease researcher designs a hospital-based case-control study investigating whether fluoroquinolone exposure (5 days in the preceding 3 months) is associated with C. difficile infection (CDI). Cases are confirmed CDI patients identified from the microbiology lab. Controls are hospitalized patients without CDI from the same wards during the same time period.

Parameters: Fluoroquinolone use among hospitalized non-CDI controls 22% (p = 0.22, from hospital antibiotic stewardship records). Prior studies: OR 3.2. +/- = 0.05, power = 80%, r = 2.

Derived p: (3.2 - 0.22) / (1 0.22 + 3.2 - 0.22) = 0.704 / (0.78 + 0.704) = 0.704 / 1.484 = 0.474

Result: n 42 cases, 84 controls = 126 total. Adding 10% buffer 47 cases + 94 controls = 141 total. The 2:1 ratio is appropriate here because CDI cases are relatively rare in any given time window, while non-CDI hospitalized controls are abundant. This design efficiently captures the required power without unnecessarily constraining case recruitment.

Matched vs Unmatched Case-Control Studies: Different Formulas

The Kelsey formula applies to unmatched (or frequency-matched) case-control studies. When individual matching is used each case is paired with one or more specific controls matched on age, sex, or other confounders a different formula is required.

For a 1:1 individually matched case-control study, the McNemar-based formula is appropriate:

McNemar Formula 1:1 Matched Case-Control Study
n = (Z+/-/2 + Z2)2 / (p p)2 - (p + p)
n = Number of matched pairs required
p = Probability case exposed, control unexposed
p = Probability case unexposed, control exposed
OR = p / p (rearranged from this ratio)
When to Use Matched Designs Individual matching is appropriate when there are strong confounders (age, sex, ethnicity) that must be controlled and when the sample is small enough that statistical adjustment alone would be unreliable. Matched designs are more efficient the same power is achieved with fewer participants but the matching must be accounted for in analysis using conditional logistic regression, not ordinary logistic regression. Ignoring the matching in the analysis undermines the entire design.

Understanding Confidence Intervals in Case-Control Sample Size

The confidence level in the sample size formula (typically 95%, corresponding to +/- = 0.05) determines the width of the confidence interval around the estimated odds ratio in your final results. A 95% CI means that if the study were repeated 100 times under identical conditions, 95 of the resulting intervals would contain the true population OR.

Wider confidence intervals (from small samples) include more plausible values for the OR and therefore provide less precise estimates of the true association. A study reporting OR = 2.4, 95% CI [0.87.3] is virtually uninterpretable in clinical terms the true OR could be anywhere from a weak protective to a strong harmful association. Adequate sample size narrows the CI to a range that permits confident clinical interpretation, such as OR = 2.4, 95% CI [1.44.1].

Precision vs Power: Two Different Concepts Power determines whether you can detect that an association is statistically significant. Confidence interval width determines how precisely you can estimate the magnitude of that association. A study can have adequate power (correctly detect that OR 1) while still producing a wide CI (imprecise estimate of the OR value). Some researchers calculate sample size based on desired CI width rather than power this approach is called precision-based sample size estimation and typically requires larger samples than power-based methods for the same OR.

Common Mistakes During Sample Size Estimation for Case-Control Studies

Mistake 1: Overestimating the Expected Odds Ratio

The most damaging error. Researchers routinely select the highest OR they can find in the literature or the OR from the single most optimistic study to justify a smaller, more feasible sample. A study powered for OR = 3.0 when the true OR is 1.8 will return a non-significant result with approximately 75% probability. This is called "winner's curse" bias, and it is rampant in medical thesis sample size sections.

Fix: Use the median OR from systematic reviews or meta-analyses of comparable populations, not the maximum OR from any single study. If published estimates span a wide range (OR 1.54.0), choose a conservative value (OR 1.52.0) and acknowledge in your proposal that you are powering for the lower bound. Conduct a sensitivity analysis showing sample sizes required for each plausible OR scenario.

Mistake 2: Using the Wrong Formula for Matched Designs

Many researchers apply the Kelsey (unmatched) formula to an individually matched study design or use the McNemar formula for an unmatched study. Beyond the formula mismatch, they then analyze matched data using ordinary logistic regression instead of conditional logistic regression, which produces biased odds ratios and incorrect p-values. The design, sample size formula, and analysis method must all be internally consistent.

Fix: Decide whether your design is matched or unmatched before calculating sample size. For unmatched designs: use the Kelsey formula and analyze with unconditional logistic regression. For individually matched designs: use the McNemar formula and analyze with conditional logistic regression. Document the matching scheme and analytical approach in your methods section.

Mistake 3: Ignoring Non-Response and Exclusions

The calculated sample size is the minimum needed for adequate power it assumes complete data from every participant. In reality, some recruited participants will decline consent, fail eligibility screening, have missing data, or withdraw. Failing to add a buffer for expected attrition means the actual analyzable sample falls below the required number.

Fix: Inflate the calculated sample by the anticipated non-response rate: adjusted n = calculated n / (1 non-response rate). For hospital-based studies: add 1015%. For community-based studies or studies requiring multiple clinic visits: add 2025%. For studies with complex inclusion/exclusion criteria: review prior similar studies' screening-to-enrollment ratios and apply accordingly. Always state the non-response adjustment explicitly in your methods section.

Mistake 4: Using Exposure Prevalence from an Unrepresentative Population

The exposure prevalence among controls (p) must reflect the true background exposure in the source population from which your cases arise. Using a national smoking prevalence figure of 22% when your study population is a rural Egyptian governorate with 40% smoking prevalence among males can result in substantial underestimation of the required sample.

Fix: Source p from local or regional epidemiological surveys, national registries, or hospital-based studies in the same demographic. Where no local data exists, conduct a brief pilot survey or use a range of p values in a sensitivity analysis. State the source of your p estimate explicitly and justify its applicability to your study population.

Mistake 5: Calculating Sample Size for the Primary Outcome Only

A case-control study often tests multiple exposure-outcome associations as secondary objectives. Each analysis has its own power requirement. A study adequately powered for the primary OR (2.5) may be severely underpowered for a secondary exposure with OR 1.5. Reporting non-significant secondary findings as meaningful "negative results" from an underpowered sub-analysis is a common and serious methodological error.

Fix: Either power the study for the smallest clinically meaningful effect across all planned analyses (conservative approach, larger n), or explicitly acknowledge in your thesis which secondary analyses are exploratory and underpowered. Apply a Bonferroni correction or adjust the alpha level when testing multiple hypotheses to maintain overall Type I error control.

Mistake 6: Presenting the Sample Size Calculation Without a Literature Reference

Thesis committees and journal reviewers will ask where the OR, p, and other inputs came from. "Assumed" or "estimated" without citation is unacceptable. Every parameter in the sample size formula must be traceable to a published source or explicitly justified as a clinically meaningful threshold.

Fix: For each input parameter, cite the specific publication and state the relevant value extracted from it. Example: "The expected OR of 2.4 was derived from the case-control study of Al-Shafi et al. (2022) conducted in a comparable Egyptian hospital population. The exposure prevalence among controls (p = 0.18) was obtained from the Egyptian Demographic and Health Survey 2021." This level of specificity is expected in any peer-reviewed submission or thesis committee review.

Frequently Asked Questions

What is the minimum sample size for a case-control study? +
There is no universal minimum the required sample depends entirely on your specific parameters: the expected odds ratio, the exposure prevalence among controls, the desired power (typically 8090%), and the alpha level (typically 0.05). Very strong associations (OR 3.0) with a common exposure (prevalence 3040%) may require as few as 3060 cases. Weak associations (OR 1.31.5) with a rare exposure (prevalence 510%) may require 500 or more cases per group. Always calculate sample size formally rather than using a rule of thumb, and always add 1020% to the final figure to account for non-response, exclusions, and missing data.
What is the Kelsey formula for case-control sample size? +
The Kelsey formula (Kelsey et al., 1996) is the most widely used method for unmatched case-control sample size estimation. The number of cases required is: n = [Z+/-/2 - ((1+1/r) - p(1-p)) + Z2 - (p(1-p) + p(1-p)/r)]2 / (p p)2. Where p = exposure prevalence in controls; p = expected exposure in cases, derived from the OR as p = (OR - p)/(1 p + OR - p); p = weighted average = (p + r-p)/(1+r); r = ratio of controls to cases; Z+/-/2 = 1.96; Z2 = 0.842. Multiply n by r to get the number of controls and add 1015% for expected non-response.
How does the odds ratio affect sample size in case-control studies? +
The odds ratio is the single most influential parameter. As the expected OR increases, the required sample decreases dramatically. With exposure prevalence among controls of 25% and 80% power at +/- = 0.05: an OR of 1.5 requires ~310 cases; OR of 2.0 requires ~95 cases; OR of 3.0 requires ~32 cases; OR of 4.0 requires ~18 cases. This means studies investigating weak associations (OR < 2.0) demand substantially larger samples and are most vulnerable to underpowering if the OR is overestimated during planning.
What is the difference between matched and unmatched case-control sample size calculations? +
Unmatched case-control studies use the Kelsey formula, treating cases and controls as independent groups. Matched designs (1:1 individual matching) require the McNemar formula: n = (Z+/-/2 + Z2)2 - (p10 + p01) / (p10 p01)2, where p10 is the probability a case is exposed but their matched control is not, and p01 is the reverse. Matched designs are generally more efficient (smaller sample for same power) because matching removes confounding variation. However, the matching must be accounted for in analysis using conditional logistic regression ignoring it in an ordinary logistic regression model produces incorrect odds ratios and undermines the entire design benefit.
How much power should I use for a case-control study? +
The conventional minimum is 80% power (2 = 0.20), meaning a 20% chance of missing a real effect. Most grant-funded clinical and epidemiological research targets 80% or 90% power. Power of 90% requires approximately 30% more participants than 80% power for the same effect size and alpha. For studies where missing the association could have serious public health consequences identifying a risk factor for a severe or fatal disease 90% power is strongly preferred. The power choice should be explicitly stated and justified in your methods section and ethical approval application.
Can I use G*Power or OpenEpi for case-control sample size calculation? +
Yes. G*Power (free, gpower.hhu.de) is widely used select Test family = z-test; Statistical test = Proportion: Inequality, Two Groups; enter alpha, power, p, and p (derived from your OR). OpenEpi (free, openepi.com) has a dedicated case-control sample size module that uses the Kelsey and Fleiss formulas directly and allows you to enter the OR and p as inputs. EpiInfo (free, CDC) also includes case-control sample size under the StatCalc module. Always verify the formula underlying your chosen tool (matched vs unmatched) and confirm the result against a manual calculation for at least one scenario before trusting it for your thesis submission.

Need Help With Your Sample Size Calculation?

StatClinic's AI Statistical Assistant helps you identify the right formula, input the correct parameters, and generate a sample size that will satisfy your thesis committee and ethics board in minutes, for free.

Try StatClinic AI Statistical Assistant