Run This Analysis
Categorical Statistics

Chi-Square vs Fisher's Exact Test in Medical Research: When Should You Use Each One?

- 15 min read ... June 2025 Updated June 2025
S
StatClinic Editorial Team Statistical content for medical researchers and clinicians
Every medical researcher who works with categorical outcomes cure rates, complication rates, mortality, disease presence or absence will eventually face the same decision: Chi-Square test or Fisher's Exact Test? The two tests answer the same fundamental question (is there a statistically significant association between two categorical variables?), but they use completely different mathematical approaches and carry different requirements. Choosing incorrectly is one of the most common methodological errors flagged in peer review. This guide explains exactly how each test works, when one must replace the other, and how to interpret and report your results correctly.

Categorical Data in Medical Research

Categorical data is any variable that falls into distinct, non-overlapping groups. In clinical research, this is the most common type of outcome data you will encounter: a patient either developed a surgical site infection or did not; a drug trial participant either responded to treatment or did not; a histopathology specimen is either positive or negative for a biomarker. When you want to test whether the distribution of these categorical outcomes differs between two or more groups, you need a test designed specifically for categorical data not a t test, not ANOVA.

The two workhorse tests for this purpose are the Chi-Square test of independence and Fisher's Exact Test. Both are applied to data arranged in a contingency table a grid that cross-tabulates two categorical variables and shows the count of observations in each combination of categories.

When do you need these tests? Use Chi-Square or Fisher's Exact when both your grouping variable (treatment vs. control, exposed vs. unexposed) AND your outcome variable (recovered vs. not, complication vs. no complication, positive vs. negative) are categorical. If your outcome is continuous (blood pressure, weight, HbA1c), use t-test or ANOVA instead.

What Is the Chi-Square Test?

The Chi-Square test of independence was developed by Karl Pearson in 1900 and remains the most widely used statistical test for categorical data in medical research. It tests whether the observed distribution of counts across categories differs significantly from what would be expected if the two variables were completely independent of each other.

The test works by calculating an expected frequency for every cell in the contingency table under the null hypothesis of independence, then quantifying how far the observed counts deviate from these expected values.

2 = [ (O E)2 / E ]
O = Observed cell frequency  |  E = Expected cell frequency  |  = Sum across all cells
Expected frequency: E = (Row Total - Column Total) / Grand Total

A large 2 value means the observed counts deviate substantially from what independence would predict suggesting a real association between the two variables. The 2 statistic follows a chi-square distribution with degrees of freedom = (number of rows 1) - (number of columns 1). For a standard 2-2 table, df = 1.

Assumptions of the Chi-Square Test

1

Independence of observations

Each subject contributes to one and only one cell of the table. Chi-Square cannot be used for paired or matched data that requires McNemar's test.

2

Expected cell frequencies 5 in all cells

This is the most critical assumption. The Chi-Square test is based on a mathematical approximation that only holds when expected frequencies are large enough. When any expected frequency drops below 5, the approximation breaks down and the p-value becomes unreliable.

3

Adequate total sample size

For 2-2 tables, a minimum total n of 20 is generally recommended. For larger tables, at least 80% of cells should have expected frequencies 5, and no cell should have an expected frequency of zero.

4

Categorical, not continuous, data

Both variables must be categorical (nominal or ordinal). Do not use Chi-Square for continuous outcomes those require t-tests, ANOVA, or correlation.

What Is Fisher's Exact Test?

Fisher's Exact Test was developed by Ronald Fisher in 1922, originally to analyze the famous "Lady Tasting Tea" experiment. Unlike Chi-Square, which calculates an approximated probability from a continuous distribution, Fisher's Exact Test calculates the exact probability of observing your specific table configuration and all configurations more extreme directly from the hypergeometric distribution. No approximation is involved, which is why the word "exact" appears in its name.

The calculation considers all possible 2-2 tables that could be formed with the same row and column totals as your observed data, then computes the exact probability of obtaining your observed result (or something more extreme) purely by chance. This makes it mathematically precise regardless of how small the expected cell frequencies are.

The key advantage Fisher's Exact Test has no minimum expected frequency requirement. It produces a valid, exact p-value even when expected cell frequencies are below 5, equal to 1, or approaching 0. This makes it the appropriate test for small pilot studies, rare disease research, and any categorical analysis with sparse data.

Assumptions of Fisher's Exact Test

1

Independence of observations

Same as Chi-Square each subject contributes to exactly one cell. Fisher's Exact Test cannot be used for paired data either.

2

Fixed marginal totals

Fisher's Exact Test was originally derived under the assumption that both row and column totals are fixed by design. In practice, this assumption is often relaxed, and the test is widely accepted and appropriate for any 2-2 categorical comparison with small samples.

3

Categorical, not continuous, data

Applies to binary or categorical outcomes only. Best suited to 2-2 tables; extensions to larger tables (Freeman-Halton test) exist but are less commonly used.

Key Differences Between Chi-Square and Fisher's Exact Test

Understanding these differences will make the choice between the two tests straightforward in any research context.

FeatureChi-Square TestFisher's Exact Test
Mathematical basisApproximation via chi-square distributionExact probability via hypergeometric distribution
Expected frequency requirementAll cells must have E 5No minimum requirement
Best for sample sizeLarger samples (n > 20 per group)Small samples, rare events, pilot studies
P-value typeApproximateExact
Table sizeAny size (2-2 to m-n)Best for 2-2; extensions exist for larger
Statistical powerHigher when assumptions are metSlightly more conservative (lower power)
Test statistic reported2, df, p-valuep-value (Fisher's exact); OR, 95% CI
Effect sizeCram(c)r's V or phi ()Odds Ratio (OR) with 95% CI
Software availabilityAll major packages (SPSS, R, SAS, Stata)All major packages
Paired data variantMcNemar's test (not Chi-Square)McNemar's test (not Fisher's)

The Expected Frequency Rule How to Choose

The decision between Chi-Square and Fisher's Exact Test is governed by a single rule that applies universally, regardless of your study design or research question. This rule is based on the expected cell frequencies not the observed frequencies you collected, but the theoretical frequencies you would expect if there were no association at all.

E = (Row Total - Column Total) / Grand Total
Calculate this for every cell in your table. Compare all expected values to the threshold of 5.
Apply this calculation BEFORE running either test.
The universal decision rule Calculate the expected frequency for every cell. If ALL expected frequencies 5 use Chi-Square. If ANY expected frequency < 5 use Fisher's Exact Test. This rule applies regardless of what your observed frequencies look like. A cell could have an observed count of 20 but an expected frequency of 3 that would still require Fisher's Exact.

Step-by-Step Decision Process

1

Set up your contingency table with observed frequencies

Arrange your data in a 2-2 (or larger) table with row totals, column totals, and grand total clearly calculated.

2

Calculate the expected frequency for every cell

E = (Row Total - Column Total) / Grand Total. Do this for all cells not just the ones you suspect might be small.

3

Check every expected frequency against the threshold

If every single expected frequency is 5 or above proceed with Chi-Square. If even one expected frequency is below 5 use Fisher's Exact Test for the entire analysis.

4

Additional check: total sample size

If the total n in your 2-2 table is below 20 (regardless of expected frequencies), use Fisher's Exact Test as an additional safety measure. Very small totals make even technically adequate expected frequencies unreliable.

Contingency Table Examples With Expected Frequencies

Example 1: Chi-Square Is Appropriate

A surgeon investigates whether antibiotic prophylaxis reduces surgical site infections (SSI) in 200 elective abdominal procedures. Patients are randomized to antibiotic prophylaxis (n = 100) or no prophylaxis (control, n = 100).

Observed Frequencies Antibiotic Prophylaxis Trial (n = 200)
Group SSI Developed No SSI Row Total
Antibiotic Prophylaxis 8E = 13.0 92E = 87.0 100
No Prophylaxis (Control) 18E = 13.0 82E = 87.0 100
Column Total 26 174 200
Expected frequencies shown in teal below each observed count. E = (Row Total - Column Total) / 200.
Smallest expected frequency = 13.0 all cells 5 Chi-Square Test is appropriate.
Analysis Chi-Square Appropriate

Expected frequency calculation: E(Antibiotic/SSI) = (100 - 26) / 200 = 13.0. All four expected frequencies equal 13.0 or 87.0 well above the threshold of 5.

Result: 2(1) = 4.88, p = 0.027, = 0.156. Odds Ratio = 0.37 (95% CI: 0.15 to 0.88).

Interpretation: Antibiotic prophylaxis was associated with a statistically significant reduction in SSI rate (8% vs. 18%, 2(1) = 4.88, p = 0.027). The odds of developing SSI were 63% lower in the prophylaxis group (OR = 0.37, 95% CI: 0.150.88). The effect size was small-to-medium ( = 0.156).

Example 2: Fisher's Exact Test Is Required

A pilot study tests a novel immunotherapy in 15 patients with a rare autoimmune condition 8 receive treatment, 7 serve as controls. The outcome is clinical remission at 6 months.

Observed Frequencies Immunotherapy Pilot Study (n = 15)
Group Remission No Remission Row Total
Immunotherapy 5E = 3.73 3E = 4.27 8
Control 2E = 3.27 5E = 3.73 7
Column Total 7 8 15
All four expected frequencies are below 5 (range: 3.274.27). Total n = 15 (below 20).
Chi-Square is NOT valid here Fisher's Exact Test is required.
Analysis Fisher's Exact Required

Expected frequency check: E(Treatment/Remission) = (8-7)/15 = 3.73; E(Treatment/No Remission) = (8-8)/15 = 4.27; E(Control/Remission) = (7-7)/15 = 3.27; E(Control/No Remission) = (7-8)/15 = 3.73. Every single cell falls below 5.

Result: Fisher's exact p = 0.282 (two-tailed). OR = 4.17 (95% CI: 0.4751.3).

Interpretation: The immunotherapy group showed a higher remission rate (62.5% vs. 28.6%), but this difference did not reach statistical significance in this small pilot study (Fisher's exact p = 0.282). The wide confidence interval (OR 0.4751.3) reflects the limited precision of a 15-patient trial. A larger confirmatory trial is warranted the non-significant result does not exclude a clinically meaningful treatment effect.

Example 3: Larger Table 3-2 Contingency

A pathologist categorizes 90 colorectal cancer specimens into three differentiation grades (well, moderate, poor) and compares lymph node metastasis rates across grades. This is a 3-2 contingency table.

3-2 Table Chi-Square for Larger Tables

With a 3-2 table and adequate expected cell frequencies (all 5), Chi-Square remains the correct test. Fisher's Exact Test in its standard form is limited to 2-2 tables the Freeman-Halton extension exists for larger tables but is not available in all software packages and is rarely reported in clinical journals.

Key rule for larger tables: Chi-Square requires that no more than 20% of cells have expected frequencies < 5, and that no cell has an expected frequency of zero. If sparse cells are a problem, consider collapsing adjacent categories (e.g., combining moderate and poor differentiation) only when clinically and biologically justified.

Always state in your methods section the number of cells with expected frequency < 5, and whether this influenced your choice of test. Transparency is a requirement, not an option.

Real Clinical Research Scenarios

Scenario 1 Parallel-Arm RCT With Adequate Sample (Chi-Square)

Randomized Controlled Trial - n = 240

A pneumologist randomizes 240 community-acquired pneumonia patients to azithromycin (n = 120) versus amoxicillin-clavulanate (n = 120). Primary outcome: clinical cure at day 7 (yes/no).

Result: Azithromycin: 98/120 cured (81.7%). Amoxicillin-clavulanate: 91/120 cured (75.8%). All expected cell frequencies exceed 5 (minimum expected = 22.4). Chi-Square is appropriate.

2(1) = 1.32, p = 0.251, = 0.074, OR = 1.43 (95% CI: 0.772.65). No statistically significant difference in cure rates between the two regimens at day 7.

Scenario 2 Small Case-Control Study of Rare Disease (Fisher's Exact)

Case-Control Study - n = 28

An oncologist investigates whether BRCA1 mutation carrier status is associated with triple-negative breast cancer in a pilot case-control study. Cases: 14 triple-negative patients. Controls: 14 hormone-receptor-positive patients.

Observed: Cases: 9 BRCA1+, 5 BRCA1. Controls: 3 BRCA1+, 11 BRCA1. Expected frequency for Controls/BRCA1+ = (14 - 12) / 28 = 6.0. Expected frequency for Controls/BRCA1 = (14 - 16) / 28 = 8.0. Cases/BRCA1+: E = 6.0. Cases/BRCA1: E = 8.0. All 5, but n = 28 total borderline.

With n = 28 and borderline expected frequencies, Fisher's Exact Test is the safer and more defensible choice. Fisher's exact p = 0.037, OR = 6.6 (95% CI: 1.250.1). BRCA1 positivity was significantly more frequent in triple-negative cases though the wide CI reflects the small sample.

Scenario 3 Postgraduate Thesis Comparing Complication Rates (Chi-Square)

Retrospective Cohort - n = 180

A surgical resident compares wound dehiscence rates between laparoscopic (n = 90) and open appendectomy (n = 90) in a retrospective cohort study for their master's thesis. Dehiscence rates: Laparoscopic 3/90 (3.3%), Open 11/90 (12.2%).

Expected frequency check: E(Laparoscopic/Dehiscence) = (90 - 14) / 180 = 7.0. All expected frequencies 5. Chi-Square is valid. 2(1) = 4.91, p = 0.027, = 0.165, OR = 0.24 (95% CI: 0.060.88).

Laparoscopic appendectomy was associated with significantly lower wound dehiscence compared to the open approach (3.3% vs. 12.2%, p = 0.027). Document the expected frequency check in your methods section examiners and reviewers will look for it.

Interpreting the P Value for Chi-Square and Fisher's Exact

The interpretation of the p-value follows the same logic regardless of which test you use. Both tests produce a p-value that represents the probability of observing a table configuration as extreme as or more extreme than the one you obtained, assuming the null hypothesis of no association is true. For a deeper understanding of p-value interpretation beyond categorical tests, read our guide on how to interpret p values in medical research.

Do not stop at the p-value A statistically significant Chi-Square tells you an association exists it says nothing about how strong or clinically meaningful it is. A large study (n = 2,000) can find 2(1) p = 0.003 for an association between two variables where the absolute risk difference is 0.8 percentage points. Always report effect size (phi, Cram(c)r's V) and the Odds Ratio with 95% confidence interval alongside your p-value.

Common Mistakes Researchers Make

Mistake 1: Using Chi-Square without checking expected cell frequencies

The most frequent error in categorical analysis. Researchers run Chi-Square by default without ever calculating expected frequencies then the test returns an invalid p-value that may appear significant even when Fisher's Exact would not be.

... Fix: Before running any Chi-Square test, calculate E = (Row Total - Column Total) / Grand Total for every cell. Document the minimum expected frequency in your methods section. Most software (SPSS, R) will flag this for you if you know where to look.

Mistake 2: Confusing observed and expected frequencies

The decision rule is based on expected frequencies, not observed ones. A cell with 2 observed events could have an expected frequency of 8 (no problem for Chi-Square), or a cell with 15 observed events could have an expected frequency of 3 (Fisher's Exact required). These are different numbers.

... Fix: Never look at your observed cell counts to decide which test to use. Always compute the expected frequencies explicitly using the formula before making your decision.

Mistake 3: Using Chi-Square or Fisher's Exact for paired data

Both Chi-Square and Fisher's Exact Test require independent observations. If the same patient appears in both rows (e.g., comparing a positive test result before and after treatment in the same patients), neither test is appropriate. Using them on paired data produces incorrect results.

... Fix: For paired binary data (same subject provides both measurements), use McNemar's test. This is a common oversight in before-after study designs with binary outcomes.

Mistake 4: Applying Yates' continuity correction and then comparing with uncorrected Chi-Square

Some researchers apply Yates' continuity correction to their 2-2 Chi-Square then compare the p-value to one without correction from a reference study. The corrected test is systematically more conservative, making direct p-value comparisons invalid.

... Fix: Current consensus discourages routine use of Yates' correction. If your expected frequencies are adequate, use standard Chi-Square. If any expected frequency is < 5, use Fisher's Exact Test directly. There is rarely a reason to apply Yates' correction.

Mistake 5: Reporting only "Chi-Square, p = 0.03" without effect size

A significant Chi-Square p-value without an accompanying effect size measure gives the reader no information about the strength of the association. This omission is now a grounds for rejection in most high-impact journals.

... Fix: For 2-2 tables: report the phi coefficient ( = (2/n)) and/or Odds Ratio with 95% CI. For larger tables: report Cram(c)r's V. Interpretation or V: 0.1 = small, 0.3 = medium, 0.5 = large effect.

Mistake 6: Collapsing table categories post hoc to rescue a Chi-Square with small expected frequencies

Merging categories to increase expected cell counts after seeing your data is a form of post-hoc data manipulation that inflates Type I error and must be declared in any honest methods section. If discovered in peer review, it will require reanalysis or rejection.

... Fix: Decide on category structure before data collection. If you anticipate small samples, pre-specify Fisher's Exact as your primary test in your study protocol or thesis proposal.

How to Report Results Correctly

Clear, complete reporting of categorical test results is a requirement at every stage thesis, conference abstract, or peer-reviewed publication. The following formats reflect APA 7th edition and CONSORT-aligned reporting standards.

Chi-Square Test APA Format
2(1, N = 200) = 4.88, p = 0.027, = 0.156
Full statement: "Antibiotic prophylaxis was associated with a significantly lower SSI rate compared to no prophylaxis (8% vs. 18%; 2(1, N=200) = 4.88, p = 0.027, OR = 0.37, 95% CI [0.15, 0.88])."
Fisher's Exact Test APA Format
Fisher's exact p = 0.282 (two-tailed), OR = 4.17, 95% CI [0.47, 51.3]
Full statement: "No statistically significant difference in remission rates was observed between immunotherapy and control groups (62.5% vs. 28.6%; Fisher's exact p = 0.282). The wide confidence interval (OR = 4.17, 95% CI [0.4751.3]) indicates insufficient precision and warrants a larger confirmatory study."
What to include in your methods section State which test you used AND why. Example: "Chi-Square test was used to compare categorical outcomes between groups, as all expected cell frequencies exceeded 5. Fisher's Exact Test was applied where any expected cell frequency was below 5 or total n was less than 20."

Frequently Asked Questions

What is the main difference between Chi-Square and Fisher's Exact Test? +
Chi-Square is based on a mathematical approximation of the sampling distribution and requires adequate sample sizes specifically, all expected cell frequencies must be 5 to be valid. Fisher's Exact Test calculates the exact probability of observing your results using the hypergeometric distribution, without any approximation and without any minimum expected frequency requirement. Chi-Square is appropriate for larger samples where its assumptions are met; Fisher's Exact is the correct choice whenever sample sizes are small or any expected cell frequency falls below 5. Both tests answer the same question: is there a statistically significant association between two categorical variables?
When should I use Fisher's Exact Test instead of Chi-Square? +
Use Fisher's Exact Test whenever any expected cell frequency in your contingency table is less than 5. This rule applies to expected frequencies not the observed counts you collected. Additionally, for any 2-2 table with a total sample size below 20 (regardless of individual expected frequencies), Fisher's Exact is the safer choice. Typical scenarios requiring Fisher's Exact in medical research include: pilot studies (n < 30), rare disease case series, small case-control studies with matched groups, any subgroup analysis within a larger study where cell counts become small, and retrospective analyses of rare complications or adverse events.
How do I calculate expected cell frequencies? +
The expected frequency for any cell is: E = (Row Total - Column Total) / Grand Total. For a 2-2 table with 100 treated patients (8 events) and 100 control patients (18 events): the expected frequency for the 'treated/event' cell = (100 - 26) / 200 = 13.0. Calculate this for every cell before deciding between Chi-Square and Fisher's Exact. In SPSS, the Chi-Square output table automatically shows expected counts look for any value marked as "< 5" in the footnote. In R, use chisq.test()$expected to extract all expected frequencies. The decision must be made based on these expected values, not the observed ones.
Can I use Fisher's Exact Test for a 3-3 or larger contingency table? +
Technically yes the Freeman-Halton extension generalizes Fisher's Exact Test to tables larger than 2-2. However, it is computationally demanding and not available in all statistical software packages. For larger tables (3-2, 3-3, etc.) with small expected frequencies, the most common approaches in clinical research are: (1) collapse categories to a 2-2 table where biologically and clinically justified, then apply standard Fisher's Exact; (2) use Chi-Square and acknowledge in your methods that a proportion of cells had expected frequencies < 5; or (3) consult a biostatistician for a simulation-based or exact permutation approach. The choice must be pre-specified in your protocol, not selected post-hoc.
What is Yates' continuity correction and should I use it? +
Yates' continuity correction modifies the Chi-Square formula by subtracting 0.5 from the absolute difference between observed and expected frequencies before squaring: 2 = [(|O E| 0.5)2 / E]. It was developed to make the Chi-Square approximation more accurate (closer to Fisher's Exact) for small 2-2 tables. However, current statistical consensus generally discourages it because it tends to be overly conservative it can push p-values above 0.05 when the uncorrected Chi-Square is significant, increasing the Type II error rate. The recommendation from most biostatisticians today is: if your expected frequencies are adequate (all 5), use standard Chi-Square. If any expected frequency is below 5, use Fisher's Exact Test not Yates'-corrected Chi-Square.
What effect size should I report alongside Chi-Square or Fisher's Exact? +
For a 2-2 table, report the phi coefficient (): = (2/n). Phi ranges from 0 to 1. Interpretation: 0.1 = small effect, 0.3 = medium effect, 0.5 = large effect. For tables larger than 2-2, report Cram(c)r's V: V = (2 / [n - min(rows1, cols1)]). Alongside the effect size, always report the Odds Ratio (OR) with 95% CI this is the clinically interpretable measure that tells you the magnitude of the association in practical terms. Most high-impact medical journals now require effect size reporting as a standard condition of acceptance, not an optional addition. An OR of 1.0 means no association; CI that excludes 1.0 means statistical significance.

Final Summary

The choice between Chi-Square and Fisher's Exact Test is determined by one rule: calculate the expected frequency for every cell in your contingency table. If all expected frequencies are 5 or above, Chi-Square is valid and appropriate. If any expected frequency falls below 5 or if your total sample size is below 20 use Fisher's Exact Test.

Both tests answer the same question: is there a statistically significant association between two categorical variables? The difference is that Chi-Square relies on an approximation that requires sufficient data to be accurate, while Fisher's Exact Test calculates the probability directly and precisely for any sample size.

Neither test should be reported with the p-value alone. Always accompany your result with the Odds Ratio and 95% confidence interval, and the phi coefficient or Cram(c)r's V as an effect size measure. Document your expected frequency check in your methods section this is not optional, and peer reviewers will look for it. For paired binary data (same subject measured twice), use McNemar's test, not Chi-Square or Fisher's Exact.

Need help choosing the correct statistical test?

Try StatClinic AI Statistical Assistant run Chi-Square, Fisher's Exact, and McNemar tests instantly. Get APA-formatted output with effect sizes, Odds Ratios, and a written interpretation ready for your thesis or paper.

Analyze My Study