Analyze My Study
Medical Statistics

How to Interpret P Value in Medical Research (With Practical Examples)

- 13 min read ... June 2025 Updated June 2025
S
StatClinic Editorial Team Statistical content for medical researchers and clinicians
The p-value is the most cited number in medical research and the most consistently misread. It appears in every randomized trial, every observational study, and every thesis defense. Journals use it as a gate for publication. Clinicians use it to decide whether an intervention works. And yet a landmark survey found that the majority of researchers, including those who regularly publish, cannot state its correct definition. This guide will fix that. We cover what the p-value actually measures, why the p < 0.05 threshold is far less meaningful than most researchers assume, and how to extract genuine insight from a p-value rather than just a binary verdict.

What Is a P Value?

The p-value was formalized by statistician Ronald Fisher in the 1920s as part of his null hypothesis significance testing framework. Its formal definition requires understanding one concept first: the null hypothesis (H).

The null hypothesis is the assumption of no effect that your treatment and control groups come from the same population, that the intervention made no difference, that the correlation you observed is zero. The p-value is then calculated by asking: if the null hypothesis were completely true, what is the probability of observing data at least as extreme as what we actually observed?

P = P(data | H is true)
The probability of obtaining results this extreme or more extreme, assuming the null hypothesis is true.
It is NOT the probability that the null hypothesis is true or false.

That definition seems technical, but it has a concrete implication: the p-value is a statement about your data given an assumed world, not a statement about the truth of your hypothesis. A low p-value does not mean your hypothesis is correct. It means your data would be very unlikely to occur in a world where your treatment had no effect at all.

Concrete analogy Imagine flipping a coin 100 times and getting 73 heads. You want to know if the coin is biased. The null hypothesis is "the coin is fair (50/50)." The p-value asks: if the coin really were fair, how often would we see 73 or more heads in 100 flips? The answer (p 0.0001) tells us this would happen less than 0.01% of the time under a fair coin strong evidence the coin is not fair. It does not prove it, but it makes the fair-coin hypothesis very hard to maintain.

Why Researchers Misunderstand P Value

In 2016, the American Statistical Association published an official statement on p-values an unprecedented move that reflected how widespread misinterpretation had become. Studies have shown that misunderstanding extends into published literature, with errors appearing in journals ranging from student theses to major clinical trials.

The root of the problem is that the p-value answers a question researchers did not actually ask. Researchers want to know: "Is my hypothesis true?" The p-value answers: "If my hypothesis were false, would my data look like this?" These are very different questions, and the cognitive leap between them creates a predictable set of misconceptions.

What researchers believe

"p = 0.03 means there is a 3% probability that our result is due to chance."

What it actually means

"If there were truly no effect, we would see results this extreme 3% of the time."

What researchers believe

"p = 0.03 means there is a 97% probability that the treatment is effective."

What it actually means

"The p-value says nothing about the probability that the treatment is effective. That requires Bayesian methods and prior probability estimates."

What researchers believe

"p = 0.06 means the treatment does not work the result is non-significant."

What it actually means

"p = 0.06 means the evidence against H did not reach the arbitrary 0.05 threshold. The study may simply be underpowered. No conclusion about the absence of effect is justified."

What researchers believe

"p = 0.001 means the treatment effect is large and clinically important."

What it actually means

"p = 0.001 in a trial of 10,000 patients might reflect a mean difference of 0.3 mmHg in blood pressure. Small p-values in large samples say nothing about effect magnitude."

What Does P < 0.05 Mean?

The 0.05 threshold is the most influential number in the history of medical research and one of the most misunderstood. Its origin is often traced to Fisher himself, who wrote in 1925 that a result could be considered "significant" if it would occur by chance fewer than 1 in 20 times. This was intended as a rough practical guideline, not a universal law. Fisher himself later clarified that the threshold should vary by context and that p-values should be reported exactly, not categorized as pass/fail.

What p < 0.05 means in practice: if the null hypothesis is true (no real treatment effect), there is less than a 5% probability of observing data as extreme as yours. It is a threshold for statistical surprise, not a threshold for truth. Cross it and the convention says you can reject the null hypothesis. Fail to cross it and you do not reject it but you have not proven it either.

P Value
Statistical Interpretation
How to Describe
p < 0.001
Result is extremely unlikely under H; very strong evidence against null
Very Strong Evidence
p = 0.0010.01
Strong evidence against null hypothesis; highly significant by convention
Strong Evidence
p = 0.010.05
Meets conventional significance threshold; evidence against H present
Significant
p = 0.050.10
Below threshold but worth noting; may reflect insufficient power
Borderline
p > 0.10
Insufficient evidence to reject H; does NOT prove no effect exists
Not Significant
The 0.05 threshold is arbitrary Ronald Fisher himself wrote that "no scientific worker has a fixed level of significance at which from year to year, and in all circumstances, he rejects hypotheses." The 0.05 convention stuck not because it is scientifically correct, but because it was printed in the most widely used statistics textbook of the 1930s. Several leading journals including the New England Journal of Medicine now explicitly discourage binary significant/non-significant language in favour of exact p-values with effect sizes and confidence intervals.

Statistical Significance vs Clinical Significance

This is the most important distinction you can make in medical research, and it is the one most often collapsed into a single yes/no interpretation. Statistical significance tells you whether an observed effect is larger than what random variation would typically produce. Clinical significance tells you whether that effect is large enough to matter for patient care, treatment decisions, or health policy.

These two dimensions are entirely independent. A result can be:

The reason statistical significance can mislead is sample size. In a trial with 15,000 participants, even a difference so small it would never influence a prescribing decision can produce p < 0.001. The mathematics of the t test or chi-square guarantee this: as n grows, the standard error shrinks, and even tiny signal-to-noise ratios become "detectable." Meanwhile, a properly designed pilot study showing a 15 mmHg blood pressure reduction might return p = 0.07 simply because it enrolled 32 patients instead of 120.

What to report instead The 95% confidence interval is more informative than the p-value for clinical interpretation. A mean blood pressure reduction of 12 mmHg (95% CI: 18 to 6) tells you the most likely size of the effect and the plausible range. The p-value simply confirms whether zero is inside or outside that interval. Always report both and always interpret the confidence interval in the context of what is clinically meaningful.

Practical Medical Research Examples

Example 1: A Clear Significant Result

Antihypertensive Drug Trial - n = 120 - Parallel-Arm RCT

A cardiologist randomizes 120 hypertensive patients to either amlodipine 5 mg (treatment group, n = 60) or placebo (control group, n = 60). Systolic blood pressure (SBP) is measured after 8 weeks.

Treatment group: Mean SBP = 136 +/- 12 mmHg. Control group: Mean SBP = 158 +/- 14 mmHg. Mean difference = 22 mmHg (95% CI: 27 to 17). Independent t-test: t(118) = 8.5, p < 0.001, Cohen's d = 1.7.

Interpretation: The p-value is far below 0.05, and critically the effect size is large (d = 1.7) and the confidence interval does not include zero. This result is both statistically and clinically significant. A 22 mmHg reduction in SBP is a treatment-level effect that would change clinical decisions. The p-value and the clinical magnitude agree.

Example 2: Statistically Significant but Clinically Trivial

HbA1c Trial - n = 8,000 - Large Parallel-Arm RCT

A large multinational trial randomizes 8,000 patients with Type 2 diabetes to a new oral antidiabetic agent (n = 4,000) versus standard care (n = 4,000). Primary outcome: HbA1c at 6 months.

Treatment group: Mean HbA1c = 7.42%. Control group: Mean HbA1c = 7.51%. Mean difference = 0.09% (95% CI: 0.14 to 0.04). t(7998) = 3.6, p = 0.0003, Cohen's d = 0.08.

Interpretation: Highly statistically significant, but the effect is clinically meaningless. A 0.09% reduction in HbA1c would not influence any treatment decision and falls far below the 0.30.5% threshold considered clinically meaningful in diabetes research. The enormous sample size generated statistical significance for a trivially small effect. Reporting only "p = 0.0003" without the mean difference or effect size would be actively misleading.

Example 3: Non-Significant but Clinically Important

Pilot Mortality Study - n = 48 - Underpowered RCT

An intensive care physician runs a small pilot RCT testing a new sepsis management protocol in 48 patients (treatment n = 24, control n = 24). Primary outcome: 28-day mortality.

Treatment group: Mortality = 5/24 (20.8%). Control group: Mortality = 9/24 (37.5%). Absolute risk reduction = 16.7 percentage points (95% CI: 6.2 to +39.6). Chi-square: 2(1) = 2.4, p = 0.12.

Interpretation: Not statistically significant at p = 0.12. But the 95% CI includes a potential absolute risk reduction of up to 39.6 percentage points a mortality benefit that would be one of the largest ever documented in sepsis research. The study is simply underpowered. A negative p-value here does not mean the protocol is ineffective. It means the sample size was too small to rule out chance. A larger confirmatory trial of ~300 patients is warranted. Stopping here would be a research error.

Example 4: Borderline P Value What to Do

Antibiotic Resistance Study - n = 84

A microbiologist compares antibiotic resistance rates between two hospital wards. Ward A: 31% resistant (n = 42). Ward B: 17% resistant (n = 42). Difference = 14 percentage points. Chi-square: 2(1) = 2.8, p = 0.051.

Interpretation: The result sits at p = 0.051 one unit above the conventional threshold. The instinctive response is to declare "not significant." The correct response is to report the exact p-value, the absolute risk difference, and its 95% confidence interval, and to note that the study may be underpowered for a 14-percentage-point difference at this sample size. Whether 14% more resistance in Ward A is clinically important is an infection control judgment, not a statistical one. Never interpret p = 0.051 as "no association." Never adjust your methodology post hoc to reach p = 0.049.

Common Interpretation Mistakes

These errors appear in student dissertations, conference presentations, and published papers alike. Recognizing them will make you a sharper reader of the literature and a more credible researcher.

Mistake 1: Treating p < 0.05 as proof of an effect

A p-value below 0.05 means the result is statistically unusual under the null hypothesis not that the treatment definitively works. Type I errors (false positives) occur 5% of the time by definition. In a field publishing thousands of papers per year, a meaningful proportion of "significant" findings are false positives, particularly in small, single-centre studies.

... Fix: Treat significant results as evidence that warrants further investigation, not as proof. Look for effect size, biological plausibility, consistency with prior work, and replication.

Mistake 2: Treating p > 0.05 as proof that there is no effect

A non-significant result is frequently described as "the treatment had no effect" or "no difference was found." This is a logical error known as accepting the null hypothesis. A non-significant p-value means only that the data did not provide enough evidence to reject H not that H is true.

... Fix: Write "the study did not detect a statistically significant difference" not "there is no difference." Examine the confidence interval to determine whether a clinically important effect can be ruled out given your sample size.

Mistake 3: P-hacking running multiple tests until significance appears

Testing 20 different outcomes, subgroups, or time points without pre-specification virtually guarantees that at least one will return p < 0.05 by chance alone (1 in 20 false positives at alpha = 0.05). This practice whether deliberate or naive is a major driver of irreproducible research in medicine.

... Fix: Register your primary and secondary outcomes before data collection. Apply Bonferroni or FDR correction when conducting multiple comparisons. Report all tests performed, not only the significant ones.

Mistake 4: Ignoring effect size and reporting only the p-value

"The new drug significantly reduced cholesterol (p = 0.001)" is an incomplete and potentially misleading statement. Without the mean reduction, confidence interval, and effect size, the reader cannot judge whether the result is clinically meaningful or powered-up noise.

... Fix: Always pair the p-value with the mean difference (or relative risk / odds ratio for binary outcomes), 95% CI, and effect size (Cohen's d, eta-squared, or Cram(c)r's V depending on the test). This triad p value, effect size, CI gives a complete picture.

Mistake 5: Comparing p-values to rank treatment effects

Researchers sometimes write "Drug A was more effective than Drug B because its p-value was smaller (p = 0.001 vs p = 0.04)." This is invalid. The p-value is influenced by sample size: a study with 5,000 patients will produce smaller p-values than one with 50 patients, regardless of which drug works better.

... Fix: Compare treatment effects using mean differences, effect sizes, or head-to-head trials not p-values. A study with p = 0.001 and n = 10,000 may reflect a smaller real-world effect than one with p = 0.04 and n = 30.

How to Report P Value in Research Papers

Correct reporting of p-values is a journal requirement and an ethical standard. The following rules reflect APA 7th edition standards, CONSORT guidelines for clinical trials, and the current consensus from the American Statistical Association.

1

Report exact p-values, not inequalities

Write p = 0.032, not p < 0.05. Write p = 0.008, not p < 0.01. The convention of "p < 0.05" conceals whether you actually found p = 0.049 or p = 0.00001 information that matters to readers and reviewers.

2

Use p < 0.001 for very small values only

When your software returns p = 0.0000 or p < 0.0001, the conventional reporting is p < 0.001. Never write p = 0.000 this implies impossible certainty. p < 0.001 is both accurate and conventional.

3

Always report the full statistical result, not just the p-value

Pair the p-value with the test statistic, degrees of freedom, mean difference or effect measure, and 95% confidence interval. This is now a standard requirement for most high-impact journals.

4

Avoid binary language in the results section

Do not write "the result was significant" without context. Write "the treatment group showed a significantly greater reduction in SBP (mean difference 22 mmHg, 95% CI 27 to 17, p < 0.001, Cohen's d = 1.7)." The number does the work the word "significant" is secondary.

APA 7th Edition Reporting Examples

T test: t(118) = 8.52, p < 0.001, d = 1.71, 95% CI [27.2, 16.8]
ANOVA: F(2, 87) = 18.4, p < 0.001, -2 = 0.30
Chi-square: 2(1, N = 84) = 2.80, p = 0.094, = 0.18
Borderline result: t(46) = 1.98, p = 0.053, d = 0.58, 95% CI [0.2, 18.4] the study may be underpowered for this effect size

Frequently Asked Questions

Does a p value of 0.001 mean the treatment definitely works? +
No. A p value of 0.001 means that if the null hypothesis were true (no real treatment effect), there is only a 0.1% probability of observing results this extreme or more extreme by chance. It does not guarantee the treatment works, and it says nothing about the size of the effect. A very large trial can produce p = 0.001 for a clinically meaningless difference such as a 0.3 mmHg reduction in blood pressure. Always report the mean difference and 95% confidence interval alongside the p value to convey the actual magnitude of the effect. A small p-value in a large sample is expected, not inherently impressive.
Is p = 0.051 really different from p = 0.049? +
Statistically, no not in any practically meaningful sense. The 0.05 threshold is an arbitrary convention, not a law of nature. A result with p = 0.051 is not meaningfully different from p = 0.049. The American Statistical Association and many leading journals explicitly discourage binary "significant/non-significant" categorizations and instead request exact p values alongside effect size and confidence intervals. If your result is p = 0.051 with a clinically important effect size and a confidence interval that barely crosses zero, that is far more interesting than p = 0.049 with a tiny effect. Report the number and let the clinical context lead the interpretation.
Can a non-significant p value (p > 0.05) mean there is no effect? +
Absolutely not and treating it this way is one of the most common and consequential errors in medical research. A non-significant p value means the data did not provide sufficient evidence to reject the null hypothesis. It does not prove the null hypothesis is true. A study may return p = 0.12 simply because it was underpowered the sample was too small to reliably detect a real effect that exists. To properly interpret a non-significant result, examine the 95% confidence interval: if it includes effect sizes that would be clinically important, you cannot rule out a meaningful treatment effect. The honest conclusion is: "We did not detect a significant difference. A larger study would be needed to confirm or exclude a clinically meaningful effect of this magnitude."
What is the difference between statistical significance and clinical significance? +
Statistical significance (p < 0.05) means the result is unlikely to be due to chance alone. Clinical significance means the result is large enough to matter to patients, clinicians, or health systems. These two concepts are completely independent. A trial with 8,000 patients might detect a statistically significant HbA1c reduction of 0.09% real, but so small no endocrinologist would change a treatment based on it. Conversely, a small pilot RCT might show a 20 mmHg blood pressure reduction with p = 0.08 not statistically significant at alpha 0.05, but potentially a very important clinical finding that warrants a properly powered confirmatory trial. Always interpret statistical results in the context of what is clinically meaningful for your specific outcome and patient population.
How should I report p values in a research paper or thesis? +
Always report exact p values rather than inequalities wherever possible write p = 0.032 rather than p < 0.05. For very small values, use p < 0.001 as the conventional lower bound. Never write p = 0.000 this is a rounding artifact in your software, not a real value. Alongside the p value, always report the test statistic (t, F, 2), degrees of freedom, mean difference or effect measure (OR, RR, mean difference), and 95% confidence interval. APA 7th edition format for a t test: t(48) = 3.42, p = 0.001, d = 0.97, 95% CI [8.2, 31.4]. Most high-impact medical journals now require exact p values and effect sizes as a condition of acceptance reporting only "p < 0.05" will result in a revision request.

Final Summary

The p-value is a tool that answers one narrow question: how often would data this extreme arise if there were truly no effect? It does not tell you whether your hypothesis is correct, how large an effect is, or whether a result is clinically meaningful. These are the questions that matter most and they require additional information that the p-value alone cannot provide.

Use the p-value as one component of a complete result: pair it with the mean difference or effect measure, the 95% confidence interval, and the effect size. Interpret the confidence interval against the threshold of clinical relevance for your specific outcome. Resist binary thinking a result of p = 0.051 in a well-designed study with a large effect size carries more scientific weight than p = 0.049 in an underpowered study finding a trivial difference.

Report exact p-values, contextualize every result, and let the clinical interpretation not the number guide your conclusions.

Need help interpreting your statistical results?

Use StatClinic AI Statistical Assistant run your analysis, get APA-formatted output, and receive a written interpretation of your p-values, effect sizes, and confidence intervals.

Analyze My Study