What Is Publication Bias?
Publication bias is the tendency for studies with statistically significant, positive, or clinically favorable results to be published more often, more quickly, and in more visible journals than studies with null, negative, or unfavorable results. When a systematic review searches the published literature and pools whatever it finds, it is not sampling from all studies that were ever conducted — it is sampling from the subset that made it to print, which is systematically skewed toward "positive" findings.
This matters enormously for meta-analysis specifically, because the entire method rests on the assumption that the included studies represent a fair, unbiased sample of the available evidence. If the missing studies would have shown a smaller effect, no effect, or even a harmful effect, the pooled estimate from only the published studies will be biased — usually toward overstating the benefit of a treatment or the strength of an association.
Why Publication Bias Occurs
Publication bias arises from decisions made by multiple parties across the research pipeline, not from any single cause.
| Source | Typical Behavior |
|---|---|
| Authors | Less motivated to write up and submit a "boring" null result; may file it away instead ("the file drawer problem") |
| Journal editors and reviewers | May perceive null results as less novel or less interesting, reducing acceptance rates for such submissions |
| Funders and sponsors | Industry-funded studies with unfavorable results for the sponsor's product may be delayed or never submitted for publication |
| Time | Significant results tend to be published faster than null results ("time-lag bias"), so recent reviews can be skewed toward early positive findings |
Why It Matters
A meta-analysis exists specifically to give clinicians and guideline committees a more reliable, precise answer than any single study can provide — but that promise depends entirely on the input evidence being representative. If the published literature systematically overrepresents favorable results, a meta-analysis built on it will confidently report an effect that is larger than the truth, sometimes drastically so, and clinical guidelines built on that meta-analysis inherit the same distortion.
Several well-documented historical cases — including selectively published antidepressant trials and specific cardiovascular drug trials — have shown pooled effects shrink substantially, or even reverse direction, once unpublished or delayed trials were later located and added to the analysis. This is precisely why publication bias assessment is now an expected, not optional, part of a rigorous systematic review.
The consequence is not purely theoretical for a treatment decision either. A clinician reading a meta-analysis with an inflated pooled effect may reasonably conclude a treatment is more effective, or a risk factor more dangerous, than it actually is — leading to overuse of an intervention with a smaller real benefit than believed, underappreciation of its harms relative to that benefit, or misallocated research funding chasing an effect that is partly an artifact of what happened to get published rather than what is clinically true.
Small-Study Effects
"Small-study effects" describes the broader observed pattern where smaller studies in a meta-analysis tend to show systematically larger (and more variable) effect sizes than larger studies. Publication bias is one common cause of this pattern — a small study needs a larger effect size to reach statistical significance than a large study does, so small studies with large effects are disproportionately likely to be published, while small studies with modest or null effects are disproportionately likely to go unpublished.
Crucially, small-study effects are not proof of publication bias on their own — genuine clinical heterogeneity, lower methodological quality in smaller studies, or chance can produce an identical visual pattern. This distinction is revisited in the misconceptions section below, because conflating the two is one of the most common errors in interpreting a funnel plot.
A useful way to think about it: publication bias is a bias in which studies you get to see at all, while some other small-study effect causes (like smaller studies more often being conducted with less rigorous methodology, or in higher-risk patient subgroups) are biases in the studies themselves, present regardless of whether they were published. Both produce the same funnel plot pattern, which is exactly why a funnel plot alone cannot tell you which explanation applies to your specific meta-analysis without further investigation.
Funnel Plot Interpretation
A funnel plot is a scatter plot with each study's effect size on the horizontal axis and a measure of its precision — usually standard error, with larger (more precise) studies plotted near the top and smaller (less precise) studies plotted toward the bottom — on the vertical axis. In the absence of bias, the scatter should resemble an inverted funnel: large, precise studies clustering tightly near the pooled estimate at the top, and smaller, less precise studies scattering more widely but roughly symmetrically around that same central value as you move down.
The Y-Axis: Study Precision
The vertical axis is almost always standard error (sometimes 1/SE, or occasionally sample size), plotted with zero at the top — meaning the most precise, usually largest, studies appear near the top of the funnel, and the least precise, usually smallest, studies appear toward the bottom.
The X-Axis: Effect Size
The horizontal axis is the study's effect size, using whichever measure the meta-analysis pools (log odds ratio, mean difference, standardized mean difference, log hazard ratio). Ratio measures are plotted on a log scale for the same reason covered in our forest plot guide — this keeps the funnel visually symmetric even when effects are naturally skewed on a raw ratio scale.
The Pseudo-Confidence-Interval Lines
The two diagonal lines forming the funnel's outer boundary represent the 95% confidence interval around the pooled estimate at each level of precision — they converge to a point at the top (where precision is highest and the interval is narrowest) and widen as you move down (where precision is lower and the interval is wider). Most points should fall within this funnel-shaped region.
Symmetrical vs Asymmetrical Funnel Plots
A symmetrical funnel plot (shown above) has studies scattered in a roughly mirror-image pattern on both sides of the pooled estimate at every level of precision — this is reassuring, though not conclusive proof of an absence of bias. An asymmetrical funnel plot shows a visible gap on one side, most classically in the bottom corner — small, imprecise studies missing from the side that would show a null or unfavorable result.
Notice the bottom-left region is empty in the second plot — small studies are present on the right (favorable/significant side) but conspicuously absent on the left (null/unfavorable side). This exact pattern is the classic visual signature that prompts a publication bias investigation, though as covered below, visual asymmetry alone is a starting point for further assessment, not a final verdict.
Egger's Test
Egger's test statistically evaluates funnel plot asymmetry by running a linear regression of each study's standardized effect size against its precision, and testing whether the regression line's intercept is significantly different from zero. In plain terms: if small, imprecise studies are systematically showing different (usually larger) effects than large, precise studies, the regression line will be tilted, and Egger's test quantifies how tilted.
A meta-analysis of 16 trials of a supplement for reducing inflammation markers reports Egger's test intercept = 2.84, 95% CI [0.91, 4.77], p = 0.008. The significant p-value and the intercept clearly different from zero suggest meaningful funnel plot asymmetry, consistent with (but not proof of) publication bias or another small-study effect.
Egger's test is generally considered more statistically powerful than Begg's test, but it is more sensitive to the specific effect measure used and can give false positives when there is genuine heterogeneity unrelated to bias — a limitation covered further below. Several modified versions exist for specific effect measures (Harbord's test and Peters' test, for example, are commonly used alternatives to standard Egger's test for binary outcomes such as odds ratios, since Egger's original formulation was developed primarily for continuous outcomes and can behave less reliably with binary data).
Begg's Test
Begg's test (formally, Begg and Mazumdar's rank correlation test) checks whether there is a significant correlation between the ranked effect sizes and the ranked variances of the included studies — if smaller, less precise studies (higher variance) systematically rank toward one extreme of effect size, the correlation will be significant.
The same 16-trial supplement meta-analysis reports Begg's test: Kendall's tau = 0.31, p = 0.06. This falls short of conventional significance despite Egger's test being significant on the same data — a common and expected occurrence, since Begg's test generally has lower statistical power, especially with a moderate number of studies.
Begg's test makes fewer distributional assumptions than Egger's test, which is sometimes cited as an advantage, but its lower power means it more often fails to detect real asymmetry, particularly with fewer than 20 studies.
Trim-and-Fill Method
Trim-and-fill is a method that estimates how many studies appear to be "missing" from the asymmetric side of a funnel plot, then imputes mirror-image studies on the opposite side to restore symmetry, and recalculates the pooled effect including these imputed studies. It provides an adjusted estimate showing what the pooled effect might look like if the suspected missing studies existed and were included.
Original pooled estimate: OR = 0.58, 95% CI [0.47, 0.71]. Trim-and-fill imputes 4 missing studies on the null side of the funnel and recalculates: adjusted OR = 0.67, 95% CI [0.54, 0.83]. The adjusted estimate still favors treatment, but the effect is meaningfully smaller once the suspected missing studies are accounted for — this is reported as a sensitivity analysis alongside, not instead of, the original estimate.
Trim-and-fill's imputed studies are statistical constructs, not real discovered studies — the method should always be presented as a sensitivity check on how robust the conclusion is to plausible missing data, never as if it definitively reconstructed the true unbiased literature. Two variants exist (the L0 and R0 estimators, if your software asks you to choose), which can occasionally impute a noticeably different number of studies on the same dataset — reporting which estimator was used, alongside the software name and version, keeps your sensitivity analysis fully reproducible.
Limitations of These Methods
- Low power with few studies: Both funnel plot visual inspection and formal tests are unreliable with fewer than 10 studies — there simply isn't enough data to distinguish a real asymmetric pattern from chance scatter.
- Cannot distinguish the cause of asymmetry: None of these methods can definitively tell you whether asymmetry is due to publication bias specifically, versus genuine heterogeneity, methodological differences between small and large studies, or chance.
- Sensitive to the choice of effect measure: The same data can appear more or less symmetric depending on which effect measure (OR vs RR, MD vs SMD) is used for the funnel plot.
- Trim-and-fill can both under- and over-adjust: Simulation studies have shown trim-and-fill can add too many or too few imputed studies depending on the true underlying pattern of missingness, which is never actually known.
- No method detects bias within a single study: These techniques assess whether studies appear to be missing from the pooled literature; they say nothing about selective outcome reporting within an individual included study, which requires a separate risk-of-bias assessment tool entirely.
- A perfectly symmetric funnel plot does not guarantee an unbiased result: If an entire body of small, unfavorable studies was suppressed evenly on both sides (unlikely but possible), or if all available studies happen to still look symmetric by chance, the absence of visible asymmetry is reassuring but not a formal guarantee.
Common Misconceptions
"The funnel plot looks asymmetric, so publication bias is proven and the pooled estimate should be discarded."
Asymmetry suggests possible small-study effects, of which publication bias is one plausible cause among several — report it as a limitation and consider trim-and-fill as a sensitivity check, not grounds to discard the analysis.
"Egger's test was not significant, so there is definitely no publication bias in this meta-analysis."
A non-significant test with few studies may simply reflect low statistical power, not genuine absence of bias — interpret a non-significant result cautiously when the study count is small.
"We ran Egger's test on our 6 included studies and it wasn't significant, so bias isn't a concern."
Formal tests are not recommended below roughly 10 studies at all — with only 6 studies, this result is essentially uninformative and should not have been used to draw any conclusion.
Software Examples: RevMan, Stata, and R
Checklist Before Reporting Publication Bias
Confirm you have at least 10 included studies
Before running or reporting any formal statistical test for asymmetry.
Show the funnel plot itself
Never report a formal test result without the visual plot it is based on.
Report both Egger's and Begg's test if feasible
Note any disagreement between them rather than cherry-picking the more favorable result.
Consider trim-and-fill as a sensitivity analysis
Report the adjusted pooled estimate alongside, not instead of, the original.
Discuss alternative explanations for any asymmetry
Genuine heterogeneity and methodological quality differences, not only publication bias.
Describe your literature search strategy
Including any attempt to find unpublished or gray literature (trial registries, conference abstracts), which itself reduces publication bias risk.
State your conclusion in proportion to the evidence
"Suggestive of," not "proves," unless the evidence is genuinely overwhelming.
Frequently Asked Questions
Pool Your Studies With Confidence
Use StatClinic's free meta-analysis calculator to compute your pooled estimate, heterogeneity statistics, and prepare the data behind your funnel plot. Free, no registration required.
Try StatClinic Free →