Open StatClinic →
🌏 Epidemiology Methods

Confounding vs Effect Modification in Medical Research:
Complete Epidemiologist's Guide

🕑 26 min read 📅 July 2026 ✅ Peer-reviewed content 📚 3800+ words
S
StatClinic Editorial Team Statistical content for medical researchers and clinicians
A 2004 observational study appeared to show that post-menopausal hormone replacement therapy (HRT) reduces the risk of coronary heart disease by about 40% in women who use it. When the Women's Health Initiative RCT was published, HRT users showed a 29% increased risk. The contradiction was not a measurement error or a random fluke — it was confounding. Women who chose HRT in observational studies were healthier, wealthier, and more health-conscious than non-users. These characteristics — independently protective against heart disease — were the confounders that made HRT appear beneficial when it was not. This is one of the most consequential confounding errors in the history of medicine. Understanding the difference between confounding (a bias to remove) and effect modification (a finding to report) is not theoretical training — it is the practitioner's most important defence against drawing false conclusions from observational data.

What Is Confounding?

Confounding is a form of bias in observational research in which a third variable — the confounder — distorts the observed association between an exposure and an outcome. The confounder creates the illusion of a causal relationship that does not exist, masks a real relationship, or changes the apparent magnitude of a true one. It is the primary reason that correlation does not imply causation, and why randomised controlled trials hold the highest position in the evidence hierarchy: randomisation distributes all confounders — measured and unmeasured — equally between groups.

The Three Criteria for Confounding

A variable qualifies as a true confounder if and only if it satisfies all three of the following criteria simultaneously:

C₁
Associated with the Exposure
The confounder must be more (or less) prevalent in exposed individuals than in unexposed individuals. If the variable is equally distributed across exposure groups, it cannot confound the exposure-outcome association.
C₂
Independent Risk Factor for the Outcome
The confounder must be independently associated with the outcome — it must have its own causal or predictive relationship with the disease, beyond its relationship with the exposure.
C₃
Not an Intermediate (Mediator)
The confounder must NOT lie on the causal pathway between the exposure and the outcome. A variable on the causal pathway is a mediator — adjusting for it would block the very mechanism of interest, creating different bias. Confounders precede the exposure or are caused by a common upstream variable.

Classic example: In a study of coffee consumption and cardiovascular disease (CVD), smoking is a confounder if: (C₁) coffee drinkers are more likely to smoke than non-coffee-drinkers; (C₂) smoking independently increases CVD risk; and (C₃) smoking does not cause coffee drinking or lie on the causal path from coffee to CVD. All three conditions are met — and indeed, many early studies incorrectly reported coffee as a cardiovascular risk factor before adjusting for smoking.

Positive vs Negative Confounding

Positive confounding inflates the apparent association between exposure and outcome: the crude (unadjusted) OR or RR is larger than the true effect. The HRT-heart disease example is positive confounding — healthier women chose HRT, making HRT appear more protective than it actually was (or was reversed after adjustment).

Negative confounding attenuates or reverses the apparent association: the crude estimate is smaller than the true effect, or the true direction of association is masked. A treatment that is preferentially given to the sickest patients may appear less effective than it truly is because the confounding by indication makes the treated group inherently at higher risk of the outcome.

Detecting Confounding: The 10% Change-in-Estimate Rule

The most practical and widely used method for detecting confounding in observational research is the 10% change-in-estimate rule: a variable is considered a meaningful confounder if including it in the regression model changes the exposure-outcome estimate (OR, RR, regression coefficient) by ≥ 10% relative to the crude estimate.

The 10% Change-in-Estimate Rule

% Change = |ORᵣᵚᵛᵈᵉ − ORᵚᵈʲ‌‌‌| / ORᵚᵈʲ‌‌‌ × 100
If % Change ≥ 10% → variable IS a meaningful confounder → include in final model
If % Change < 10% → variable is NOT a meaningful confounder → may exclude from final model

Example: Crude OR = 2.0, Adjusted OR = 1.5
(2.0 − 1.5) / 1.5 × 100 = 33% change → Meaningful confounder ✓
The 10% rule has limitations: It is a pragmatic heuristic, not a mathematically rigorous criterion. A variable may qualify as a confounder by biological reasoning (satisfying all three criteria) but produce <10% change in a particular sample due to insufficient power or collinearity with another included variable. Always use the 10% rule together with subject-matter knowledge and a DAG-based analysis plan, never as a standalone algorithmic criterion.

Stratified Analysis for Confounding Detection

Before multivariable regression, stratified analysis provides a transparent picture of whether confounding is present. Divide the sample into strata of the potential confounder and compute stratum-specific effect estimates. If the stratum-specific estimates are similar to each other but differ from the crude estimate, confounding is present. The Mantel-Haenszel method provides a weighted pooled estimate that removes the confounding while preserving the association structure across strata.

The key diagnostic pattern for confounding is: crude estimate ≠ adjusted estimate, but stratum-specific estimates are approximately equal to each other. When the stratum-specific estimates differ substantially from each other, that is effect modification — the topic of the next section.

Controlling for Confounding

Confounding can be addressed at two stages: during study design and during statistical analysis.

Design-Stage Strategies

Analysis-Stage Strategies

What Is Effect Modification?

Effect modification — also called statistical interaction or heterogeneity of effect — occurs when the magnitude or direction of the association between an exposure and outcome genuinely differs across strata of a third variable. Unlike confounding, effect modification is not a bias. It is a real biological or clinical phenomenon: the exposure truly has a different effect in one subgroup than another.

The classic epidemiological terminology is: "variable M modifies the effect of exposure E on outcome O." M is the effect modifier. The effect of E on O is heterogeneous across levels of M. This is biologically important — it means that a treatment or risk factor may be beneficial in some patients and harmful (or neutral) in others, which has direct implications for clinical practice and public health targeting.

⚠ Confounding = Bias to Remove

A third variable creates a spurious or distorted association between exposure and outcome. The true effect is the same across strata of the confounder; the crude estimate is simply wrong. Adjustment produces the correct estimate; the adjusted analysis supersedes the crude analysis.

Correct response: Control it — adjust in multivariable model

🔎 Effect Modification = Finding to Report

The true association between exposure and outcome genuinely differs across subgroups. The overall adjusted estimate is a misleading average. Each stratum tells a different biological story. This heterogeneity is a scientific discovery that should drive subgroup-specific clinical guidance.

Correct response: Report it — present stratum-specific estimates

Additive vs Multiplicative Interaction

Effect modification can be assessed on two different scales, and they can give different conclusions — a source of substantial confusion in the literature:

A critical distinction: It is mathematically possible to have additive interaction present while multiplicative interaction is absent (and vice versa). In practice, most epidemiologists now recommend reporting both scales. For clinical decision-making and cost-effectiveness analyses, additive interaction is more actionable because it reflects the absolute excess risk — the number of preventable cases — attributable to the combination.

Directed Acyclic Graphs and Collider Bias

Directed Acyclic Graphs (DAGs) — also called causal diagrams — provide a formal graphical language for representing causal assumptions and identifying which variables to adjust for. A DAG consists of nodes (variables) and directed arrows (causal effects). The absence of an arrow is an explicit causal assumption: the researcher asserts that no direct causal effect exists along that path.

DAG: Confounding vs Effect Modification vs Collider

CONFOUNDING Exposure Confounder Outcome backdoor path EFFECT MODIFICATION Exposure Outcome Modifier modifies → COLLIDER BIAS Exposure Outcome Collider ⚠ Adjusting opens spurious path

The DAG reveals a critical third concept: collider bias. A collider is a variable that has arrows pointing into it from both the exposure and the outcome — it is a common effect of both. Counterintuitively, adjusting for a collider opens a spurious non-causal path between the exposure and outcome, introducing bias. This means that conditioning on variables without causal reasoning (including all available variables "just in case") can make estimates worse, not better. The famous Berkson's bias — where two unrelated diseases appear negatively correlated in a hospital sample because both can lead to hospitalisation — is a collider bias caused by conditioning on the collider "being hospitalised."

DAGs are drawn before analysis based on domain knowledge, not statistical testing. Software such as dagitty.net (free) can identify the minimal sufficient adjustment set and highlight colliders given the assumed causal structure.

Detecting Effect Modification

Effect modification is assessed by two complementary approaches:

1. Stratified Analysis

Compute the exposure-outcome association separately within each stratum of the potential modifier. If the stratum-specific estimates differ in magnitude or direction, effect modification is present. The diagnostic pattern for effect modification is: stratum-specific estimates differ from each other — in contrast to confounding, where stratum-specific estimates agree with each other but differ from the crude estimate.

2. Interaction Term in Regression

Add the product of the exposure and modifier to the regression model and test its coefficient. In logistic regression:

logit(P) = β₀ + β₁ · Exposure + β₂ · Modifier + β₃ · (Exposure × Modifier)

The coefficient β₃ is the interaction term. A statistically significant interaction term (p < 0.05, or more conservatively p < 0.10 given low power for interaction tests) indicates that the effect of the exposure depends on the level of the modifier. The stratum-specific ORs are recovered from the model parameters: for Modifier = 0, OR = e𝚽₁; for Modifier = 1, OR = e𝚽₁⁺𝚽₃.

Low power for interaction tests: Statistical tests for interactions require approximately 4× the sample size needed to detect a main effect of the same magnitude. Interaction tests are underpowered in most clinical studies. A non-significant p-value for the interaction term does NOT confirm the absence of effect modification — it may simply reflect insufficient power. Always report the stratum-specific estimates and their 95% CIs even when the formal interaction test is non-significant, and state that the study was not adequately powered to detect interaction.

Clinical Examples

1
Aspirin and GI Bleeding: Age as an Effect Modifier
A pharmacoepidemiological cohort study examines the association between low-dose aspirin use and upper gastrointestinal (GI) bleeding in 8,400 adults. The research team investigates whether age modifies the aspirin-bleeding association, given that elderly patients are known to have higher baseline GI bleeding risk and altered platelet function.
8,400
Total cohort (n)
3,210
Aspirin users
5,190
Non-users
p = 0.008
Interaction test (age × aspirin)
StratumAge GroupOR (Aspirin vs No)95% CIInterpretation
Crude (overall)All ages1.821.52–2.18No adjustment
Age-adjustedAll ages1.791.49–2.14Age is NOT a confounder (10% rule: 1.7%)
Young stratum< 60 years1.310.96–1.79n.s. (p = 0.09)
Old stratum≥ 60 years2.742.18–3.44p < 0.001
Checking for Confounding vs Effect Modification Age-adjusted OR = 1.79 vs Crude OR = 1.82 % Change = |1.82 − 1.79| / 1.79 × 100 = 1.7% → below 10% threshold Conclusion: Age is NOT a confounder of the aspirin-bleeding association Interaction Test (Logistic Regression) logit(GI bleed) = β₀ + β₁·Aspirin + β₂·AgeGroup + β₃·(Aspirin × AgeGroup) β₃ coefficient: OR = 2.09 (95% CI: 1.21–3.62), p = 0.008 → Statistically significant interaction: age IS an effect modifier Ratio of Stratum-Specific ORs (Multiplicative Interaction) OR(elderly) / OR(young) = 2.74 / 1.31 = 2.09 → same as interaction OR above
Age is an effect modifier (not a confounder) of the aspirin-GI bleeding association. Elderly patients face 2.74× higher bleeding risk with aspirin; younger patients face only 1.31× risk — a non-significant increase.
Clinical implication: The overall (age-adjusted) OR of 1.79 masks the clinically critical fact that most of the bleeding risk is concentrated in those aged ≥ 60. Reporting only the pooled adjusted estimate would lead a clinician to conclude that all patients face approximately equal bleeding risk on aspirin — which is wrong. The correct action is to report stratum-specific ORs separately for younger and older patients and discuss the age-specific risk-benefit calculation for aspirin prescribing. In the elderly, the absolute bleeding risk increase per 100 patients treated would need to be weighed against the absolute cardiovascular events prevented.
2
Antihypertensive Therapy and Stroke: Smoking as a Confounder
A retrospective cohort study investigates whether antihypertensive treatment reduces 5-year stroke risk in hypertensive adults. Crude analysis suggests treatment has a modest protective effect. The research team suspects smoking confounds the association because smokers in this population are paradoxically more likely to be on antihypertensive therapy (due to higher cardiovascular risk driving more medication use) and also at higher independent stroke risk.
2,840
Total cohort (n)
1,520
Treated (antihypertensive)
1,320
Untreated
41%
Treated patients who smoke vs 27% untreated
Step 1 — Verify Confounder Criteria for Smoking C₁: Smokers more common in treated group (41% vs 27%, p < 0.001) ✓ C₂: Smoking independently increases stroke risk (OR = 2.1, p < 0.001) ✓ C₃: Smoking is not on the pathway between treatment and stroke ✓ → All three criteria met: smoking is a confounder Step 2 — Apply the 10% Change-in-Estimate Rule Crude OR (treatment vs no treatment): 0.76 (24% apparent risk reduction) Smoking-adjusted OR: 0.54 (46% risk reduction after adjustment) % Change = |0.76 − 0.54| / 0.54 × 100 = 40.7% → far above 10% threshold Step 3 — Stratified Analysis Non-smokers: OR = 0.55 (95% CI 0.40–0.76) Smokers: OR = 0.53 (95% CI 0.37–0.75) Stratum-specific ORs are similar → confirms confounding, not effect modification Mantel-Haenszel pooled OR = 0.54 (consistent with adjusted regression) Interaction Test Smoking × Treatment: OR = 0.96 (95% CI 0.58–1.60), p = 0.87 → no interaction
Smoking was a negative confounder: it masked the true protective effect of antihypertensives. Crude OR = 0.76 understated the benefit; adjusted OR = 0.54 reveals the true ~46% risk reduction.
Why negative confounding is clinically dangerous: Negative confounding caused by confounding-by-indication makes effective treatments appear less effective than they are. In this case, antihypertensives were more commonly prescribed to higher-risk smoking patients, diluting the apparent protective effect in crude analysis. Without confounder adjustment, a clinician reviewing the crude OR of 0.76 would underestimate the benefit of treatment. The adjusted OR of 0.54 represents the true causal effect after accounting for the differential smoking distribution. Note the stratified analysis confirmation: when the stratum-specific ORs (0.55 and 0.53) are approximately equal to each other but differ from the crude (0.76), this is the hallmark pattern of confounding — not effect modification.
3
Exercise Intervention and Weight Loss: Sex as an Effect Modifier
A 6-month supervised aerobic exercise RCT randomises 380 overweight adults (BMI 28–35) to exercise (n = 190) or usual activity control (n = 190). Primary outcome is change in body weight (kg). The protocol pre-specifies sex as a potential effect modifier based on prior physiological evidence that hormonal differences may attenuate the weight-loss response to aerobic exercise in women.
380
Total (n)
−2.1 kg
Overall treatment effect (95% CI −3.2 to −1.0)
p = 0.021
Sex × exercise interaction
Pre-specified
Subgroup analysis (protocol)
Subgroupn (exercise / control)Weight change (kg)95% CIp-value
Overall190 / 190−2.1−3.2 to −1.00.001
Men96 / 92−3.6−5.3 to −1.9<0.001
Women94 / 98−0.7−2.1 to +0.70.318
Interaction testpᵢ​ₙₜₘₙₘ = 0.021 — sex significantly modifies the effect of exercise on weight loss
Sex is a significant effect modifier (p = 0.021). Exercise produces meaningful weight loss in men (−3.6 kg, p < 0.001) but not in women (−0.7 kg, p = 0.318).
Why the overall estimate is misleading: The pooled treatment effect of −2.1 kg is a weighted average that accurately describes neither the male nor female response to exercise. A clinician using the pooled estimate to counsel patients would underestimate the benefit for men and overestimate it for women. Because this subgroup analysis was pre-specified in the protocol, it constitutes confirmatory evidence — not exploratory data dredging — and should be reported as a primary subgroup finding with the interaction test result. The biological explanation (hormonal adaptation, relative fat oxidation efficiency, or compensatory food intake) would be discussed in the paper as the mechanism driving the effect modification. Note: this is an RCT, so sex is not a confounder — both sexes were randomised. The sex-differential response is pure effect modification.

Thesis Writing Recommendations

Confounding and effect modification analysis should be explicitly addressed in the Statistical Analysis section of any observational thesis chapter. For RCT chapters, confounding is less relevant (handled by randomisation) but pre-specified subgroup analyses for effect modification must still be reported.

Statistical Analysis Section

State: (1) which variables were considered potential confounders and the criterion for including them (prior knowledge, DAG, literature, or 10% rule); (2) which variables were examined as potential effect modifiers and the pre-specified rationale; (3) the regression model used for confounder adjustment and which covariates were included; (4) how interaction terms were tested (product term in regression) and the significance level used for declaring effect modification (recommend p < 0.05 for pre-specified analyses, p < 0.10 for exploratory); (5) that stratum-specific estimates will be presented when effect modification is detected.

Model Reporting — Confounding Adjustment
"Crude and adjusted associations between antihypertensive treatment and 5-year stroke were estimated using logistic regression. Potential confounders were identified a priori based on published literature and causal diagram review: age, sex, diabetes status, smoking history, baseline systolic blood pressure, and statin use. Variables were retained in the adjusted model if they changed the exposure coefficient by ≥ 10% using the change-in-estimate criterion. Smoking was identified as a meaningful confounder (40.7% change in estimate). The fully adjusted OR was 0.54 (95% CI: 0.43–0.68), compared to the crude OR of 0.76 (95% CI: 0.62–0.94)."
Model Reporting — Effect Modification
"Effect modification by sex was pre-specified in the statistical analysis plan. An interaction term (sex × exercise allocation) was added to the primary linear regression model. The interaction was statistically significant (p = 0.021), indicating that sex modified the effect of exercise on weight change. Stratum-specific estimates were: men −3.6 kg (95% CI: −5.3 to −1.9, p < 0.001) and women −0.7 kg (95% CI: −2.1 to +0.7, p = 0.318). The pooled estimate (−2.1 kg) is not presented as the primary finding given the significant effect modification by sex."

Common Mistakes Researchers Make

Mistake 1: Treating Effect Modifiers as Confounders and Adjusting Them Away

The most consequential conceptual error in applied epidemiology. When sex, age group, or disease severity modifies the effect of a treatment, adjusting for that variable in a main-effects-only model produces a pooled estimate that represents no actual patient. The biologically real difference between subgroups is erased, and clinically important heterogeneity is concealed from readers and decision-makers.

Fix: Test for interaction before deciding how to handle a variable. If the interaction term is significant (or biologically expected and pre-specified), report stratum-specific estimates and do NOT include the variable only as a main effect. If interaction is absent, include the variable as a covariate in the adjusted model.

Mistake 2: Adjusting for Mediators

A mediator is a variable that lies on the causal pathway between exposure and outcome (E → M → O). Adjusting for a mediator in a regression model blocks the very mechanism of interest, producing a biased estimate of the total effect. A study of diet → obesity → diabetes that adjusts for obesity (the mediator) estimates only the direct diet effect on diabetes not mediated through obesity — which is usually not the primary research question.

Fix: Draw a DAG before analysis to distinguish confounders (variables not on the causal pathway) from mediators (variables on the pathway). If estimating the total effect of an exposure, do not include mediators as covariates. If the mediated pathway is itself of interest, use mediation analysis (counterfactual or SEM-based), not simple covariate adjustment.

Mistake 3: Conditioning on a Collider

Including a collider in the adjustment set (or restricting the analysis to a subset defined by the collider) opens spurious associations between the exposure and outcome. Hospital-based studies are particularly vulnerable: hospitalisation is a collider for almost any disease pair, creating apparent negative associations (Berkson's bias). Adjusting for a collider makes bias worse, not better.

Fix: Draw a DAG before including any variable in the adjustment set. Identify which variables are colliders (common effects of the exposure and outcome, or their descendants) and explicitly exclude them from the adjustment set. Colliders are identifiable only through subject-matter causal reasoning — not through statistical tests.

Mistake 4: Conducting Post-Hoc Subgroup Analyses Without Correction

Running 10 subgroup analyses after seeing the overall result and reporting only the significant one is a major form of selective reporting that inflates Type I error dramatically (10 tests at α = 0.05 gives a 40.1% familywise error rate). This practice is a well-documented source of non-reproducible findings in clinical research and has been highlighted by the CONSORT statement as requiring explicit safeguards.

Fix: Pre-specify all subgroup analyses in the protocol or statistical analysis plan before data collection. Limit to biologically motivated hypotheses. Present all pre-specified subgroup results regardless of significance. Apply Bonferroni or Holm correction to p-values when multiple non-pre-specified subgroup analyses are conducted. State clearly in the paper which analyses were pre-specified and which were exploratory.

Mistake 5: Using Statistical Criteria Alone to Select Confounders

Some researchers include variables as confounders only if they reach statistical significance (p < 0.05) in univariate screening. This is methodologically incorrect: a variable that satisfies all three epidemiological confounding criteria should be included in the adjusted model regardless of its univariate p-value. Statistical non-significance in univariate testing may reflect low power — the variable can still confound the exposure-outcome association in multivariable analysis.

Fix: Identify potential confounders a priori using subject-matter knowledge and DAG analysis. Use the 10% change-in-estimate criterion (applied in the multivariable model, not univariate screening) to determine which variables are empirically meaningful confounders for retention in the final model. Variables identified by prior literature as confounders should be included regardless of their p-value in your specific sample.

Mistake 6: Interpreting Non-Significant Interaction as Absence of Effect Modification

A non-significant interaction term (p ≥ 0.05) does not confirm that the true effect is homogeneous across subgroups. Interaction tests are inherently underpowered — requiring 4× the sample needed for the main effect — and most clinical studies are too small to formally detect biologically real effect modification. Concluding "there is no interaction" based solely on a non-significant p-value when sample sizes are small is scientifically incorrect.

Fix: Always report stratum-specific estimates and their 95% CIs even when the formal interaction test is non-significant. If the stratum-specific estimates appear clinically different (e.g., OR = 2.1 in elderly vs OR = 1.2 in young patients), acknowledge that the study may be underpowered to detect interaction and interpret cautiously. State this limitation explicitly in the Discussion.

Scientific Reporting Standards

The STROBE statement (Strengthening the Reporting of Observational Studies in Epidemiology) provides the most authoritative reporting guidelines for confounding and effect modification:

Practical Guidance for Medical Researchers

Draw Your DAG Before Touching the Data

Before opening any dataset, draw a DAG representing your assumed causal structure. Use dagitty.net (free) to identify the minimal sufficient adjustment set for your exposure-outcome relationship and to flag colliders. The DAG forces you to make your causal assumptions explicit — and reviewers can critique those assumptions rather than the arbitrary variable selection typical of "stepwise regression" approaches.

Pre-Specify Your Interaction Hypotheses

Identify potential effect modifiers from the biological literature before data collection and register them in your study protocol. Pre-specified interaction analyses have far greater credibility than post-hoc subgroup comparisons. Limit to 2–3 biologically motivated modifiers per analysis to avoid inflating the Type I error from multiple interaction tests.

Always Report Both Crude and Adjusted Estimates

The crude (unadjusted) estimate tells the reader what the raw data show; the adjusted estimate tells them what the analysis suggests after controlling for confounding. Presenting both side by side quantifies the impact of confounding adjustment and demonstrates methodological transparency. Reporting only the adjusted estimate without the crude removes the reader's ability to evaluate the confounding impact.

Test for Additive Interaction, Not Only Multiplicative

Standard regression interaction terms test multiplicative interaction. For public health relevance, also calculate RERI (Relative Excess Risk due to Interaction) and its 95% CI to assess additive interaction. The interactionR package in R computes RERI, AP, and S (synergy index) with bootstrap CIs from logistic regression output. These metrics directly answer "how many excess cases are attributable to the combination of both exposures?"

Distinguish "Adjusted for" from "Controlled for"

Multivariable regression adjusts for measured confounders in the model — it cannot control for unmeasured or unknown confounders. This is a fundamental limitation of all observational research. Always include a paragraph in the Discussion acknowledging residual confounding from unmeasured variables as a limitation, and propose what the most likely direction of any residual confounding would be (towards or away from the null).

Use Propensity Score Methods for Confounding by Indication

When studying the effects of treatments that are prescribed preferentially to sicker patients (confounding by indication), propensity score matching or inverse probability weighting creates more comparable groups than covariate adjustment alone. Propensity scores can balance many confounders simultaneously and produce more interpretable treatment-effect estimates than a regression with 15 covariates. The R MatchIt and WeightIt packages provide accessible implementations.

Frequently Asked Questions

What is confounding in medical research? +
Confounding occurs when a third variable distorts the apparent association between an exposure and outcome, making the relationship appear stronger, weaker, or in the wrong direction compared to the true effect. A confounder must be: (1) associated with the exposure; (2) an independent risk factor for the outcome; (3) not on the causal pathway between exposure and outcome. Example: smoking confounds the apparent coffee-CVD association because smokers drink more coffee AND independently have higher CVD risk, creating a false association between coffee and CVD that disappears after adjusting for smoking.
What is effect modification (interaction)? +
Effect modification occurs when the association between an exposure and outcome genuinely differs across strata of a third variable — the effect modifier. Unlike confounding (a bias to remove), effect modification is a real biological phenomenon to discover and report. The protective effect of aspirin against MI may be stronger in men than women (sex is an effect modifier); the risk of GI bleeding from aspirin may be greater in the elderly (age is an effect modifier). Effect modification requires presenting stratum-specific results, not adjusting away the modifier.
What is the critical difference between a confounder and an effect modifier? +
The fundamental distinction: confounding is a bias to control; effect modification is a scientific finding to report. The diagnostic pattern in stratified analysis distinguishes them: if stratum-specific estimates are similar to each other but differ from the crude estimate → confounding (adjust for it). If stratum-specific estimates differ from each other → effect modification (report them separately). Never "adjust for" an effect modifier in a main-effects-only model — this produces a misleading pooled average that represents no actual patient.
How do I detect confounding using the 10% rule? +
Calculate (1) the crude association (unadjusted OR or regression coefficient); (2) the adjusted association after adding the potential confounder. If |Crude − Adjusted| / Adjusted × 100 ≥ 10%, the variable is a meaningful confounder and should be retained in the model. Example: crude OR = 2.0, adjusted OR = 1.5 → 33% change → meaningful confounder. The direction matters too: a crude OR that falls after adjustment indicates positive confounding (crude overestimated the effect); a crude OR that rises indicates negative confounding.
How is effect modification tested statistically? +
Two approaches: (1) Stratified analysis — compute exposure-outcome associations separately in each stratum of the modifier; substantially different stratum-specific estimates indicate effect modification; (2) Interaction term in regression — add the product (Exposure × Modifier) to the model; a significant interaction coefficient (p < 0.05) confirms effect modification. Note: interaction tests are underpowered in most clinical studies — always report stratum-specific estimates and 95% CIs even when the p-value is non-significant.
What is a DAG and why should I use one? +
A Directed Acyclic Graph (DAG) is a causal diagram using arrows to represent assumed causal effects between variables. DAGs identify: which variables are confounders (block backdoor paths), which are mediators (on the causal pathway — do NOT adjust), and which are colliders (common effects — adjusting for them opens spurious associations). Use dagitty.net (free) to draw your DAG, identify the minimal sufficient adjustment set, and detect colliders before beginning analysis. DAGs are drawn from subject-matter knowledge, not statistical tests.
What is collider bias? +
Collider bias occurs when you adjust for (or restrict to) a variable that is a common effect of both the exposure and the outcome (a collider). Conditioning on a collider opens a spurious non-causal path between exposure and outcome, introducing bias. Berkson's paradox is the classic example: in a hospital-based study, conditioning on hospitalisation (a collider for any two conditions that can lead to hospital admission) creates a spurious negative association between otherwise unrelated diseases. Colliders are identified from DAGs — statistical tests cannot detect them.
Should I adjust for an effect modifier in multivariable regression? +
No — do not adjust away an effect modifier in a main-effects-only model. When effect modification is present, the adjusted main effect (without an interaction term) is a misleading weighted average of the subgroup-specific effects and does not represent any actual patient. The correct approach: add the Exposure × Modifier interaction term to the model, then compute and report the stratum-specific effect estimates for each level of the modifier. The interaction term itself is not "the finding" — the stratum-specific estimates are.
What is the difference between additive and multiplicative interaction? +
Multiplicative interaction tests whether the combined effect of two exposures exceeds their product (standard regression interaction terms test this scale). Additive interaction tests whether the combined effect exceeds their sum — measured by RERI (Relative Excess Risk due to Interaction). The two can give opposite conclusions: absence of multiplicative interaction does not imply absence of additive interaction. For public health decisions, additive interaction is more relevant because it quantifies the absolute excess cases attributable to the combination, determining preventable burden.
How should I report confounding and effect modification in a thesis? +
For confounding: present a table with both crude and adjusted estimates; list all confounders included and the criterion used to select them; quantify the change-in-estimate for each confounder. For effect modification: report the interaction test p-value; present stratum-specific estimates with 95% CIs in a table; state whether the interaction was pre-specified or exploratory; interpret the clinical meaning of the heterogeneity. STROBE item 16(a) requires reporting all subgroup and interaction analyses. For RCTs, CONSORT item 18 requires distinguishing pre-specified from exploratory subgroup findings.

Run Your Adjusted Analysis Online

Use StatClinic to compute crude and adjusted odds ratios, test for interaction terms, and explore stratum-specific effect estimates with 95% confidence intervals for your epidemiological or clinical dataset.

Open StatClinic →