Open StatClinic →
🏭 Clinical Trial Design & Analysis

Intention-to-Treat vs Per-Protocol
Analysis in Clinical Trials

🕑 28 min read 📅 July 2026 ✅ Peer-reviewed content 📚 3900+ words
S
StatClinic Editorial Team Statistical content for medical researchers and clinicians
The ACCORD trial enrolled 10,251 patients with Type 2 diabetes and randomly assigned them to intensive glycaemic control (target HbA1c <6.0%) or standard control (target HbA1c 7.0–7.9%). The intention-to-treat analysis, which included all randomized participants regardless of protocol adherence, showed that intensive control unexpectedly increased all-cause mortality by 22% — a result that ended the intensive arm early. A subsequent per-protocol analysis, restricted to participants who maintained their target HbA1c, told a different story. The divergence between these two analyses has been debated in diabetes medicine for over a decade, because they answer different questions: ITT asks “what happens when we assign this strategy to an unselected clinical population?” while per-protocol asks “what happens in patients who achieve the target?” Choosing the wrong analysis for the wrong question produces misleading evidence that shapes clinical practice. Choosing the right analysis — or reporting both correctly — is one of the most consequential methodological decisions in a clinical trial.

Intention-to-Treat Analysis

The intention-to-treat (ITT) principle requires that all randomized participants be included in the analysis, assigned to the group to which they were originally randomized, regardless of what subsequently happened: whether they received the assigned treatment, adhered to the protocol, crossed over to the other arm, withdrew from the study, or were lost to follow-up. The operative phrase is analyze as randomized.

ITT answers a specific, policy-relevant question: What is the overall effect of assigning this treatment strategy to this patient population? This includes the full messy reality of clinical practice — patients who don’t take their medication, patients who drop out when they feel better (or worse), patients who cross over because their condition deteriorates. In this sense, ITT estimates effectiveness — what happens in a real-world healthcare system, not a controlled ideal environment.

Intention-to-Treat (ITT)

Effectiveness Analysis
  • All randomized participants included
  • Analyzed in the group they were assigned to
  • Regardless of actual treatment received
  • Regardless of adherence or crossover
  • Missing outcomes handled by imputation or MMRM
  • Preserves randomization — confounders balanced
❓ “What happens when we assign this treatment to real patients?”

Per-Protocol (PP)

Biological Efficacy Analysis
  • Only adherers analyzed (received treatment as specified)
  • Excludes crossovers, major protocol violators
  • Excludes participants below minimum dose threshold
  • Excludes those who didn’t complete minimum follow-up
  • Breaks randomization — confounders may be unbalanced
  • Susceptible to selection bias (healthy adherer effect)
❓ “What happens when patients actually take the treatment?”

Why ITT Preserves the Validity of Randomization

The single most important property of a randomized controlled trial is that randomization distributes both measured and unmeasured confounders equally between treatment arms. This balance is what allows causal inference — the ability to attribute a difference in outcomes between arms to the treatment, rather than to pre-existing differences between the groups.

When per-protocol analysis excludes non-adherers, this balance is destroyed. The excluded participants are not missing at random — they are systematically different from adherers in multiple ways:

The healthy adherer bias is the most studied manifestation: participants who adhere to protocol are healthier, more motivated, and better-prognosis than non-adherers — regardless of which arm they are in. This has been demonstrated in multiple large trials where patients randomized to placebo who adhered to taking their placebo tablets lived longer than non-adherent patients — simply because adherence itself is a proxy for health behaviors and prognosis.

The healthy adherer effect in practice: A landmark illustration came from observational pharmacoepidemiology studies showing that patients who filled their statin prescriptions regularly had lower mortality than non-adherent patients — but this apparent statin benefit partially persisted even for patients who stopped taking statins, suggesting that what was being measured was the prognostic health behavior of medication adherence, not the pharmacological statin effect. In per-protocol analyses, this same confounding inflates apparent treatment efficacy.

When Per-Protocol Analysis Is Essential

ITT is the primary analysis for most confirmatory RCTs, but PP analysis fills indispensable roles that ITT cannot serve.

Non-Inferiority and Equivalence Trials

This is where the analytic implications reverse, and ITT’s conservative properties become a liability. In superiority trials, ITT is conservative — non-adherence dilutes the treatment effect toward the null, reducing Type I error. In non-inferiority trials, this dilution is anti-conservative: non-adherence makes both arms look more alike, pushing them toward equivalence and falsely suggesting non-inferiority of the new treatment. Both ITT and PP must confirm non-inferiority for the conclusion to be scientifically valid.

⚠ Non-Inferiority Trial Rule: Both ITT and PP Must Agree

If ITT shows non-inferiority (new ≥ standard − margin) but PP does not: the ITT result may be driven by non-adherence diluting the true inferiority of the new drug toward apparent equivalence. Conclusion: non-inferiority NOT established.

If PP shows non-inferiority but ITT does not: high non-adherence in the new drug arm (possibly due to side effects or poor palatability) may be causing real-world failure even when the drug works pharmacologically. Conclusion: non-inferiority NOT established in the real-world sense.

Only when both ITT and PP analyses independently confirm non-inferiority is the conclusion robust. ICH E9 and EMA guidelines explicitly require both analyses for non-inferiority trials submitted for regulatory approval.

Understanding Biological Mechanism

When a trial is designed to establish whether a treatment has a pharmacological or biological effect — independent of real-world adherence challenges — PP provides the cleaner signal. Phase II dose-finding trials, mechanistic pharmacokinetic-pharmacodynamic trials, and proof-of-concept studies often prioritise PP analysis because they are asking “does this molecule work?” rather than “does this treatment strategy work in clinical practice?”

Vaccine Efficacy Trials

Vaccine efficacy is conventionally defined in the per-protocol population: participants who received the full vaccine series according to schedule, had no pre-existing immunity at baseline, and were followed for the pre-specified exposure period. The ITT estimate includes individuals who received partial vaccination, missed second doses, or had pre-existing immunity — diluting the efficacy estimate below the true pharmacological effect of the vaccine regimen.

Modified ITT (mITT)

Modified ITT (mITT) is a pre-specified variation that allows a narrowly defined exclusion from the ITT population — typically, participants who were randomized but never received any study intervention and contributed no outcome data. This is the most legitimate and commonly used exclusion.

The requirements for an mITT exclusion to be scientifically acceptable are strict: the exclusion criteria must be (1) pre-specified in the protocol before randomization occurs; (2) defined entirely by events or information that precede randomization or treatment initiation; and (3) independent of any outcome measurement. If any outcome information influences who is included in the mITT, the mITT becomes an unacknowledged PP analysis.

mITT abuse: The mITT label has been extensively misused in published trials. A survey of published mITT definitions found that approximately 40% included at least one criterion that could introduce outcome-dependent selection bias — effectively converting an ITT analysis into a PP analysis while calling it ITT. Common abuse: excluding patients who “did not reach the minimum efficacy-evaluable threshold” (which is outcome-dependent) or excluding patients with “insufficient data for analysis” (which conflates missing outcomes with protocol non-adherence). Reviewers should scrutinise any mITT definition carefully, and thesis examiners will expect you to justify every exclusion.

Types of Protocol Deviations

Non-Adherence to Treatment

Participant received less than the minimum specified dose, missed ≥ X% of doses (pre-specified threshold), stopped treatment early, or received an unapproved concomitant intervention that interacts with the study treatment.

Crossover Between Arms

Participant assigned to control arm received the experimental treatment (or vice versa) during the follow-up period — either by clinician decision or self-initiation. Most consequential deviation for both ITT and PP analysis.

Lost to Follow-Up

Participant cannot be located or contacted for outcome assessment. Distinguished from withdrawal: the participant did not actively withdraw consent — they are simply missing. ITT must account for these through imputation or mixed-model approaches.

Ineligibility Post-Randomization

A participant is found after randomization to not meet inclusion criteria — e.g., baseline lab result confirms exclusion criterion missed at screening. Must be disclosed in the CONSORT flow diagram; inclusion in mITT exclusion only if pre-specified.

Consent Withdrawal

Participant actively withdraws consent for continued participation. By law, their data cannot be used if they explicitly withdraw consent for data use. This is the only exclusion from ITT that is non-discretionary and regulatory.

Major Protocol Violation

A deviation serious enough to compromise data integrity or participant safety — e.g., administration of a prohibited medication, failure to perform a required procedure, breach of blinding. Major violations are typically identified by the Data Safety Monitoring Board (DSMB) and documented in the trial master file.

Missing Data in ITT Analysis

The inclusion of all randomized participants in ITT is conceptually simple but practically complex: outcome data will inevitably be missing for some participants (dropouts, lost to follow-up, deaths from competing causes). How these missing outcomes are handled substantially affects the validity of ITT estimates.

MethodAssumptionWhen to UseStatus
MMRM
(Mixed-Model Repeated Measures)
Missing at Random (MAR) — dropout depends on observed data, not future unobserved outcomes Continuous longitudinal outcomes in confirmatory trials; uses all observed data without explicit imputation Recommended
Multiple Imputation (MI) MAR — missing values imputed from a regression model using observed data; analysis repeated across imputed datasets; pooled via Rubin’s rules When MMRM is not directly applicable; binary or time-to-event outcomes with missing data; required by ICH E9(R1) Recommended
LOCF
(Last Observation Carried Forward)
Missing Not at Random (MNAR) assumption implicit: the outcome stays constant at its last measured value after dropout Historically dominant; still used in some regulatory settings; biased when outcomes trend post-dropout (common in most clinical conditions) Use with caution
BOCF
(Baseline Observation Carried Forward)
Extreme conservative assumption: assumes no benefit at all after dropout Appropriate when dropout is informative (patients drop out because they are getting worse); worst-case sensitivity analysis Sensitivity only
Complete Case Analysis Missing Completely at Random (MCAR) — assumes dropout is unrelated to outcome. Virtually never met in clinical trials. Only valid if MCAR assumption can be demonstrated (nearly impossible to verify). Equivalent to PP for completers. Generally avoid

Clinical Examples

1
Antidepressant RCT: Why ITT and PP Diverge — and Why the Gap Is Informative
A double-blind RCT compares escitalopram vs placebo for major depressive disorder. n = 120 (60 per arm). Primary outcome: PHQ-9 score reduction at 12 weeks (MCID = 5 points). Differential dropout: 7/60 (12%) in the treatment arm withdrew due to side effects; 17/60 (28%) in the placebo arm withdrew citing lack of perceived benefit.
120
Randomized (60/arm)
12%
Dropout: treatment arm
28%
Dropout: placebo arm
96
Completers (PP eligible)
ITT Analysis (all 120 randomized; missing outcomes by MMRM) Treatment arm (n=60): mean PHQ-9 reduction −8.4 (SD 5.1) Placebo arm (n=60): mean PHQ-9 reduction −5.3 (SD 5.6) Difference: −3.1 points (95% CI: −5.1, −1.1) t(118) = 3.04, p = 0.003, Cohen's d = 0.57 PP Analysis (completers only: 53 treatment, 43 placebo) Treatment arm (n=53): mean PHQ-9 reduction −10.2 (SD 4.2) Placebo arm (n=43): mean PHQ-9 reduction −4.6 (SD 4.8) Difference: −5.6 points (95% CI: −7.3, −3.9) t(94) = 6.51, p < 0.001, Cohen's d = 1.30 Why the PP Effect is More Than Twice the ITT Effect 17 placebo patients dropped out "for lack of benefit" → They were the worst-responders in the placebo arm → Removing them from PP artificially improves the placebo arm's mean → The apparent placebo response in PP (−4.6) looks better than in ITT (−5.3) because the worst cases are gone from the denominator Dropout reason in treatment arm: primarily side effects (nausea, insomnia) → These patients were excluded from PP despite having early treatment response → Removing them from PP artificially improves the treatment arm estimate too → But the control arm purification effect dominates The Dropout Pattern Is Itself Clinical Evidence Placebo withdrawal rate: 28% — "lack of benefit" Treatment withdrawal rate: 12% — side effects (suggests drug IS acting) Differential: 16 percentage points This differential is evidence the drug is working — patients know it's working or not working, even in a double-blind trial (through perceived symptom change)
📊 ITT (Primary) — Report This First
Δ = −3.1 pts (95% CI: −5.1, −1.1)
p = 0.003, d = 0.57
n = 120 (all randomized)
📋 PP (Sensitivity) — Report Alongside
Δ = −5.6 pts (95% CI: −7.3, −3.9)
p < 0.001, d = 1.30
n = 96 (completers only)
The key interpretive insight: The ITT-PP gap (d = 0.57 vs d = 1.30) is not a failure of one analysis or the other — it is clinically meaningful information. The large PP effect tells us the drug has a substantial biological effect when taken. The smaller ITT effect reflects real-world dropout, primarily side effects, which are genuine treatment consequences that must be represented in the effectiveness estimate. The differential dropout pattern (28% placebo vs 12% treatment) itself constitutes evidence of treatment efficacy — patients in a clinical setting recognize whether they are improving, even under double-blind conditions. Reporting only the PP result would overstate the benefit clinicians can realistically expect for an unselected patient population; reporting only the ITT result without acknowledging the PP would understate the biological effect in patients who can tolerate and persist with the treatment.
2
Antibiotic Prophylaxis Non-Inferiority Trial: Why Both Analyses Are Mandated
A surgical prophylaxis trial randomizes 500 patients (250/arm) undergoing elective colorectal surgery to: oral antibiotic A (new, cheaper, outpatient-compatible) vs IV antibiotic B (established standard). Primary outcome: surgical site infection (SSI) within 30 days. Non-inferiority margin: ΔNI = 3% absolute (if A’s SSI rate is within 3% of B’s, it is declared non-inferior). A total 28/250 patients in arm A and 12/250 in arm B had major protocol adherence deviations.
500
Randomized (250/arm)
3.0%
Non-inferiority margin
11.2%
Arm A SSI rate (ITT)
8.8%
Arm B SSI rate (ITT)
ITT Analysis (all 250/arm) Arm A (oral, n=250): SSI 28/250 = 11.2% Arm B (IV, n=250): SSI 22/250 = 8.8% Difference (A−B) = +2.4% (Arm A has higher SSI rate) 95% CI for difference: [−0.7%, +5.5%] Non-inferiority test: upper CI limit (5.5%) > NI margin (3.0%) → ITT does NOT confirm non-inferiority PP Analysis (received ≥90% of assigned prophylaxis course) PP-eligible: Arm A 222/250 (88.8%), Arm B 238/250 (95.2%) [Arm A has higher non-adherence: 28 patients did not complete full oral course] Arm A PP (n=222): SSI 18/222 = 8.1% Arm B PP (n=238): SSI 18/238 = 7.6% Difference (A−B) PP = +0.5% 95% CI: [−2.1%, +3.1%] Non-inferiority test: upper CI limit (3.1%) > NI margin (3.0%) — marginally fails → PP also does NOT confirm non-inferiority (upper CI just exceeds margin) Why the ITT-PP Gap Exists Here 28 patients in Arm A had non-adherence (oral course not completed) → They received partial prophylaxis → higher SSI rates → Including them in ITT worsens Arm A's apparent performance vs PP → The ITT vs PP discrepancy reveals an adherence problem with the oral regimen Regulatory Conclusion Neither ITT nor PP confirms non-inferiority. The 28-patient adherence failure in the oral arm is clinically informative: the oral regimen may be pharmacologically equivalent but adherence-dependent. Recommendation: redesign with adherence support; repeat trial with n inflated for the observed 12% non-adherence rate.
Neither the ITT analysis (upper CI 5.5% > margin 3.0%) nor the PP analysis (upper CI 3.1% > margin 3.0%) confirmed non-inferiority of oral prophylaxis over IV. The higher non-adherence in the oral arm (11.2% vs 4.8%) suggests a real-world compliance challenge that ITT correctly captures and that investigators must address before the oral regimen can be considered clinically equivalent.
The non-inferiority lesson: Had the investigators reported only the PP result — which shows near-identical SSI rates (8.1% vs 7.6%) — they might have concluded equivalence. But the ITT analysis, by capturing the real-world adherence failure, correctly flags that the oral regimen does not perform as well as the IV standard when all patients in an unselected population are included. This is exactly the scenario where regulatory bodies (EMA, FDA) mandate both analyses: PP tests pharmacological non-inferiority under ideal conditions; ITT tests clinical non-inferiority in real-world practice. When they disagree, as here, the new treatment cannot be considered non-inferior without addressing the source of the disagreement.
3
T2DM Lifestyle Intervention: Complier-Average Causal Effect (CACE) Analysis
An RCT evaluates an intensive lifestyle programme (eight group sessions, diet counselling, and structured exercise) vs usual care for Type 2 diabetes control. n = 160 (80/arm). Primary outcome: HbA1c reduction at 6 months. However, 35/80 (44%) of intervention participants attended fewer than 50% of sessions and are classified as non-compliers. The researchers want to know: what was the true effect of the programme for participants who engaged with it — but without the selection bias of standard PP?
160
Randomized (80/arm)
44%
Non-compliers (intervention)
−0.9%
ITT HbA1c difference
−1.5%
PP HbA1c difference
ITT Analysis (all 80/arm) Intervention: mean HbA1c reduction −1.2% (SD 0.8%) Usual care: mean HbA1c reduction −0.3% (SD 0.6%) ITT difference: −0.9% (95% CI: −1.12%, −0.68%), p < 0.001 Per-Protocol Analysis (≥50% sessions attended; n=45 intervention, 80 control) Intervention PP: mean HbA1c reduction −1.8% (SD 0.6%) Usual care: mean HbA1c reduction −0.3% (SD 0.6%) PP difference: −1.5% (95% CI: −1.72%, −1.28%), p < 0.001 [Larger than ITT — selection bias: PP participants are more motivated] CACE Analysis (Complier-Average Causal Effect) CACE = ITT effect / proportion of compliers in intervention arm Complier proportion = 45/80 = 0.5625 (56.25%) CACE = −0.9% / 0.5625 = −1.60% Interpretation: the true causal effect of the lifestyle programme in patients who would comply with it if assigned to it is estimated at −1.60% HbA1c reduction. The CACE (−1.60%) is close to the PP estimate (−1.50%) — both suggest the programme has a genuine ≈1.5% HbA1c effect in compliers. The difference is the method of estimation: PP: selects observed compliers (selection bias possible) CACE: uses randomization as an instrument — unbiased if assumptions hold CACE Assumptions (Instrumental Variable) 1. Relevance: randomization assignment predicts treatment receipt (✓ — partially) 2. Exclusion restriction: assignment only affects outcome through treatment (✓) 3. Monotonicity: no "defiers" — no one assigned to usual care who would have attended sessions if assigned to intervention (approximately ✓)
ITT: −0.9% HbA1c (p < 0.001). PP: −1.5% (selection bias risk). CACE: −1.6% — the unbiased causal effect estimate for compliers, using randomization as an instrumental variable. All three analyses are reported with their assumptions, limitations, and the clinical recommendation: invest in engagement support to improve real-world compliance from 56% toward the PP-equivalent population.
When CACE adds value over PP: CACE (also called the Local Average Treatment Effect or LATE) is particularly valuable in behavioural and lifestyle intervention trials where non-compliance is high and heterogeneous — patients who attend 10% of sessions and patients who attend 90% are both “non-compliant” by a binary threshold, but their expected outcomes differ substantially. Standard PP conflates non-compliance with the choice to attend; CACE uses the random assignment as an instrument to isolate the causal effect of compliance on outcome, avoiding the selection bias inherent in self-selected compliers. The convergence of PP (−1.5%) and CACE (−1.6%) here provides reassurance that the PP estimate was not substantially biased — but this cannot be assumed without the CACE check.

Thesis Writing Recommendations

For students conducting or secondary-analysing an RCT, the analysis strategy section is one of the most technically demanding and most closely scrutinised parts of the thesis. Examiners at postgraduate level will specifically check whether the ITT principle was correctly applied, whether any mITT exclusions were pre-specified, and whether PP was correctly framed as a secondary sensitivity analysis.

Model Statistical Analysis Paragraph — RCT Analysis Strategy
“The primary analysis followed the intention-to-treat (ITT) principle, including all participants as randomized regardless of protocol adherence, treatment discontinuation, or crossover. Missing outcome data at 12 weeks were handled using a mixed-model repeated measures (MMRM) approach, which used all observed longitudinal data and assumed data were missing at random (MAR). A pre-specified per-protocol (PP) sensitivity analysis was conducted in participants who attended ≥ 80% of scheduled sessions and had no major protocol violations. The PP population was defined before database lock, with the protocol deviation list adjudicated by a blinded committee independent of outcome assessment. The non-inferiority margin was pre-specified as ΔNI = 3%, and non-inferiority was to be declared only if both the ITT and PP 95% CIs excluded this margin.”
Model Results Disclosure — ITT and PP Reporting
“Of 120 randomized participants, 7 (12%) in the escitalopram arm and 17 (28%) in the placebo arm withdrew before 12 weeks; all 120 were included in the ITT primary analysis. In the ITT analysis, escitalopram produced a significantly greater PHQ-9 reduction than placebo (mean difference −3.1 points, 95% CI: −5.1 to −1.1; p = 0.003; Cohen’s d = 0.57). In the pre-specified PP sensitivity analysis (n = 96 completers), the treatment difference was larger (−5.6 points, 95% CI: −7.3 to −3.9; d = 1.30). The greater PP effect is consistent with the differential dropout pattern — placebo-arm withdrawal was predominantly for perceived lack of benefit — and should be interpreted as an upper-bound estimate of biological efficacy rather than the expected real-world effectiveness.”

Common Mistakes Researchers Make

Mistake 1: Reporting Per-Protocol as the Primary Analysis in a Superiority Trial

When researchers report only the PP analysis — or present the PP result prominently with ITT as a footnote — in a superiority trial, the finding is typically inflated. This occurs because PP removes the worst-responding non-adherers from both arms (but disproportionately from the control arm), artificially improving the treatment-control contrast. Journal reviewers and ethics committee statisticians will immediately question a superiority trial that reports PP as primary.

Fix: For superiority trials, designate ITT as the primary analysis in the protocol, pre-register this decision, and report PP as a pre-specified sensitivity analysis. In Results, lead with the ITT result and present PP as context. Never reverse this hierarchy post-hoc to make a non-significant ITT result appear significant through PP.

Mistake 2: Treating Non-Inferiority Trials as If Only ITT Is Needed

The conservative nature of ITT is well-known for superiority trials. What is less well-understood — and more consequential — is that ITT is anti-conservative for non-inferiority trials. Non-adherence dilutes the apparent difference between arms toward zero, pushing toward equivalence and falsely supporting non-inferiority of an inferior new treatment. A new drug that is actually inferior to the standard can achieve ITT non-inferiority simply if enough patients in both arms fail to adhere.

Fix: For every non-inferiority or equivalence trial, pre-specify that both ITT and PP analyses are required and that non-inferiority will only be declared when both analyses independently confirm it. Report both sets of CIs against the non-inferiority margin. If they diverge, investigate the source of divergence — differential non-adherence between arms is the most common explanation.

Mistake 3: Defining mITT Exclusions Post-Hoc or Using Outcome-Dependent Criteria

The abuse of the mITT label is one of the most common methodological problems in published RCTs. Post-hoc mITT exclusions — defined after seeing the data to exclude participants who are inconvenient for the analysis — can systematically remove the worst-outcome participants and inflate apparent efficacy. Outcome-dependent exclusions (“patients without evaluable efficacy data”) are particularly problematic because missing data and treatment failure are often correlated.

Fix: Define all mITT exclusions in the protocol before the first participant is randomized. Ensure all exclusion criteria are based on pre-randomization information (eligibility confirmed post-screening, no treatment ever received, consent withdrawn pre-intervention). Report the full ITT analysis (all randomized) alongside the mITT in all submissions. Any mITT that differs substantially from the full ITT requires explicit investigation and justification.

Mistake 4: Using Complete Case Analysis as a Substitute for ITT

Complete case analysis — analyzing only participants with complete outcome data — is frequently presented as ITT in papers where the authors have confused “including all participants who completed the study” with “including all participants as randomized.” Complete case analysis is statistically equivalent to PP for completers; it violates the ITT principle and introduces the same selection biases as PP analysis, while being labelled as ITT.

Fix: The denominator in ITT Results must match the denominator in the randomized row of the CONSORT flow diagram. Any reduction in n between randomization and analysis must be accounted for by an explicit missing data handling method (MMRM, multiple imputation) — not by dropping participants from the analysis. Peer reviewers will cross-check the flow diagram denominator against the analysis denominator.

Mistake 5: Ignoring the Differential Dropout Pattern When Interpreting ITT vs PP

When ITT and PP analyses produce different effect size estimates, many researchers report the discrepancy without investigating or explaining why the gap exists. The ITT-PP gap is not a statistical nuisance — it is clinically informative. Understanding why participants dropped out of each arm, and whether the dropout patterns were differential, provides essential context for interpreting both the ITT estimate (real-world effectiveness) and the PP estimate (biological efficacy).

Fix: Report the dropout rate and primary dropout reasons for each arm separately. If the rates differ substantially between arms, analyse whether dropout was associated with baseline characteristics or early outcome measures. Discuss how the differential dropout pattern explains the ITT-PP gap. A table comparing baseline characteristics of completers vs non-completers within each arm helps readers assess the direction and magnitude of healthy adherer bias.

Mistake 6: Using LOCF as the Default Missing Data Approach for ITT

Last Observation Carried Forward (LOCF) was the dominant missing data approach in RCTs for decades and remains common in older trial reports. Its central assumption — that a patient’s outcome remains at its last measured value after dropout — is nearly always incorrect for clinical outcomes. For a condition that tends to improve (depression, HbA1c in treated diabetes), LOCF for placebo dropouts overstates placebo response; for conditions that worsen (progressive neurological disease), LOCF understates decline. LOCF biases the treatment comparison in the direction most favourable to the drug if it is applied asymmetrically to the two arms.

Fix: Use MMRM as the primary missing data approach for continuous longitudinal outcomes — it correctly models the trajectory of all observed data without explicit imputation. Use multiple imputation as an alternative or for non-continuous outcomes. LOCF should be reported only as a pre-specified sensitivity analysis for historical comparability or regulatory requests — never as the primary ITT method. Justify missing data approach in the Statistical Analysis section by citing ICH E9(R1).

Scientific Reporting Standards

Practical Guidance

Pre-Specify Everything Before Randomization

Register the ITT population definition, any mITT exclusion criteria, the PP eligibility threshold (e.g., ≥80% adherence), the missing data approach, and the order of primary and secondary analyses in the trial protocol and Statistical Analysis Plan — before any participant is randomized. Post-hoc decisions about which participants to include or exclude are the single greatest source of bias in RCT reporting. ClinicalTrials.gov and ISRCTN registration with a pre-specified analysis plan creates an immutable public record.

Use MMRM as Your Default for Continuous Longitudinal Outcomes

For the ITT analysis of continuous longitudinal outcomes (e.g., PHQ-9 over 4 time points, HbA1c at 3, 6, and 12 months), Mixed-Model Repeated Measures (MMRM) is the current gold standard. It uses all observed data from all participants at all time points, handles missing data under the MAR assumption without separate imputation, and is less sensitive to model misspecification than single-imputation methods. Specify MMRM in the SAP with the full covariance structure (unstructured, compound symmetry, or AR1) and the fixed effects included.

Report the CONSORT Flow Diagram With Correct Denominators

Every RCT paper and thesis chapter must include a CONSORT flow diagram. The denominator in every row of the Results table must be traceable to the flow diagram. The most common reviewer critique of ITT violations is a mismatch between the number randomized (flow diagram) and the number analysed (Results). If you used MMRM or multiple imputation, the analysed n equals the randomized n — all participants are included, with missing data handled by the model rather than exclusion.

Investigate and Report Differential Dropout

For any trial where dropout rates differ between arms by more than 5 percentage points, conduct and report a formal dropout analysis: compare baseline characteristics of completers vs dropouts within each arm; test whether dropout was associated with early outcomes; examine the primary dropout reason in each arm (side effects vs lack of benefit vs unrelated events). This analysis belongs in the Results, not the Supplementary. The dropout pattern contextualises both the ITT and PP results for readers trying to understand the clinical implications.

For Lifestyle and Behavioural Trials, Consider Reporting CACE

When adherence to a behavioural or lifestyle intervention is variable and you expect a large ITT-PP gap, consider adding a CACE (complier-average causal effect) analysis as a secondary analysis alongside the standard PP. CACE uses randomization as an instrumental variable to estimate the effect in compliers without the selection bias of observed PP. Report the proportion of compliers, the ITT estimate, and the CACE = ITT / compliance rate. The convergence of CACE and PP provides reassurance about the absence of healthy adherer bias.

Know Your Estimand Before Designing the Analysis

The ICH E9(R1) estimand framework asks: what exactly do you want to estimate, and for whom, and under what conditions? Before choosing ITT or PP, define the estimand: “The effect of assigning drug A vs placebo on PHQ-9 at 12 weeks in adults with MDD, regardless of treatment discontinuation (treatment policy estimand).” This forces a precise answer to the research question before choosing the analysis. Different estimands — treatment policy, hypothetical (what if everyone adhered?), principal stratum — require different analytic approaches. Clarifying the estimand prevents the most common error: choosing the analysis first and the question second.

Frequently Asked Questions

What is intention-to-treat (ITT) analysis? +
ITT analysis includes all randomized participants in the analysis, assigned to the group they were originally randomized to, regardless of whether they received the assigned treatment, adhered to the protocol, crossed over, or withdrew. The operative rule is: analyze as randomized. ITT preserves the statistical balance created by randomization — ensuring confounders are equally distributed between arms — and estimates the real-world effectiveness of the treatment policy, not the biological efficacy under ideal conditions. It is the primary analysis for most confirmatory RCTs, required by CONSORT 2010 and ICH E9.
What is per-protocol (PP) analysis? +
PP analysis restricts the analysis to participants who adhered to the protocol — they received the assigned treatment as specified, did not cross over, completed minimum follow-up, and had no major protocol violations. Also called on-treatment analysis, efficacy-evaluable analysis, or completer analysis. PP estimates biological efficacy — does the treatment work when taken as prescribed? The critical weakness: participants who adhere are not a random subset of those randomized. They tend to be healthier, more motivated, and better-prognosis (healthy adherer effect), inflating the apparent treatment-control difference.
What is the difference between ITT and per-protocol analysis? +
The fundamental difference is what question each answers: ITT asks “does assigning this treatment to this population improve outcomes?” (effectiveness, real-world); PP asks “does this treatment work when patients take it as prescribed?” (efficacy, ideal conditions). Statistically: ITT includes all randomized participants (preserving randomization balance); PP excludes non-adherers (breaking randomization, introducing selection bias). PP effect sizes are typically larger than ITT because non-responding dropouts — who have worse outcomes — are excluded from PP but correctly included in ITT. The ITT-PP gap itself is clinically informative about the real-world impact of non-adherence.
Why is ITT the preferred primary analysis for RCTs? +
ITT is preferred because it preserves the statistical properties of randomization — the equal distribution of measured and unmeasured confounders between arms that is the RCT’s core advantage over observational research. PP analysis breaks this balance by excluding non-adherers who are systematically different from adherers: they tend to have worse prognosis, more comorbidities, lower health literacy, and greater disease burden. This creates a confounded comparison indistinguishable from a poorly matched observational study. ITT also represents the real-world policy question most relevant to clinical practice: not “what happens in an ideally compliant patient?” but “what happens when we offer this treatment to an unselected clinical population?”
When is per-protocol analysis required or preferred? +
PP analysis is required or preferred: (1) Non-inferiority and equivalence trials — ITT is anti-conservative for non-inferiority because non-adherence dilutes differences toward equivalence. Both ITT and PP must confirm non-inferiority. (2) Vaccine efficacy trials — efficacy is defined in the per-protocol (full vaccination series) population. (3) Understanding biological mechanism — Phase II proof-of-concept trials asking “does the drug work pharmacologically?” (4) When adherence to the assigned treatment is the intervention itself — and the estimand of interest is the hypothetical effect under full compliance. In all cases, PP must be pre-specified and reported alongside ITT.
What is modified ITT (mITT)? +
Modified ITT is a pre-specified variant that allows a narrow, pre-defined exclusion from the ITT population — most commonly, participants who were randomized but never received any study intervention. Requirements: exclusion criteria must be pre-specified in the protocol before randomization; criteria must be based on pre-randomization information only; no outcome data can influence who is included. The full ITT analysis (all randomized) must be reported alongside mITT. mITT has been widely misused: any post-hoc mITT exclusion or outcome-dependent exclusion constitutes an unacknowledged PP analysis and should be identified as such by peer reviewers and examiners.
Why do ITT and per-protocol analyses produce different results? +
They differ because non-adherers are systematically different from adherers, and their exclusion from PP distorts the comparison. In the treatment arm: dropouts often leave due to adverse effects or lack of perceived benefit — both correlated with worse outcomes. In the control arm: dropouts often leave due to perceived lack of efficacy — they are the worst responders. Removing the worst control-arm responders from PP inflates the control arm’s apparent performance, widening the treatment-control gap. The healthy adherer effect — where adherers in both arms are healthier regardless of which treatment they receive — further confounds PP comparisons.
What is the Complier-Average Causal Effect (CACE)? +
CACE is the causal effect of treatment in the subgroup who would comply with their assignment regardless of which arm they were randomized to. It is estimated using randomization as an instrumental variable: CACE = ITT effect / proportion of compliers in the treatment arm. CACE avoids the selection bias of standard PP because it uses the randomization — not self-selection into adherence — to identify the complier effect. It is most useful in behavioural and lifestyle trials with high and variable adherence. Requires assumptions: randomization predicts treatment receipt; assignment only affects outcome through treatment; no ‘defiers.’
How should ITT and PP analyses be reported in a thesis or paper? +
CONSORT 2010 requires: (1) flow diagram showing n randomized, n allocated, n lost, n analysed in each arm; (2) primary analysis ITT with denominator = n randomized; (3) pre-specified PP as sensitivity analysis with explicit labelling and its n. In Methods: define ITT population, any mITT exclusions (pre-specified only), PP eligibility criteria, and missing data approach. In Results: report ITT as primary with effect size + 95% CI + p-value; PP as sensitivity analysis with explicit “pre-specified secondary analysis” label. Discrepancies between ITT and PP must be interpreted — not reported silently.
How is missing data handled in ITT analysis? +
Key approaches: (1) MMRM — uses all observed longitudinal data under MAR assumption; gold-standard for continuous outcomes; recommended for confirmatory trials. (2) Multiple Imputation (MI) — creates multiple complete datasets, analyzes each, pools via Rubin’s rules; required by ICH E9(R1) for many regulatory submissions. (3) LOCF — carries the last observed value forward; historically common but biased; assumes no change after dropout (almost never true). (4) BOCF — returns all dropouts to baseline; extremely conservative; appropriate for worst-case sensitivity analyses only. Complete case analysis (excluding dropouts) violates ITT and is equivalent to PP for completers — it should not be presented as ITT.

Analyse Your Clinical Trial Data

Statistical tools for RCT analysis, effect size calculation, missing data assessment, and CONSORT-compliant reporting.

Open StatClinic →