Intention-to-Treat (ITT) Analysis in RWE and Target Trials
An analysis strategy that locks each person to the treatment strategy assigned or emulated at time zero and counts follow-up and outcomes under that initial strategy regardless of later discontinuation, switching, non-adherence, dose changes, or add-on therapy.
On this page
Intention-to-treat analysis keeps each patient in the treatment group they started in, even if they later stop, switch, or add another treatment. In a randomized trial this protects the original random assignment. In real-world data it is better described as the effect of the initial treatment decision under usual-care adherence, because clinicians and patients chose the starting treatment rather than being randomized.
Intention-to-treat (ITT) analysis
estimates the effect of initiating or being assigned to a treatment strategy, not the effect of actually taking treatment continuously. In a randomized trial, the ITT principle protects the baseline exchangeability created by randomization: participants are analyzed in the groups to which they were assigned, and post-randomization non-adherence, switching, and rescue treatment are treated as part of the treatment policy being compared.
In an RWE target-trial emulation there is no randomized assignment to preserve, so "ITT" is more precisely the observational analog of the treatment-policy or initiation effect: define eligible people at time zero, assign the arm from the baseline treatment decision that the data can observe, adjust for baseline confounding, then follow everyone under the initial strategy even if real-world care later diverges.
The important discipline is that ITT is an estimand choice, not a magic bias shield.
For a two-drug new-user study, an ITT-style emulation asks, "What is the effect of starting drug A rather than drug B under the adherence and switching patterns that normally follow that start?" It does not ask, "What is the effect if everyone remains on drug A for two years?" Heavy discontinuation or crossover can dilute an ITT contrast toward no difference, but that dilution may be exactly the policy-relevant answer when the decision maker cares about recommending a strategy in usual care.
Conversely, if the scientific question is biological efficacy while adherent, a per-protocol or while-on-treatment estimand is the right companion analysis.
Pros, cons, and trade-offs
- vs per-protocol analysis: ITT is simpler, avoids conditioning on post-baseline adherence, and is usually the primary policy-effect contrast. It remains interpretable when discontinuation and switching are expected parts of usual care. Cost: it can be diluted by non-adherence and cannot isolate the effect under full protocol adherence. Prefer ITT when the decision is "start this strategy" or "assign this policy"; prefer per-protocol when the decision is "what if patients actually follow the strategy?"
- vs naive as-treated analysis: ITT avoids the selection bias created by reclassifying or censoring people based on post-baseline behavior. A naive as-treated analysis often makes adherers look healthier because adherence is a prognostic post-baseline variable. Cost: ITT attributes off-treatment follow-up to the initial arm, so acute pharmacologic risk questions can be washed out. Use as-treated only with explicit risk-window construction and informative-censoring adjustment.
- vs treatment-policy estimand under ICH E9(R1): They are close but not identical labels. ICH E9(R1) frames treatment policy as an intercurrent-event strategy: outcomes are used regardless of events such as treatment discontinuation or additional medication. ITT in RWE is the operational emulation of that idea when the initial treatment decision can be observed at a single time zero.
- vs target-trial emulation generally: ITT is one causal contrast within a target-trial protocol. The target trial still must specify eligibility, strategies, time zero, follow-up, outcome, censoring, and summary measure. Calling an analysis "ITT" does not fix immortal time, prevalent-user bias, or baseline confounding if those design elements are wrong.
When NOT to use - and when it is actively misleading
- The question is the effect of sustained adherence. If stakeholders need the effect under "remain on treatment for 12 months" or "follow treat-to-target without rescue therapy," an ITT analysis answers the wrong question. Use per-protocol, clone-censor-weight, or another g-method aligned to the sustained strategy.
- Treatment is defined by future behavior. A label such as "completed 6 cycles" or "received surgery within 90 days" cannot be an ITT arm assigned at time zero unless everyone is cloned into baseline-compatible strategies. If the arm is known only after survival through a grace period, a naive ITT label creates immortal time.
- Follow-up after switching is not observable or comparable. ITT requires outcome capture after discontinuation and switching. If claims enrollment ends when patients leave a plan, or an EHR cannot observe care outside the system, "follow regardless of adherence" is not implemented by the data; informative censoring must be handled explicitly.
- Crossover is so common that the initiation decision no longer has clinical meaning. The ITT effect may still be valid for a policy contrast, but it will not communicate biological efficacy. Report adherence and switching patterns, and include a per-protocol companion rather than overselling the ITT estimate.
- Baseline exchangeability is assumed rather than built. In observational data, the initial treatment strategy is selected by clinicians and patients. ITT-style follow-up does not remove confounding by indication; it must be paired with a defensible new-user design and baseline adjustment or weighting.
Data-source operational depth
- Claims: The cleanest ITT-style RWE contrast is an active-comparator new-user cohort: index date = first qualifying NDC fill after a washout; arm = drug class on that fill; baseline covariates come only from the lookback; follow-up ignores later refill gaps, discontinuation, switching, and add-on therapy unless they are part of censoring or competing-risk rules. Require continuous medical + pharmacy enrollment across lookback and follow-up so outcome capture continues after switching. Medicare Advantage-only person-time can break ITT implementation because the absence of fills or outcomes may be missing FFS data rather than true non-use or no event.
- EHR: The arm should be anchored to an order or administration decision that is captured at time zero. Later treatment changes stay in follow-up, but encounter leakage can make post-switch outcomes invisible; define an active observation rule and model loss to follow-up if it is prognostic. An unfilled prescription order may be a strategy decision but not actual initiation, so state whether the ITT emulates assignment/order or dispensing/start.
- Registry: Strong for assigned treatment, disease severity, and adjudicated outcomes, but often weak for complete medication changes outside registry visits. ITT is feasible when outcomes remain ascertained after off-protocol care; link to claims, pharmacy, or vital records to avoid differential missingness after switching.
- Linked claims-EHR-vital records: Best substrate for ITT-style emulation because the index treatment decision, baseline severity, subsequent outcomes, and death can all be observed. The main risk is selection into the linkable subset; describe the population as linkable patients if linkage determines eligibility or follow-up completeness.
Worked claims example
Question: among adults with type 2 diabetes and stage-3 CKD, what is the 2-year ITT-style effect of initiating an SGLT2 inhibitor vs a DPP-4 inhibitor on hospitalization for heart failure? Eligibility is assessed at the first qualifying fill after 365 days of continuous medical + pharmacy enrollment. The arm is locked from the NDC on that first fill. Baseline confounding is adjusted with overlap weights from pre-index covariates.
Follow-up starts on the fill date and continues for two years regardless of later refill gaps, discontinuation, switch to the other class, add-on therapy, or dose changes. Censor only at structural loss of observable follow-up (disenroll, data end) and handle death according to the pre-specified estimand, for example as a competing event when the endpoint is non-fatal hospitalization.
Report arm-specific 2-year cumulative incidence and the risk difference, plus adherence and switching summaries so readers can see how much usual-care behavior diluted the initiation contrast.
Decision diagram
flowchart LR T0[Time zero: eligible new user fills drug A or B] --> Lock[Lock assigned arm from index treatment] Lock --> Base[Adjust only for baseline covariates] Base --> Follow[Follow for outcome regardless of stop, switch, add-on, or dose change] Follow --> End[Censor only at structural follow-up end or handle death per estimand] End --> Risk[Report arm-specific risk and risk difference]
Worked example
Scenario
A claims analyst compares new users of SGLT2 inhibitors with new users of DPP-4 inhibitors among adults with type 2 diabetes and CKD. Each patient is assigned to the arm of the first qualifying fill after a 365-day washout. The analyst follows everyone for 730 days for hospitalization for heart failure, ignoring later switching or discontinuation because the target estimand is the initiation/treatment-policy effect.
Dataset
Four simplified patients in an ITT-style active-comparator new-user emulation. The assigned arm is fixed at the index fill even when later treatment behavior changes.
| person_id | index_date | assigned_arm | later_treatment_behavior | hhf_date | structural_censor_date | counted_in_assigned_arm_through |
|---|---|---|---|---|---|---|
| P1001 | 2024-01-03 | SGLT2i | continues SGLT2i | 2026-01-03 | 2026-01-03 | |
| P1002 | 2024-02-10 | SGLT2i | switches to DPP-4i on 2024-08-20 | 2025-05-01 | 2026-02-10 | 2026-02-10 |
| P2001 | 2024-01-15 | DPP-4i | stops after first fill | 2025-06-30 | 2025-06-30 | |
| P2002 | 2024-03-01 | DPP-4i | adds SGLT2i on 2024-12-01 | 2024-10-05 | 2026-03-01 | 2026-03-01 |
Steps
Result
The analytic dataset has one row per initiator, fixed assigned arm, baseline-only adjustment variables, a two-year outcome indicator, and a structural censoring time. Later treatment behavior is summarized descriptively but does not reassign or censor the primary ITT analysis.
Trade-offs
Runnable example
Build an ITT-style analytic dataset from an active-comparator new-user cohort. Inputs: cohort : person_id, index_date, assigned_arm, baseline covariates outcomes: person_id, event_date for first endpoint censor : person_id, censor_date for disenrollment/death/data end as defined by the estimand behavior: optional...
import numpy as np
import pandas as pd
import statsmodels.api as sm
HORIZON_DAYS = 730
def make_itt_dataset(cohort, outcomes, censor, horizon_days=HORIZON_DAYS):
dat = cohort.copy()
dat["horizon_date"] = dat["index_date"] + pd.to_timedelta(horizon_days, unit="D")
dat = dat.merge(outcomes.groupby("person_id", as_index=False)["event_date"].min(),
on="person_id", how="left")
dat = dat.merge(censor[["person_id", "censor_date"]], on="person_id", how="left")
dat["analysis_end"] = dat[["horizon_date", "censor_date"]].min(axis=1)
dat["event"] = ((dat["event_date"].notna()) &
(dat["event_date"] <= dat["analysis_end"])).astype(int)
dat["followup_days"] = (dat["analysis_end"] - dat["index_date"]).dt.days.clip(lower=0)
return dat
def overlap_weights(dat, covariates):
x = sm.add_constant(pd.get_dummies(dat[covariates], drop_first=True), has_constant="add")
y = (dat["assigned_arm"] == dat["assigned_arm"].sort_values().unique()[1]).astype(int)
ps = sm.Logit(y, x).fit(disp=False).predict(x).clip(0.01, 0.99)
dat = dat.copy()
dat["ps"] = ps
dat["ow"] = np.where(y == 1, 1 - ps, ps)
return dat
def weighted_itt_risk_difference(dat):
risks = (dat.groupby("assigned_arm")
.apply(lambda g: np.average(g["event"], weights=g["ow"]))
.rename("risk"))
return {"risk_by_arm": risks.to_dict(),
"risk_difference": float(risks.iloc[1] - risks.iloc[0])}ITT-style treatment-policy dataset and overlap-weighted risk difference in R. Post-index switching or stopping is not used to reassign arms or censor follow-up.
library(data.table)
make_itt_dataset <- function(cohort, outcomes, censor, horizon_days = 730L) {
setDT(cohort); setDT(outcomes); setDT(censor)
ev <- outcomes[, .(event_date = min(event_date)), by = person_id]
dat <- merge(copy(cohort), ev, by = "person_id", all.x = TRUE)
dat <- merge(dat, censor[, .(person_id, censor_date)], by = "person_id", all.x = TRUE)
dat[, horizon_date := index_date + horizon_days]
dat[, analysis_end := pmin(horizon_date, censor_date, na.rm = TRUE)]
dat[, event := as.integer(!is.na(event_date) & event_date <= analysis_end)]
dat[, followup_days := pmax(as.integer(analysis_end - index_date), 0L)]
dat[]
}
add_overlap_weights <- function(dat, covariates) {
f <- as.formula(paste("I(assigned_arm == sort(unique(assigned_arm))[2]) ~",
paste(covariates, collapse = " + ")))
ps <- pmin(pmax(predict(glm(f, data = dat, family = binomial()), type = "response"), 0.01), 0.99)
dat[, ps := ps]
dat[, ow := ifelse(assigned_arm == sort(unique(assigned_arm))[2], 1 - ps, ps)]
dat[]
}
weighted_itt_risk_difference <- function(dat) {
risk <- dat[, .(risk = weighted.mean(event, ow)), by = assigned_arm][order(assigned_arm)]
list(risk_by_arm = risk, risk_difference = risk$risk[2] - risk$risk[1])
}SAS ITT-style analytic file and weighted outcome model. The assigned arm comes from the index treatment and is not changed by later switching, discontinuation, or add-on therapy.
%let horizon_days = 730;
proc sql;
create table first_event as
select person_id, min(event_date) as event_date format=date9.
from work.outcomes
group by person_id;
create table itt0 as
select c.*, e.event_date, z.censor_date,
c.index_date + &horizon_days as horizon_date format=date9.,
min(calculated horizon_date, coalesce(z.censor_date, calculated horizon_date)) as analysis_end format=date9.
from work.cohort c
left join first_event e on c.person_id = e.person_id
left join work.censor z on c.person_id = z.person_id;
quit;
data itt;
set itt0;
event = (event_date ne . and event_date <= analysis_end);
followup_days = max(analysis_end - index_date, 0);
run;
proc logistic data=itt noprint;
class assigned_arm(ref='DPP4i') sex(ref='F') / param=ref;
model assigned_arm(event='SGLT2i') = age sex baseline_hosp baseline_ckd baseline_hf;
output out=ps p=ps;
run;
data itt_weighted;
set ps;
ps = min(max(ps, 0.01), 0.99);
if assigned_arm = 'SGLT2i' then ow = 1 - ps;
else ow = ps;
run;
proc genmod data=itt_weighted;
class assigned_arm;
weight ow;
model event = assigned_arm / dist=bin link=identity;
estimate 'SGLT2i minus DPP4i risk difference' assigned_arm 1 -1;
run;Citations
- [1]International Council for Harmonisation. ICH E9(R1) addendum on estimands and sensitivity analysis in clinical trials to the guideline on statistical principles for clinical trials. EMA scientific guideline page.
- [2]Hernan MA, Hernandez-Diaz S. Beyond the intention-to-treat in comparative effectiveness research. Clinical Trials. 2012;9(1):48-55.
- [3]Hernan MA, Robins JM. Per-protocol analyses of pragmatic trials. New England Journal of Medicine. 2017;377(14):1391-1398.