Clone-Censor-Weight for Per-Protocol Target-Trial Emulation
A target-trial-emulation technique that estimates the per-protocol effect of treatment strategies not distinguishable at baseline by replicating each person once per strategy (cloning), artificially censoring each clone when its observed data first deviate from its assigned strategy, and applying inverse-probability-of-censoring weights to remove the selection bias that the artificial censoring induces.
On this page
Clone-censor-weight (CCW) is a technique for comparing two treatment strategies using real-world data when everyone in the study could plausibly follow either strategy on day one. The trick is that each person is placed into both groups simultaneously by creating two copies called clones; each clone is followed over time and removed from its group the moment the real person's behavior diverges from that group's rule, and then a statistical weight is applied to compensate for those removals so the comparison stays fair. This lets researchers answer a per-protocol question: what would the outcome have been if patients actually stuck to their assigned strategy from start to finish?
Clone-censor-weight (CCW)
is the standard device for estimating the per-protocol effect of treatment strategies that two or more eligible patients could still satisfy at time zero but that diverge only later — the setting where a naive design has no clean way to assign an arm. Classic examples are sustained strategies ("initiate treatment within a grace period and stay on it" vs "never initiate"), duration strategies ("treat for 12 months" vs "treat for 6 months"), and threshold/dynamic strategies ("start when eGFR crosses X").
Because every eligible person is, at baseline, compatible with every strategy, CCW makes that compatibility explicit: it clones each person into one copy per strategy, follows each clone under its assigned rule, artificially censors a clone the instant its observed behavior departs from that rule, and then weights the surviving clone-time by the inverse probability of remaining uncensored so the censored clones are represented by similar still-adherent clones.
Core estimand distinction
CCW targets the observational analog of the per-protocol effect of a sustained strategy — what the cumulative incidence (or hazard) would have been had everyone followed the assigned strategy exactly. This is not the intention-to-treat (initiation) contrast, which an active-comparator new-user design with baseline propensity-score adjustment already estimates cleanly. It is also not the naive "as-treated" or "ever-exposed" contrast, which classifies person-time by post-baseline behavior and so re-imports immortal time, selection on adherence, and time-varying confounding.
The cloning step exists precisely because the strategies are not distinguishable at t0: assigning the arm from baseline data alone (as one would in active-comparator new-user) is impossible for "start within 30 days and persist," so instead of choosing an arm we place each person in both arms and let the data censor the incompatible clone later.
The inverse-probability-of-censoring weighting (IPCW) that follows is, mechanically, a marginal structural model fit on a cloned scaffold; CCW is best understood as an MSM whose person-time has been duplicated to encode strategy membership.
Pros, cons, and trade-offs
- vs naive as-treated / ever-exposed: CCW eliminates immortal-time bias (no person-time is counted before the deviation that defines non-adherence), selection bias from post-baseline arm assignment, and confounding by time-varying factors that drive both adherence and the outcome. Cost: it answers a hypothetical full-adherence question that can diverge from real-world adherence, and it requires a correctly specified IPCW model. Prefer CCW for any sustained/dynamic strategy where as-treated is the only alternative — as-treated is almost always biased here.
- vs landmark analysis: Landmark conditions on survival and exposure status up to a fixed landmark time, discards everyone who had the event or deviated before it, and answers a question about the selected survivors. CCW clones everyone at baseline and recovers the full eligible population through weighting, so it avoids the survivor selection landmark builds in and accommodates grace periods naturally. Cost: heavier specification (two weighting models, weight diagnostics) and bootstrap inference. Prefer landmark only when the scientific question really is about a post-baseline conditioning set; otherwise CCW is the more defensible per-protocol estimator.
- vs marginal structural models / g-methods without cloning: A standard MSM or the g-formula can estimate sustained-strategy effects directly without cloning. CCW's advantage is interpretability and protocol transparency — the cloning + censoring rules force you to write down the strategy, the grace period, and the deviation definition exactly as a trial protocol would, which is why regulators and reviewers favor it. Cost: the clone expansion multiplies the dataset, induces within-person correlation (requiring robust/bootstrap variance), and can be less statistically efficient than a well-specified g-formula. Prefer CCW when explicit protocol/estimand specification and auditability matter (regulatory submissions, head-to-head sustained strategies); consider plain g-methods or TMLE on the cloned data when efficiency or double robustness is the priority.
When NOT to use — and when it is actively misleading
- The strategies are fully distinguishable at baseline (e.g., drug A initiator vs drug B initiator, both single point decisions). Then an active-comparator new-user design with IPTW is simpler, more efficient, and equally valid; cloning adds dataset size and variance for no gain.
- Deviation is a collider / informed by the outcome process. If patients stop because of early toxicity or declining health that itself predicts the outcome, and those drivers are unmeasured, the IPCW "no unmeasured determinants of censoring" assumption fails and CCW is biased — frequently more convincingly biased than a crude analysis because the machinery lends false credibility.
- The grace period is so short (or the strategy so demanding) that the eligible-and-still-adherent stratum collapses. A handful of persistently adherent clones near the end of follow-up receive enormous weights; effective sample size craters and the estimate becomes a high-variance artifact of a few people.
- Competing risks differ by strategy and are ignored. In elderly claims cohorts, death competes with the event and may itself depend on the strategy; a cause-specific-hazard CCW that censors at the competing event answers a different (and often less policy-relevant) question than one targeting the subdistribution / cumulative incidence. Pre-specify which.
- Adherence is unobservable in the data. In Medicare Advantage-only person-time, fills are not captured, so the deviation rule cannot fire correctly and clones are censored (or not) spuriously — restrict to fee-for-service / full-benefit person-time before applying CCW.
Data-source operational depth
- Claims (FFS / commercial): Treatment status over time is reconstructed from pharmacy fills (`fill_date` + `days_supply`, stitched into on-treatment episodes with a stockpiling/grace rule). Require continuous medical + pharmacy enrollment so a "no fill" gap is a true gap, not missingness. Failure modes: MA-only person-time lacks FFS pharmacy/medical claims, so the adherence/deviation rule misfires — exclude it; differential competing risks (death) by strategy in elderly cohorts bias a cause-specific CCW unless handled; immortal time sneaks back in if the grace period is treated as "guaranteed survival" rather than encoded as eligible-but-uncensored clone-time; 90-day mail-order and free samples distort `days_supply` and therefore the deviation timing.
- EHR: Orders + administrations give finer "on-treatment" granularity and capture reasons for deviation (toxicity, response), which is invaluable for arguing the IPCW assumption — but visit-driven capture means a patient who leaves the system looks like a deviation/censoring event when they merely changed providers. Link to fills to confirm the patient actually started, and model loss to follow-up as a separate censoring process.
- Registry: Structured start/stop dates and adjudicated outcomes support clean cloning and censoring rules and are common in embedded/registry-based emulations; weak for complete longitudinal drug exposure, so link to claims for the full fill history and to a death index to firm up the competing-risk handling.
- Linked claims-EHR-vital records: The ideal substrate (EHR reasons-for-deviation + claims completeness + reliable mortality for competing risks), at the cost of linkage selection and reconciling order/fill/service-date discrepancies before time-zero and deviation timing are assigned.
Worked claims example
Question: does initiating a statin within 6 months of a first MI and persisting vs never initiating reduce 1-year all-cause mortality, in a commercial + Medicare fee-for-service database? (1) Eligibility: adults with an incident MI hospitalization, 365 days of continuous A/B/D (or commercial medical+pharmacy) enrollment before discharge, and no statin fill in that lookback.
Time zero = discharge date. (2) Strategies: S1 = fill a statin within a 180-day grace period and remain covered thereafter; S0 = never fill a statin. (3) Cloning: each eligible person becomes two clones, one assigned S1 and one assigned S0, both starting at t0.
(4) Artificial censoring: the S0 clone is censored on the first day a statin fill is observed; the S1 clone is censored at day 180 if no statin has been filled by then (failed the grace period) and, after a fill, on the first day the patient's covered supply lapses beyond a 30-day permissible gap (failed persistence). A clone that has the outcome before deviating is not censored — its event counts.
(5) IPCW: fit a pooled logistic model per arm for the probability of remaining uncensored in each month given time-varying covariates (recent hospitalizations, cardiac procedures, comorbidity flags, prior adherence proxies) and baseline covariates; the stabilized weight for a clone-month is the cumulative product of those inverse probabilities. (6) Estimation: fit a weighted pooled logistic outcome model on the clone-months with a flexible function of time and a strategy indicator; convert the fitted discrete hazards into standardized 1-year cumulative-incidence curves and report the risk difference and risk ratio at 12 months.
(7) Inference & diagnostics: nonparametric bootstrap over persons (resample whole persons, re-clone, refit weights and outcome model) for confidence intervals; report unique N vs expanded clone-months, the stabilized-weight distribution and effective sample size, and sensitivity analyses to the grace-period length (90 vs 180 vs 270 days), the permissible-gap rule, and weight truncation at the 1st/99th percentiles.
Interpreting the output
Using the worked example: Clone S1 (initiate-and-persist) contributes 365 uncensored days; Clone S0 (never-initiate) is artificially censored at day 100 when Maria fills her first statin. After IPCW reweighting across the full cohort, a per-protocol comparison yields — for illustration — HR = 0.74 (95% CI 0.61–0.90) for 1-year mortality.
Formal interpretation: The clone-censor-weight HR of 0.74 is the per-protocol effect of the initiate-and-persist strategy versus the never-initiate strategy — the causal effect that would be observed if all patients in the target population had adhered to their assigned strategy throughout follow-up. It is not an intent-to-treat estimate (which includes non-adherers) and does not apply to the as-treated population. The IPCW step is essential: without reweighting, patients who deviated (like Maria's Clone S0 at day 100) would be dropped, leaving a selected, healthier subgroup.
The estimate is valid under two untestable conditions: no unmeasured confounding conditional on covariates entering the IPCW model, and correct specification of the censoring model. Confidence intervals must be obtained by cluster bootstrap over original patients — not over expanded clone-months — to respect the within-patient correlation introduced by cloning.
Practical interpretation: Patients who initiated a statin within 180 days and remained on it would, on average, have died at a 26% lower rate than those who never started — if the adherence assumptions and no-unmeasured-confounding conditions hold. The clone-censor-weight framework makes this "what if everyone had followed the strategy" question answerable from observational data without requiring treatment to have been randomized.
Decision diagram
flowchart TD
Elig[Eligible at t0: satisfies BOTH strategies<br/>no prior statin, continuous enrollment] --> Clone{Clone once per strategy}
Clone -->|assign S1: initiate within grace + persist| A0[Clone A starts at t0 under S1]
Clone -->|assign S0: never initiate| B0[Clone B starts at t0 under S0]
A0 --> Acens[Artificially censor clone A at<br/>grace-period failure OR persistence lapse]
B0 --> Bcens[Artificially censor clone B at<br/>first statin fill]
Acens --> IPCW[Stabilized IPCW per arm<br/>pooled logistic: P uncensored each month]
Bcens --> IPCW
IPCW --> Out[Weighted pooled-logistic outcome model<br/>standardize to 1-year cumulative incidence]
Out --> Boot[Person-level bootstrap CIs<br/>weight + ESS diagnostics, sensitivity to grace/gap/truncation]gantt title One observed patient becomes two clones with different censor times dateFormat YYYY-MM-DD axisFormat %b %Y section Observed (single patient) MI discharge = time zero :milestone, t0, 2024-01-01, 0d No statin until day 90, then persistent fills :obs, 2024-01-01, 365d section Clone S1 (initiate within 180d + persist) Eligible, uncensored: filled day 90 within grace, stays covered :active, s1, 2024-01-01, 365d Outcome or admin censoring ends follow-up :crit, s1e, 2024-12-31, 1d section Clone S0 (never initiate) Uncensored until first fill :done, s0, 2024-01-01, 89d Artificially censored at first statin fill (day 90) :crit, s0c, 2024-03-31, 1d
Worked example
Scenario
A researcher wants to know whether starting a statin within 180 days of a heart attack and staying on it reduces one-year mortality compared to never starting one. Using a claims database, every eligible patient is cloned into two copies at hospital discharge (day 0). Clone S1 is assigned to the initiate-and-persist strategy; Clone S0 is assigned to the never-initiate strategy. We follow one patient, Maria, to see exactly when and why each of her clones gets artificially censored.
Dataset
Maria's statin fill history after her MI discharge on 2024-01-01, as it would appear in a pharmacy claims table.
| person_id | fill_date | drug | days_supply |
|---|---|---|---|
| M001 | 2024-04-10 | atorvastatin | 90 |
| M001 | 2024-07-05 | atorvastatin | 90 |
Steps
Result
Clone S1 contributes 365 uncensored days under the initiate-and-persist strategy. Clone S0 contributes 100 days before artificial censoring at day 100 (2024-04-10). The per-protocol analysis weights Clone S0 records to account for this early removal, then compares 1-year mortality rates across both arms using those weights.
Trade-offs
Runnable example
Clone-censor-weight per-protocol emulation of a sustained "initiate within grace period and persist" vs "never initiate" strategy from claims-style inputs. Required inputs (already cleaned, de-duplicated): cohort : one row per eligible new-MI patient -> person_id, t0 (index/discharge date), <baseline covariates>...
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
GRACE_DAYS = 180 # window to initiate the statin and remain "adherent" to S1
GAP_DAYS = 30 # permissible gap before a persistence lapse counts as deviation
HORIZON_M = 12 # months of follow-up for the 1-year risk contrast
def build_person_periods(cohort, rx, outcome, fup_end):
"""One row per person-month from t0 to t0+HORIZON_M, with time-varying on-treatment + event flags."""
rows = []
rx_by = {p: g.sort_values("fill_date") for p, g in rx.groupby("person_id")}
for r in cohort.itertuples(index=False):
pid, t0 = r.person_id, r.t0
ev = outcome.loc[outcome.person_id == pid, "event_date"]
ev_date = ev.min() if len(ev) else pd.NaT
adm = fup_end.loc[fup_end.person_id == pid, "fup_end_date"]
adm_date = adm.min() if len(adm) else pd.NaT
fills = rx_by.get(pid, pd.DataFrame(columns=["fill_date", "days_supply"]))
# covered-day intervals stitched with stockpiling (carry-over of surplus supply)
covered_until = t0 - pd.Timedelta(days=1)
intervals = []
for f in fills.itertuples(index=False):
start = max(f.fill_date, covered_until + pd.Timedelta(days=1))
covered_until = max(covered_until, f.fill_date) + pd.Timedelta(days=int(f.days_supply))
intervals.append((start, covered_until))
first_fill = fills["fill_date"].min() if len(fills) else pd.NaT
for m in range(HORIZON_M):
m_start = t0 + pd.Timedelta(days=30 * m)
m_end = t0 + pd.Timedelta(days=30 * (m + 1)) - pd.Timedelta(days=1)
if pd.notna(adm_date) and adm_date < m_start:
break
on_tx = any(s <= m_end and e + pd.Timedelta(days=GAP_DAYS) >= m_start for s, e in intervals)
started_by = pd.notna(first_fill) and first_fill <= m_end
event_m = int(pd.notna(ev_date) and m_start <= ev_date <= m_end)
rows.append(dict(person_id=pid, month=m, on_tx=int(on_tx),
started_by=int(started_by),
days_since_t0=(m_start - t0).days, event=event_m))
if event_m:
break
pp = pd.DataFrame(rows)
return pp.merge(cohort, on="person_id", how="left")
def expand_and_censor(pp):
"""Duplicate each person-month into clone S1 and clone S0; set artificial censoring per the strategy."""
out = []
for strat in ("S1", "S0"):
c = pp.copy()
c["strategy"] = strat
c["clone_id"] = c["person_id"].astype(str) + "_" + strat
if strat == "S0": # never initiate -> censor at first treatment
c["artif_cens"] = (c["on_tx"] == 1).astype(int)
else: # initiate within grace, then persist
grace_fail = (c["days_since_t0"] > GRACE_DAYS) & (c["started_by"] == 0)
persist_fail = (c["started_by"] == 1) & (c["on_tx"] == 0)
c["artif_cens"] = (grace_fail | persist_fail).astype(int)
out.append(c)
x = pd.concat(out, ignore_index=True)
# An event in a month overrides artificial censoring: the event is observed, not censored.
x.loc[x["event"] == 1, "artif_cens"] = 0
return x.sort_values(["clone_id", "month"])
def stabilized_ipcw(clones, tv_covs):
"""Pooled-logistic IPCW: P(uncensored this month). Stabilized weight = prod(num)/prod(denom)."""
clones = clones.copy()
clones["uncens"] = 1 - clones["artif_cens"]
rhs_den = "bs(days_since_t0, df=4) + " + " + ".join(tv_covs)
rhs_num = "bs(days_since_t0, df=4)"
w = []
for strat, g in clones.groupby("strategy"):
den = smf.logit("uncens ~ " + rhs_den, data=g).fit(disp=0)
num = smf.logit("uncens ~ " + rhs_num, data=g).fit(disp=0)
g = g.assign(p_den=den.predict(g), p_num=num.predict(g))
g = g.sort_values(["clone_id", "month"])
g["sw"] = (g.groupby("clone_id")["p_num"].cumprod() /
g.groupby("clone_id")["p_den"].cumprod())
w.append(g)
out = pd.concat(w, ignore_index=True)
out["sw"] = out["sw"].clip(upper=out["sw"].quantile(0.99)) # truncate extreme weights
return out
def run_ccw(cohort, rx, outcome, fup_end, tv_covs=("on_tx",)):
pp = build_person_periods(cohort, rx, outcome, fup_end)
clones = expand_and_censor(pp)
wdat = stabilized_ipcw(clones, list(tv_covs))
# Weighted pooled-logistic outcome model (discrete-time hazard) on uncensored clone-months.
m = smf.glm("event ~ strategy + bs(days_since_t0, df=4)",
data=wdat[wdat["artif_cens"] == 0],
family=__import__("statsmodels.api", fromlist=["families"]).families.Binomial(),
freq_weights=wdat.loc[wdat["artif_cens"] == 0, "sw"]).fit()
# Standardize: predict monthly hazard under each strategy, convert to 1-year cumulative incidence.
risks = {}
grid = wdat[["days_since_t0"]].drop_duplicates().sort_values("days_since_t0")
for strat in ("S1", "S0"):
g = grid.assign(strategy=strat)
h = m.predict(g).values
risks[strat] = 1 - np.prod(1 - h)
return dict(model=m, risk_S1=risks["S1"], risk_S0=risks["S0"],
risk_difference=risks["S1"] - risks["S0"])Clone-censor-weight per-protocol emulation in R with data.table + splines, mirroring the Python pipeline. Inputs: cohort : person_id, t0 (Date), <baseline covariates> rx : person_id, fill_date (Date), days_supply (integer) outcome : person_id, event_date (Date) # first all-cause death fup_end : person_id,...
library(data.table)
library(splines)
GRACE_DAYS <- 180L; GAP_DAYS <- 30L; HORIZON_M <- 12L
build_person_periods <- function(cohort, rx, outcome, fup_end) {
setDT(cohort); setDT(rx); setDT(outcome); setDT(fup_end); setorder(rx, person_id, fill_date)
out <- list()
for (i in seq_len(nrow(cohort))) {
pid <- cohort$person_id[i]; t0 <- cohort$t0[i]
ev <- min(outcome[person_id == pid, event_date], Inf)
adm <- min(fup_end[person_id == pid, fup_end_date], Inf)
f <- rx[person_id == pid]
# stitch covered-day intervals with stockpiling carry-over
covered_until <- t0 - 1L; iv <- list(); first_fill <- if (nrow(f)) min(f$fill_date) else as.Date(NA)
if (nrow(f)) for (j in seq_len(nrow(f))) {
start <- max(f$fill_date[j], covered_until + 1L)
covered_until <- max(covered_until, f$fill_date[j]) + f$days_supply[j]
iv[[length(iv) + 1L]] <- c(start, covered_until)
}
for (m in 0:(HORIZON_M - 1L)) {
ms <- t0 + 30L * m; me <- t0 + 30L * (m + 1L) - 1L
if (is.finite(adm) && adm < ms) break
on_tx <- any(vapply(iv, function(z) z[1] <= me && z[2] + GAP_DAYS >= ms, logical(1)))
started_by <- !is.na(first_fill) && first_fill <= me
event_m <- as.integer(is.finite(ev) && ev >= ms && ev <= me)
out[[length(out) + 1L]] <- data.table(person_id = pid, month = m,
on_tx = as.integer(on_tx), started_by = as.integer(started_by),
days_since_t0 = as.integer(ms - t0), event = event_m)
if (event_m == 1L) break
}
}
merge(rbindlist(out), cohort, by = "person_id")
}
expand_and_censor <- function(pp) {
mk <- function(strat) {
c <- copy(pp); c[, `:=`(strategy = strat, clone_id = paste0(person_id, "_", strat))]
if (strat == "S0") c[, artif_cens := as.integer(on_tx == 1L)]
else c[, artif_cens := as.integer((days_since_t0 > GRACE_DAYS & started_by == 0L) |
(started_by == 1L & on_tx == 0L))]
c
}
x <- rbindlist(list(mk("S1"), mk("S0")))
x[event == 1L, artif_cens := 0L] # observed events are not artificially censored
setorder(x, clone_id, month); x[]
}
stabilized_ipcw <- function(clones, tv_covs = c("on_tx")) {
clones[, uncens := 1L - artif_cens]
fden <- as.formula(paste("uncens ~ bs(days_since_t0, df = 4) +", paste(tv_covs, collapse = " + ")))
fnum <- uncens ~ bs(days_since_t0, df = 4)
res <- clones[, {
den <- glm(fden, data = .SD, family = binomial()); num <- glm(fnum, data = .SD, family = binomial())
pd <- predict(den, .SD, type = "response"); pn <- predict(num, .SD, type = "response")
ord <- order(clone_id, month)
.SD[ord][, sw := ave(pn[ord], clone_id[ord], FUN = cumprod) /
ave(pd[ord], clone_id[ord], FUN = cumprod)][]
}, by = strategy]
res[, sw := pmin(sw, quantile(sw, 0.99))] # truncate extreme weights
res[]
}
run_ccw <- function(cohort, rx, outcome, fup_end, tv_covs = c("on_tx")) {
pp <- build_person_periods(cohort, rx, outcome, fup_end)
w <- stabilized_ipcw(expand_and_censor(pp), tv_covs)
fit <- glm(event ~ strategy + bs(days_since_t0, df = 4),
data = w[artif_cens == 0L], family = binomial(), weights = w[artif_cens == 0L, sw])
grid <- unique(w[, .(days_since_t0)])[order(days_since_t0)]
risk <- sapply(c("S1", "S0"), function(s) {
h <- predict(fit, cbind(grid, strategy = s), type = "response"); 1 - prod(1 - h)
})
list(fit = fit, risk_S1 = risk["S1"], risk_S0 = risk["S0"],
risk_difference = unname(risk["S1"] - risk["S0"]))
}Clone-censor-weight per-protocol emulation in SAS. Required input datasets (post data-management): work.pp : LONG person-period table from claims, one row per person-month over the 12-month horizon, with person_id, t0, month, days_since_t0, on_tx (0/1 covered this month), started_by (0/1 ever filled by this month),...
%let grace = 180; %let horizon = 12;
/* (1) CLONE: duplicate each person-month into strategy S1 (initiate+persist) and S0 (never initiate). */
data clones;
set work.pp;
length strategy $2 clone_id $40;
do strategy = 'S1', 'S0';
clone_id = catx('_', put(person_id, best12.), strategy);
/* (2) ARTIFICIAL CENSORING per the assigned strategy. */
if strategy = 'S0' then artif_cens = (on_tx = 1); /* censor never-arm at first treatment */
else artif_cens = ( (days_since_t0 > &grace and started_by = 0) /* failed grace-period initiation */
or (started_by = 1 and on_tx = 0) ); /* failed persistence (gap > permissible)*/
if event = 1 then artif_cens = 0; /* observed events are not censored */
uncens = 1 - artif_cens;
output;
end;
run;
proc sort data=clones; by strategy clone_id month; run;
/* (3) IPCW denominator and numerator models, fit separately by strategy, pooled over person-months. */
proc logistic data=clones noprint;
by strategy;
model uncens(event='1') = days_since_t0 days_since_t0*days_since_t0 on_tx; /* + baseline + time-varying covs */
output out=pden p=p_den;
run;
proc logistic data=clones noprint;
by strategy;
model uncens(event='1') = days_since_t0 days_since_t0*days_since_t0; /* stabilization numerator (time) */
output out=pnum p=p_num;
run;
/* Stabilized weight = running product of numerator / running product of denominator within each clone. */
data weights;
merge pden(keep=clone_id strategy month p_den) pnum(keep=clone_id month p_num);
by strategy clone_id month; /* sorted upstream; merge aligns the two predicted-probability streams */
retain cum_num cum_den;
by strategy clone_id;
if first.clone_id then do; cum_num = 1; cum_den = 1; end;
cum_num = cum_num * p_num;
cum_den = cum_den * p_den;
sw = cum_num / cum_den;
run;
/* Truncate extreme stabilized weights at the 99th percentile. */
proc univariate data=weights noprint; var sw; output out=p99 pctlpts=99 pctlpre=p; run;
data analytic;
if _n_ = 1 then set p99;
set weights;
if sw > p99 then sw = p99;
where artif_cens = 0; /* outcome model uses uncensored clone-months only */
run;
/* join baseline covariates / event back on by clone_id month before fitting if needed */
/* (4) WEIGHTED POOLED-LOGISTIC outcome model (discrete-time hazard) -> strategy contrast. */
proc genmod data=analytic;
class strategy clone_id;
weight sw;
model event(event='1') = strategy days_since_t0 days_since_t0*days_since_t0 / dist=bin link=logit;
repeated subject=clone_id / type=ind; /* robust container; bootstrap over person_id for valid CIs */
estimate 'S1 vs S0 log-OR (per-month hazard)' strategy 1 -1;
run;
/* Standardize the fitted monthly hazards under each strategy to 1-year cumulative incidence in a DATA step:
CI_s = 1 - PROD over months of (1 - hazard_s(month)), then risk difference = CI_S1 - CI_S0. */Citations
- [1]Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183(8):758-764.
- [2]Hernán MA, Sauer BC, Hernández-Díaz S, Platt R, Shrier I. Specifying a target trial prevents immortal time bias and other self-inflicted injuries in observational analyses. J Clin Epidemiol. 2016;79:70-75.
- [3]Maringe C, Benitez Majano S, Exarchakou A, et al. Reflection on modern methods: trial emulation in the presence of immortal-time bias. Assessing the benefit of major surgery for elderly lung cancer patients using observational data. Int J Epidemiol. 2020;49(5):1719-1729.