Prevalent User Bias
The bias that arises when a drug cohort includes patients already on treatment at the start of follow-up (prevalent users) rather than restricting to incident initiators at a common time zero, conflating early discontinuers and events with later survivors and inducing depletion of susceptibles, immortal time, and adjustment for post-initiation covariates.
On this page
Prevalent user bias happens when a drug study includes patients who were already taking the medication before the study clock started, rather than starting everyone's clock at their very first fill. Those long-term users survived the earliest, riskiest months on the drug — the patients who had early side effects or stopped early are already gone, leaving a group that looks healthier than a true beginner. As a result, the study underestimates early harms and can make a drug look safer than it really is for someone just starting it. The fix is the new-user design: only enroll patients at their very first fill, so every person's early risk period is actually watched.
Prevalent user bias
(also called survivor bias or depletion-of-susceptibles bias) is the systematic error that occurs when an observational drug study counts person-time and outcomes from patients who were already taking the drug when follow-up began, instead of restricting to incident (new) users whose follow-up starts at first exposure. A prevalent cohort is a left-truncated, conditionally selected sample: to appear in it, a patient had to survive on treatment, free of the outcome and free of intolerable side effects, long enough to still be filling the drug at study entry.
The early high-risk window — the first weeks and months when adverse events, discontinuations, and the strongest treatment effects concentrate — is invisible. What remains is an enriched pool of tolerators who look healthier than a representative initiator, biasing harms toward the null and exaggerating apparent benefit.
Core conceptual distinction
. Three distinct mechanisms ride together under one label, and separating them is what makes the bias tractable.
- Depletion of susceptibles: patients destined to be harmed (or to respond poorly) have already left the population by study entry, so the surviving prevalent users are a selected, lower-risk subset — the hazard you observe is not the hazard a new initiator faces.
- Immortal time / time-zero misalignment: if follow-up starts at a landmark other than initiation (a diagnosis date, an enrollment date, the calendar start of the database), the span between true initiation and study entry is guaranteed event-free by construction, manufacturing apparent protection.
- Adjustment for post-initiation covariates: baseline covariates measured at study entry for a prevalent user are actually consequences of prior treatment (a normalized LDL, a controlled blood pressure, a stable HbA1c), so conditioning on them adjusts away part of the very effect under study or opens collider paths.
The corresponding estimand distinction: prevalent-user analyses target a vague "effect of being on the drug" averaged over an unknowable mix of durations; the new-user design targets the well-defined effect of initiating treatment versus a comparator strategy from a common time zero — the quantity a clinician and a regulator can act on.
Pros, cons, and trade-offs
- vs the new-user (incident-user) design — the canonical fix. New-user restriction sets time zero at first fill after a clean washout, eliminating depletion of susceptibles, immortal time, and post-initiation adjustment in one move. Cost: smaller cohorts and a population of initiators, who may differ from the prevalent users who dominate real-world prescribing; rare or expensive drugs may yield too few incident users. Prefer new-user for essentially every causal question about a treatment's effect, especially safety and early effects.
- vs the prevalent new-user (time-conditional PS) design (Suissa 2017) — a middle path when pure incident users are too few. It matches prevalent users to new users on duration of current use via time-conditional propensity scores, recovering sample size while restoring a defensible time zero for each matched set. Cost: more complex, assumes the duration-conditional exchangeability holds, and still cannot resurrect the unobserved early discontinuers. Prefer it over a naive prevalent cohort whenever incident-only analysis is underpowered.
- vs simply keeping the prevalent/ever-exposed cohort — larger N and superficial "real-world representativeness." But for any initiation question this is the bias you came to remove; it is rarely defensible and routinely overturned when an incident-user re-analysis is run. Acceptable only for descriptive prevalence-of-use questions, never for comparative causal estimates.
When NOT to use — and when it is actively misleading or dangerous
- Do not accept a prevalent-user cohort for any initiation or safety question. Reporting an attenuated or null harm from a prevalent cohort is the dangerous failure mode: depletion of susceptibles hides exactly the early excess risk that pharmacovigilance exists to detect (the rofecoxib and HRT histories are object lessons). A "reassuring" prevalent-user safety result can be affirmatively harmful.
- Do not "fix" a prevalent cohort by adjusting for baseline covariates measured at study entry. Those covariates are post-initiation; adjustment makes the estimate worse, not better, by controlling away the treatment effect or inducing collider bias.
- Do not impose new-user restriction blindly when the drug is genuinely never-stopped and the question is about long-term maintenance — but even then, anchor time zero at initiation and follow forward; the prevalent shortcut is still wrong.
- Do not assume "no fill in the lookback" means new use when the lookback is unobserved (Medicare Advantage-only person-time, a new health-plan enrollee, a registry whose drug-start field is blank). Misclassified prevalent users masquerading as incident users reintroduce the bias silently.
Data-source operational depth
- Claims (FFS or commercial): The defining operation is the washout. Require continuous medical AND pharmacy enrollment across the full lookback (commonly 365 days, sometimes 180) so that "no prior fill" reflects true absence, not unobserved care. The dominant failure mode is Medicare Advantage: MA encounter data are incomplete and FFS pharmacy claims are absent, so a long-time user who switched into your view looks incident — restrict to enrollees with Parts A/B/D (or a commercial medical+pharmacy benefit) and exclude MA-only person-time. Stockpiling, 90-day mail-order, and free samples distort `days_supply` and can hide a prior fill that ran into the lookback. Differential competing risks bite in elderly claims: prevalent users who survived to study entry have differentially lower competing mortality than a representative initiator, further selecting the cohort.
- EHR: Initiation is the order or administration, not the dispense — but a medication-list entry marked "active" often has no reliable start date and may be a carry-over reconciled in from an outside system, the textbook prevalent user wearing a new-user costume. Require a first order after a clean gap and, where possible, link to dispensing to confirm the patient actually started. Visit-driven capture also means a patient who leaves the system is differentially lost, compounding survivor selection.
- Registry: Enrollment date frequently does not equal treatment start; a disease registry may enroll prevalent cases years into therapy. Require an explicit drug-start-date field and treat enrollment-as-time-zero as an immortal-time trap. Link to claims for the full fill history and to a death index to characterize the competing risk.
- Linked claims–EHR–vital records: Best substrate for distinguishing prevalent from incident use (EHR start + claims fill history + mortality), but order/fill/service date discrepancies must be reconciled before assigning time zero, and only the linkable subset is observed (a selection layer on top of the survivor selection).
Worked claims example (depletion of susceptibles in action)
Question: 90-day risk of acute kidney injury (AKI) after starting an ACE inhibitor among adults with hypertension in a commercial + Medicare FFS database. A naive "current user" analyst pulls everyone with an ACEi fill overlapping 2024-01-01 (`fill_date` ≤ index ≤ `fill_date + days_supply`) and follows them 90 days. Suppose the true biology is an early hazard: AKI risk is RR ≈ 2.5 in the first 90 days of initiation, then null.
Among 1,000 true initiators (first fill in 2024 after a 365-day fill-free, continuously A/B/D-enrolled lookback), 60 develop AKI in 90 days (6.0%). But the prevalent pool that overlaps 2024-01-01 is dominated by patients who started in 2021–2023 and kept filling — they have already passed through and survived the early-hazard window (the susceptibles were depleted: those who got AKI stopped the drug, switched, or died). Among 4,000 such prevalent users, only 40 AKI events occur in the next 90 days (1.0%).
A pooled "current user" rate of (60+40)/(1,000+4,000) = 2.0% buries the 6.0% initiation risk, and an unwary safety read concludes ACEi is well tolerated.
The correct construction: index = first ACEi `fill_date` in the study window with no ACEi fill in the prior 365 days and continuous medical+pharmacy FFS enrollment spanning that entire lookback (excluding MA-only person-time so the absence of prior fills is observed, not missing); follow forward 90 days from that fill; censor at disenrollment, death, end of data, and — for an as-treated variant — last `days_supply` end plus a grace period.
The diagnostic that exposes the bias is to count, within the naive cohort, how many "current users" had a fill in the prior 365 days (the would-be-excluded prevalent users) and compare their 90-day event rate to that of the true initiators; a large gap is the depletion-of-susceptibles signature.
Interpreting the output
In the ACEi-and-AKI study, the raw dataset mixes 1,000 new initiators (60 events, 90-day risk = 6.0%) with 4,000 prevalent users (40 events, risk = 1.0%). The pooled naive event rate across all 5,000 patients is 100/5,000 = 2.0%.
(1) Formal interpretation. The pooled rate of 2.0% misrepresents the risk that matters for a prescribing decision — the risk a new patient faces when starting ACEi therapy. Prevalent users have already survived the early high-risk window; their 1.0% event rate reflects depletion of susceptibles (patients who developed early AKI, intolerance, or discontinuation are absent from the prevalent pool), not a lower underlying pharmacological hazard. Mixing these populations suppresses the true 6.0% early-initiation risk by a factor of three.
A new-user design restricted to first prescriptions recovers the 6.0% figure and provides the clinically relevant risk profile for drug-utilization policy and labeling.
(2) Practical interpretation. A drug-safety signal operating in the first 90 days of therapy will be diluted — potentially below statistical detection thresholds — in any analysis that pools new and prevalent users. The contrast of 6.0% versus 2.0% illustrates a threefold dilution: a dataset appearing to show only modest early risk is actually masking a sixfold higher rate in the population most exposed to that risk.
Restricting to new users is not a design preference; it is a precondition for estimating initiation hazards and for detecting time-limited early adverse effects.
Decision diagram
flowchart TD Init[True initiation cohort<br/>all who ever start the drug] --> Early[Early high-risk window<br/>adverse events, intolerance, discontinuation] Early -->|outcome / stop / die| Gone[Susceptibles removed<br/>events + discontinuers leave] Early -->|tolerate + keep filling| Surv[Surviving tolerators] Gone -.invisible to prevalent cohort.-> Study Surv --> Study[Prevalent cohort observed at study entry<br/>enriched for low-risk survivors] Study --> Bias[Observed hazard < initiation hazard<br/>harms biased toward null, benefits inflated] Init --> Fix[New-user restriction:<br/>time zero = first fill after washout] --> Unbiased[Early window observed<br/>estimand = effect of INITIATION]
gantt title Time-zero misalignment - prevalent vs new user (claims) dateFormat YYYY-MM-DD axisFormat %b %Y section Prevalent user (biased) Prior use - early risk UNOBSERVED :crit, prior, 2021-06-01, 2023-12-31 Wrong time zero at study entry :milestone, pt0, 2024-01-01, 0d Observed follow-up (survivors only) :active, pf, 2024-01-01, 90d section New user (correct) Fill-free + continuously enrolled washout :done, wash, 2023-01-01, 2023-12-31 Time zero = first fill :milestone, nt0, 2024-01-01, 0d Observed follow-up (full early window) :active, nf, 2024-01-01, 90d
Worked example
Scenario
A researcher wants to know whether a blood-pressure drug raises the risk of an acute kidney injury (AKI) in the first 90 days of use. The study window opens on 2024-01-01. Two patients both have a fill of the drug on record. Patient A is a new user — her very first fill was on 2024-01-15, well after the study window opened, and she had no fills in the prior 365 days. Patient B is a prevalent user — he has been filling the drug since mid-2022 and simply happened to have a fill overlapping 2024-01-01. The table below shows their pharmacy records. We then trace what each patient's timeline looks like and why including Patient B alongside Patient A produces a biased estimate of AKI risk.
Dataset
Pharmacy fill records for two patients — the columns an analyst sees in a real claims pharmacy table.
| person_id | fill_date | drug | days_supply |
|---|---|---|---|
| A-101 | 2024-01-15 | lisinopril | 30 |
| A-101 | 2024-02-13 | lisinopril | 30 |
| A-101 | 2024-03-14 | lisinopril | 30 |
| B-202 | 2022-07-01 | lisinopril | 90 |
| B-202 | 2022-09-28 | lisinopril | 90 |
| B-202 | 2022-12-26 | lisinopril | 90 |
| B-202 | 2023-03-25 | lisinopril | 90 |
| B-202 | 2023-06-22 | lisinopril | 90 |
| B-202 | 2023-09-19 | lisinopril | 90 |
| B-202 | 2023-12-01 | lisinopril | 90 |
Steps
Result
- Label
Prevalent users are survivors — their early high-risk period is unobserved, so including them biases the AKI rate toward zero (safe-looking) and away from the true initiation risk of ~6%.
- Value
biased_toward_null
Trade-offs
Runnable example
Detect prevalent-user contamination and build a clean incident-user cohort from claims-style inputs. Required inputs (already cleaned and de-duplicated): rx : pharmacy fills -> person_id, fill_date (datetime64), ndc, days_supply enroll : enrollment spans -> person_id, enroll_start, enroll_end, ma_only (bool) #...
import pandas as pd
import numpy as np
WASHOUT_DAYS = 365 # fill-free + continuously enrolled lookback that defines an incident user
FOLLOWUP_DAYS = 90 # early-hazard window in which depletion of susceptibles is most visible
def study_fills(rx: pd.DataFrame, study_ndcs: set[str]) -> pd.DataFrame:
f = rx[rx["ndc"].isin(study_ndcs)].sort_values(["person_id", "fill_date"])
return f
def covered_full_washout(enroll: pd.DataFrame, idx: pd.DataFrame) -> set:
"""person_ids with continuous, FFS-observable enrollment spanning [index-WASHOUT, index]."""
e = enroll.merge(idx[["person_id", "index_date"]], on="person_id")
e["covers"] = ((e["enroll_start"] <= e["index_date"] - pd.Timedelta(days=WASHOUT_DAYS)) &
(e["enroll_end"] >= e["index_date"]) &
(~e["ma_only"]))
return set(e.loc[e["covers"], "person_id"])
def build_incident_cohort_with_diagnostic(rx, enroll, outcomes, study_ndcs, study_start, study_end):
f = study_fills(rx, study_ndcs)
# First study-drug fill inside the study window = candidate time zero.
in_window = f[(f["fill_date"] >= study_start) & (f["fill_date"] <= study_end)]
idx = (in_window.groupby("person_id")["fill_date"].min()
.reset_index().rename(columns={"fill_date": "index_date"}))
# Prevalent flag: any study-drug fill in the WASHOUT_DAYS before the candidate index (the bias signature).
prior = f.merge(idx, on="person_id")
prevalent_ids = set(prior.loc[(prior["fill_date"] < prior["index_date"]) &
(prior["fill_date"] >= prior["index_date"] - pd.Timedelta(days=WASHOUT_DAYS)),
"person_id"])
idx["would_be_prevalent"] = idx["person_id"].isin(prevalent_ids)
# Observable washout (continuous medical+pharmacy FFS enrollment, no MA-only gaps).
observable = covered_full_washout(enroll, idx)
idx = idx[idx["person_id"].isin(observable)].copy()
# Early event within FOLLOWUP_DAYS of each person's time zero.
ev = outcomes.merge(idx[["person_id", "index_date"]], on="person_id")
ev["early"] = ((ev["event_date"] >= ev["index_date"]) &
(ev["event_date"] <= ev["index_date"] + pd.Timedelta(days=FOLLOWUP_DAYS)))
early_ids = set(ev.loc[ev["early"], "person_id"])
idx["early_event"] = idx["person_id"].isin(early_ids)
# Depletion-of-susceptibles diagnostic: incident initiators vs would-be prevalent "current users".
diag = (idx.groupby("would_be_prevalent")["early_event"]
.agg(n="size", events="sum"))
diag["rate"] = diag["events"] / diag["n"]
incident = idx.loc[~idx["would_be_prevalent"], ["person_id", "index_date"]].copy()
incident["baseline_start"] = incident["index_date"] - pd.Timedelta(days=WASHOUT_DAYS)
return incident, diagDetect prevalent-user contamination and build a clean incident-user cohort with data.table. Inputs mirror the Python version: rx : person_id, fill_date (Date), ndc, days_supply enroll : person_id, enroll_start (Date), enroll_end (Date), ma_only (logical) outcomes : person_id, event_date (Date) study_ndcs is a...
library(data.table)
WASHOUT_DAYS <- 365L
FOLLOWUP_DAYS <- 90L
build_incident_cohort_with_diagnostic <- function(rx, enroll, outcomes, study_ndcs,
study_start, study_end) {
setDT(rx); setDT(enroll); setDT(outcomes)
f <- rx[ndc %chin% study_ndcs][order(person_id, fill_date)]
# First study-drug fill inside the study window = candidate time zero.
idx <- f[fill_date >= study_start & fill_date <= study_end,
.(index_date = min(fill_date)), by = person_id]
# Prevalent flag: any study-drug fill in the washout window before candidate index.
prior <- merge(f, idx, by = "person_id")
prevalent_ids <- unique(prior[fill_date < index_date &
fill_date >= index_date - WASHOUT_DAYS, person_id])
idx[, would_be_prevalent := person_id %chin% prevalent_ids]
# Observable washout: continuous medical+pharmacy FFS enrollment, no MA-only gaps.
e <- merge(enroll, idx[, .(person_id, index_date)], by = "person_id")
observable <- e[enroll_start <= index_date - WASHOUT_DAYS &
enroll_end >= index_date & !ma_only, unique(person_id)]
idx <- idx[person_id %chin% observable]
# Early event within FOLLOWUP_DAYS of each person's time zero.
ev <- merge(outcomes, idx[, .(person_id, index_date)], by = "person_id")
early_ids <- unique(ev[event_date >= index_date &
event_date <= index_date + FOLLOWUP_DAYS, person_id])
idx[, early_event := person_id %chin% early_ids]
# Depletion-of-susceptibles diagnostic.
diag <- idx[, .(n = .N, events = sum(early_event)), by = would_be_prevalent]
diag[, rate := events / n]
incident <- idx[would_be_prevalent == FALSE, .(person_id, index_date)]
incident[, baseline_start := index_date - WASHOUT_DAYS]
list(incident = incident, diagnostic = diag)
}Detect prevalent-user contamination and build a clean incident-user cohort in SAS via PROC SQL. Required input datasets (post data-management), with study-drug NDCs already restricted into work.rx: work.rx : person_id, fill_date, ndc, days_supply (only drug-of-interest fills) work.enroll : person_id, enroll_start,...
%let washout = 365;
%let followup = 90;
%let studystart = '01JAN2024'd;
%let studyend = '31DEC2024'd;
/* Candidate time zero = first study-drug fill inside the study window. */
proc sql;
create table idx as
select person_id, min(fill_date) as index_date format=date9.
from work.rx
where fill_date between &studystart and &studyend
group by person_id;
quit;
/* Prevalent flag: any study-drug fill in the washout window before candidate index (the bias signature). */
proc sql;
create table idx_flag as
select i.person_id, i.index_date,
(exists (select 1 from work.rx p
where p.person_id = i.person_id
and p.fill_date < i.index_date
and p.fill_date >= i.index_date - &washout)) as would_be_prevalent
from idx i;
quit;
/* Keep only observable washout: continuous medical+pharmacy FFS enrollment, no MA-only spans. */
proc sql;
create table idx_obs as
select f.*
from idx_flag f
where exists (select 1 from work.enroll e
where e.person_id = f.person_id
and e.ma_only = 0
and e.enroll_start <= f.index_date - &washout
and e.enroll_end >= f.index_date);
quit;
/* Early event within the follow-up window of each person's time zero. */
proc sql;
create table idx_ev as
select o.person_id, o.index_date, o.would_be_prevalent,
(exists (select 1 from work.outcomes ev
where ev.person_id = o.person_id
and ev.event_date >= o.index_date
and ev.event_date <= o.index_date + &followup)) as early_event
from idx_obs o;
quit;
/* Depletion-of-susceptibles diagnostic: early event rate by would-be-prevalent status. */
proc sql;
create table work.dos_diagnostic as
select would_be_prevalent,
count(*) as n,
sum(early_event) as events,
mean(early_event) as rate format=percent8.1
from idx_ev
group by would_be_prevalent;
quit;
/* Clean incident cohort with baseline covariate window for downstream PS/outcome modeling. */
proc sql;
create table work.incident as
select person_id, index_date,
index_date - &washout as baseline_start format=date9.
from idx_ev
where would_be_prevalent = 0;
quit;Citations
- [1]Ray WA. Evaluating medication effects outside of clinical trials: new-user designs. American Journal of Epidemiology. 2003;158(9):915-920.
- [2]Suissa S, Moodie EEM, Dell'Aniello S. Prevalent new-user cohort designs for comparative drug effect studies by time-conditional propensity scores. Pharmacoepidemiology and Drug Safety. 2017;26(4):459-468.
- [3]Suissa S. Immortal time bias in pharmacoepidemiology. American Journal of Epidemiology. 2008;167(4):492-499.
- [4]Hernán MA, Alonso A, Logan R, et al. Observational studies analyzed like randomized experiments: an application to postmenopausal hormone therapy and coronary heart disease. Epidemiology. 2008;19(6):766-779.
- [5]Danaei G, Tavakkoli M, Hernán MA. Bias in observational studies of prevalent users: lessons for comparative effectiveness research from a meta-analysis of statins. American Journal of Epidemiology. 2012;175(4):250-262.