Prospective Cohort Study
An observational design that defines a group by exposure status at a fixed time origin and follows it forward in calendar time to ascertain incident outcomes, so that exposure and covariates are recorded before the outcome occurs.
On this page
A prospective cohort study starts with a group of people sorted by whether they did or did not get a treatment, marks that starting moment as everyone's day zero, and then watches them move forward in time to see who later develops the outcome you care about. Because you decide who's in which group before any outcomes happen, you get to measure things in the same order you want to reason about them: cause first, effect second. The payoff is that you can directly count how often the outcome happens (the actual risk), and you can track several different outcomes from the same starting group. The catch is that it's slow and wasteful when the outcome is very rare, since you have to watch a lot of people to catch a few events.
A prospective cohort study classifies people by exposure at a defined time origin (time zero) and then follows them forward to observe who develops the outcome The defining feature is temporal: exposure status and baseline covariates are fixed and recorded before any outcome is known, so the direction of measurement matches the direction of inference
This is what distinguishes it from a retrospective cohort (where both exposure and outcome have already occurred when the investigator looks) and from a case-control study (which samples on the outcome and looks backward at exposure)
In real-world data the "prospective" label is often about analytic posture rather than wall-clock timing: a study built in an administrative database is technically conducted on already-accrued records, but it is designed and analyzed prospectively when the protocol fixes eligibility, time zero, exposure, and covariate windows a priori and follows each person forward from time zero — the structure Hernán and Robins formalize as target-trial emulation.
Core conceptual / estimand distinction
A cohort is a sampling-and-follow-up frame, not an estimator The design fixes who is in the risk set, when their clock starts, and what counts as person-time at risk; the estimand (cumulative incidence, incidence rate, hazard ratio, risk difference, restricted mean survival time) and the confounding-control strategy (restriction, matching, propensity scores, g-methods) are layered on top
The cohort frame delivers two things a case-control design cannot: it lets you estimate absolute risks and rates directly (numerator events over a denominator of person-time you actually counted), and it lets you study multiple outcomes from a single exposure definition Getting the frame right — one unambiguous time zero per person, exposure assigned at time zero, covariates measured before time zero, follow-up that begins at time zero — is the work; nearly every notorious cohort bias is a violation of one of those four rules.
Pros, cons, and trade-offs
- vs case-control: A cohort yields incidence and absolute risk, handles many outcomes at once, and avoids the recall and selection biases that plague exposure ascertainment after the outcome is known. Cost: it is inefficient for rare outcomes (you must follow large denominators for few events) and for outcomes with long induction periods. Prefer a cohort when the exposure is rare or when absolute risk / multiple outcomes matter; prefer nested case-control or case-cohort sampling within the cohort when an expensive covariate (biomarker, chart abstraction, adjudication) must be collected and the outcome is rare.
- vs retrospective cohort: Prospective measurement (or prospective design in RWD) lets you specify exposure and confounders before knowing outcomes, which protects against data-driven definition tweaking and against conditioning on post-baseline variables. Cost: in primary-data collection it is slower and more expensive; in RWD the trade is that you are limited to variables the data captured, captured before time zero.
- vs cross-sectional: A cohort establishes temporality (exposure precedes outcome), the single most important ingredient for causal interpretation, which a cross-sectional snapshot cannot. Cost: follow-up infrastructure, loss to follow-up, and competing risks.
- vs RCT: A cohort can study harms, long-term outcomes, rare exposures, and populations excluded from trials, at real-world scale and cost. Cost: treatment assignment is not randomized, so confounding (especially confounding by indication) is the central threat and must be addressed by design (active comparator, new-user restriction) and analysis (PS methods, negative controls), not assumed away.
When NOT to use — and when it is actively misleading or dangerous
- Very rare outcomes with expensive covariates. Following a huge cohort to capture a handful of events wastes measurement resources; a nested case-control or case-cohort design recovers nearly all the efficiency at a fraction of the cost. Forcing a full cohort here is not wrong, but it is the wrong tool.
- Ill-defined or person-varying time zero. If time zero is set at a point that itself depends on future events (e.g., starting follow-up at diagnosis but classifying exposure by a treatment received later), you manufacture immortal time bias: exposed person-time before the drug is dispensed is guaranteed event-free and spuriously favors the exposed. This is the single most common fatal error in RWD cohorts and is actively misleading — it can invert the sign of an effect.
- Prevalent-user (ever-exposed) cohorts. Starting follow-up among current users mixes people at different points in their treatment trajectory, induces depletion of susceptibles (survivors tolerate the drug and look healthier), and forces adjustment for variables on the causal pathway. Use a new-user cohort unless initiation is too rare, in which case consider the prevalent-new-user (Suissa) extension.
- Differential loss to follow-up by exposure. If the exposed and unexposed are censored for outcome-related reasons at different rates (informative censoring), naive estimates are biased; this demands explicit observation windows and, often, inverse-probability-of-censoring weighting.
- No way to control confounding by indication. A cohort comparing treated vs untreated for a condition that itself predicts the outcome, with no active comparator and no measured confounders, produces a confounded contrast that looks quantitative but is not interpretable as causal.
Data-source operational depth
- Claims (FFS or commercial): The natural substrate for incident-user cohorts. Exposure = the pharmacy claim (NDC + `fill_date` + `days_supply`); diagnoses come from medical claims (ICD-10-CM on professional/facility lines). Require continuous medical + pharmacy enrollment across the whole baseline/washout window so that "no prior fill" is a real observation, not unobserved person-time. Failure modes: (1) Medicare Advantage / capitated person-time lacks fee-for-service claims — utilization and fills are invisible, so absence of an event or a prior fill is missingness; restrict to enrollees with Parts A/B/D (or a commercial medical+pharmacy benefit) and drop MA-only spans. (2) Differential competing risks by exposure in elderly claims — death is a competing event that is often unobserved unless a death index or Part A inpatient-discharge-status is linked; if one arm is older/sicker, ignoring the competing risk overstates the cumulative incidence of the event of interest (use Fine-Gray or report cause-specific and cumulative-incidence estimates). (3) Immortal time in procedure/treatment studies — defining the exposed group by receipt of a procedure but starting the clock at an earlier landmark builds guaranteed survival into the exposed; align time zero to the procedure or use a landmark/time-varying treatment.
- EHR: Time zero is the order or administration, not a dispensing; problem lists, labs, and notes sharpen indication and baseline severity (an advantage over claims), but visit-driven capture means a patient who leaves the system disappears — define observation windows explicitly and treat loss to follow-up as potentially informative. Linkage to pharmacy fills is preferred to confirm the patient actually started.
- Registry: Strongest for indication, disease severity, and adjudicated/validated outcomes (e.g., cancer stage, cause of death); typically weak for complete medication exposure and for non-registry comorbidity. Link to claims for the full fill history and to a death index to firm up censoring.
- Linked claims–EHR–vital records: The ideal substrate (EHR severity + claims completeness + reliable mortality), but linkage introduces selection (only the linkable subset) and date-discrepancy issues (order vs fill vs service date) that must be reconciled before time zero is assigned.
Worked claims example
Question: 2-year cumulative incidence of hospitalized GI bleed among adults initiating a non-selective NSAID, in a commercial + Medicare FFS database
- Eligibility: age ≥18 and ≥365 days of continuous medical + pharmacy enrollment before the first NSAID fill (FFS-observable, no MA-only spans)
- Washout: no NSAID fill in the 365-day lookback — this makes the cohort incident users and removes prevalent-user bias
- Time zero: the date of that first qualifying fill (the `fill_date` of the index NDC)
(4) Baseline covariates: measured only in the 365 days up to and including time zero (prior GI bleed, anticoagulant/antiplatelet use, age, utilization), so no covariate is on the causal pathway (5) Follow-up and person-time: from time zero forward to the first inpatient claim with a primary GI-bleed diagnosis; censor at disenrollment, death (from a linked death index — a competing event, not an outcome), 2 years, or end of data
Do not count post-`days_supply` time as immortal "exposed" time — if this is an as-treated analysis, the on-treatment window is the stitched `days_supply` episodes plus a pre-specified grace period (6) Estimand: report the cumulative incidence function treating death as a competing risk (not 1 − KM), plus the incidence rate per 1,000 person-years; compare exposure groups with an active comparator (e.g., a different analgesic class) and PS adjustment rather than against never-users.
Decision diagram
flowchart TD Pop[Source population in the data] --> Elig[Eligibility + continuous, exposure-observable enrollment<br/>across the full baseline/washout window] Elig --> T0[Time zero = first exposure after drug-free washout<br/>assign exposure group at time zero] T0 --> Base[Baseline covariates measured ONLY up to and including time zero<br/>no post-time-zero information] Base --> Fup[Follow-up FORWARD from time zero<br/>count person-time at risk while observable] Fup --> Out[Incident outcome ascertainment<br/>identical rules across exposure groups] Fup --> Cen[Censor at disenrollment / death competing risk / end of data] Out --> Est[Estimand: incidence rate, cumulative incidence with competing risks, hazard ratio] Cen --> Est
gantt title Prospective cohort timeline for one incident user (claims) dateFormat YYYY-MM-DD axisFormat %b %Y section Baseline Continuous enrollment + washout (no prior study fill) :done, wash, 2023-01-01, 2023-12-31 section Time zero First study fill -> exposure group assigned :milestone, t0, 2024-01-01, 0d section Follow-up (forward) Person-time at risk (forward from time zero) :active, fu, 2024-01-01, 365d First incident outcome OR censor (disenroll/death/data end) :crit, cen, 2024-12-31, 1d
Worked example
Scenario
We want the 2-year (730-day) risk of a hospitalized GI bleed among adults who start taking an NSAID pain reliever. We have a tiny claims dataset of four patients, each with the pharmacy fill that marks their first NSAID. We set each person's day zero to that first fill, then follow every patient forward to see who is hospitalized for a GI bleed before two years are up. We then count the events and divide by the four people we started with.
Dataset
The raw rows an analyst would see: one starting fill per patient, plus what happened during forward follow-up.
| person_id | fill_date | drug | days_followed | outcome |
|---|---|---|---|---|
| 1001 | 2023-01-15 | ibuprofen | 180 | GI bleed hospitalization |
| 1002 | 2023-02-03 | ibuprofen | 730 | none |
| 1003 | 2023-02-20 | naproxen | 730 | none |
| 1004 | 2023-03-11 | naproxen | 400 | none (left plan, censored) |
Steps
Result
2-year cumulative incidence = 1 GI-bleed event / 4 patients enrolled at time zero = 0.25, or about 25%.
Trade-offs
Runnable example
Prospective (incident-user) cohort construction from claims-style inputs. Required inputs (already cleaned, de-duplicated): rx : pharmacy fills -> person_id, fill_date (datetime64), ndc, days_supply enroll : enrollment spans -> person_id, enroll_start, enroll_end (datetime64), ma_only (bool) # ma_only spans lack FFS...
import pandas as pd
WASHOUT_DAYS = 365 # drug-free + continuous-enrollment lookback that makes a user "incident"
STUDY_NDCS = {...} # set of NDCs defining the exposure of interest
def build_prospective_cohort(rx: pd.DataFrame, enroll: pd.DataFrame) -> pd.DataFrame:
rx = rx.sort_values(["person_id", "fill_date"])
study = rx[rx["ndc"].isin(STUDY_NDCS)]
# Candidate time zero = first fill of the study exposure for each person.
idx = (study.groupby("person_id", as_index=False)
.first()
.rename(columns={"fill_date": "index_date"})[["person_id", "index_date"]])
# New-user restriction: no study fill in the washout window strictly before time zero.
prior = study.merge(idx, on="person_id")
prior_in_washout = prior[(prior["fill_date"] < prior["index_date"]) &
(prior["fill_date"] >= prior["index_date"] - pd.Timedelta(days=WASHOUT_DAYS))]
idx = idx[~idx["person_id"].isin(prior_in_washout["person_id"])].copy()
# Continuous, FFS-observable enrollment spanning the full washout through time zero (no MA-only gaps).
e = enroll.merge(idx, on="person_id")
e["covers"] = ((e["enroll_start"] <= e["index_date"] - pd.Timedelta(days=WASHOUT_DAYS)) &
(e["enroll_end"] >= e["index_date"]) &
(~e["ma_only"]))
eligible = e.loc[e["covers"], "person_id"].unique()
cohort = idx[idx["person_id"].isin(eligible)].copy()
# Baseline covariate window: measure confounders strictly up to and including time zero (never after).
cohort["baseline_start"] = cohort["index_date"] - pd.Timedelta(days=WASHOUT_DAYS)
# Follow-up starts AT time zero -> no immortal time. Censoring (disenroll/death/end-of-data) added downstream.
cohort["followup_start"] = cohort["index_date"]
return cohort[["person_id", "index_date", "baseline_start", "followup_start"]]Prospective (incident-user) cohort construction with data.table. Inputs mirror the Python version: rx : person_id, fill_date (Date), ndc, days_supply enroll : person_id, enroll_start, enroll_end (Date), ma_only (logical) Returns one row per eligible new initiator with time zero and baseline-window bounds.
library(data.table)
WASHOUT_DAYS <- 365L
STUDY_NDCS <- c(...) # NDCs defining the exposure of interest
build_prospective_cohort <- function(rx, enroll) {
setDT(rx); setDT(enroll)
setorder(rx, person_id, fill_date)
study <- rx[ndc %chin% STUDY_NDCS]
# Candidate time zero = first study fill per person.
idx <- study[, .(index_date = fill_date[1L]), by = person_id]
# New-user restriction: drop anyone with a study fill in the washout window before time zero.
study <- merge(study, idx, by = "person_id")
prior_ids <- unique(study[fill_date < index_date &
fill_date >= index_date - WASHOUT_DAYS, person_id])
idx <- idx[!person_id %chin% prior_ids]
# Continuous, FFS-observable enrollment across the full washout through time zero (no MA-only spans).
e <- merge(enroll, idx, by = "person_id")
ok <- e[enroll_start <= index_date - WASHOUT_DAYS &
enroll_end >= index_date & !ma_only, unique(person_id)]
cohort <- idx[person_id %chin% ok]
cohort[, baseline_start := index_date - WASHOUT_DAYS] # covariate window ends at time zero
cohort[, followup_start := index_date] # follow-up begins at time zero -> no immortal time
cohort[, .(person_id, index_date, baseline_start, followup_start)]
}Prospective (incident-user) cohort construction in SAS using PROC SQL. Required input datasets (post data-management): work.rx : person_id, fill_date (SAS date), ndc, days_supply work.enroll : person_id, enroll_start, enroll_end (SAS dates), ma_only (0/1) Produces work.cohort: one row per eligible new initiator with...
%let washout = 365;
%let study_ndcs = '00000-0000-00'; /* quoted, comma-separated NDC list defining the exposure */
/* Candidate time zero = first study-exposure fill per person. */
proc sql;
create table idx as
select person_id, min(fill_date) as index_date format=date9.
from work.rx
where ndc in (&study_ndcs)
group by person_id;
quit;
/* New-user restriction: exclude any prior study fill inside the washout window before time zero. */
proc sql;
create table newuser as
select i.*
from idx i
where not exists (
select 1 from work.rx p
where p.person_id = i.person_id
and p.ndc in (&study_ndcs)
and p.fill_date < i.index_date
and p.fill_date >= i.index_date - &washout
);
quit;
/* Continuous, FFS-observable enrollment across the full washout through time zero (no MA-only spans). */
proc sql;
create table cohort as
select n.person_id,
n.index_date,
n.index_date - &washout as baseline_start format=date9., /* covariate window ends at time zero */
n.index_date as followup_start format=date9. /* follow-up begins at time zero */
from newuser n
where exists (
select 1 from work.enroll e
where e.person_id = n.person_id
and e.ma_only = 0
and e.enroll_start <= n.index_date - &washout
and e.enroll_end >= n.index_date
);
quit;Citations
- [1]Ray WA. Evaluating medication effects outside of clinical trials: new-user designs. American Journal of Epidemiology. 2003;158(9):915-920.
- [2]Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. American Journal of Epidemiology. 2016;183(8):758-764.
- [3]Suissa S, Dell'Aniello S. Time-related biases in pharmacoepidemiology. Pharmacoepidemiology and Drug Safety. 2020;29(9):1101-1110.