ADaM for Real-World Data Submissions
The CDISC Analysis Data Model is the standardized structure for analysis datasets in an FDA submission; ADaM requires that every analysis value trace to the SDTM record it came from via ADSL and the define.xml PARAM/PARAMCD/PARAMTYP metadata, and that one variable per record carry the analysis-ready value, making the analytic dataset itself the audit-floor for a clinical- or real-world-data submission.
On this page
ADaM (Analysis Data Model) is the CDISC format for analysis-ready datasets in an FDA submission — the layer you actually run your statistical models on. Where SDTM holds raw observations normalized for review, ADaM holds one analysis value per row with explicit flags showing which records contribute to each analysis, where each value came from upstream, and how it was derived (a missing value filled in, an average over a window, etc.). Every value must trace back to an SDTM record through a documented chain so a reviewer can re-derive any result. For real-world data there is no randomization arm, so ADaM forces the sponsor to define and justify what counts as the treated and comparator group, the analysis populations, and the time zero.
The Analysis Data Model (ADaM) is the CDISC standard for analysis-ready datasets submitted to the FDA. Where SDTM is the normalized, one-row-per-observation tabulation floor, ADaM is the one-row-per-subject-per-parameter (or per-record) analysis floor. The defining rule of ADaM is traceability — every analysis value in an ADaM dataset must trace back to the SDTM record (or upstream data source) that produced it, via the ADSL subject-level analysis dataset and the define.xml metadata describing how each variable was derived.
Why it matters for RWE
In a randomized trial, analysis populations and treatment groups are defined by the protocol. In RWE there is no randomization arm: ADaM forces the sponsor to make explicit — in ADSL population flags (ITTFL, PPSFL), ARM/ACTARM values, and time-zero (TRTSDT) derivations — exactly who counts as treated vs comparator, when follow-up starts, and how intercurrent events (discontinuation, switching, death) are handled. This is where the target-trial emulation's design decisions become visible to the reviewer: a grace period implemented inconsistently between SDTM EX records and the ADaM TRTSDT derivation is precisely the kind of defect define.xml and the ADRG exist to expose.
Core ADaM structures for RWE packages
- ADSL (Subject-Level Analysis Dataset): one row per subject. Carries demographics, planned/actual treatment, study start/stop dates, population flags, and the denominators every other dataset checks against. In RWE builds, ADSL is where cohort membership, eligibility-window qualification, index-date definitions, and linked-data provenance flags live.
- BDS (Basic Data Structure): one row per subject per parameter per timepoint — ADTTE (time-to-event with AVAL/CNSR/SRCSEQ), ADLB (labs from EHR LB domains), ADVS (vitals), ADEFF (efficacy). PARAM/PARAMCD identify the parameter; AVISIT/AVAL carry the analysis value; SRCSEQ/SRCVAR point back to SDTM.
- OCCDS (Occurrence Data Structure): one row per occurrence event — adverse events (ADAE built from SDTM AE, which in RWE came from problem lists or incident diagnoses), concomitant medications summarized, medical history.
RWE-specific derivation discipline
- Time zero and follow-up windows: TRTSDT must derive deterministically from SDTM EX (dispensing/administration) records under a pre-specified rule (first dispensing within eligibility window; grace period handling documented). Any imputation is a named algorithm in define.xml.
- Censoring: ADTTE CNSRGx reasons must map to documented SDTM DS/censoring sources — administrative end of enrollment, plan switch, death — not to analyst judgment at run time.
- Endpoint derivations from claims/EHR noise: deduplication windows, lookback windows for baseline comorbidity, and outcome-algorithm definitions belong in ADaM metadata (with the source algorithm in the ADRG), so the same endpoint is reproducible across data sources.
- Cost and utilization analyses: claims payments become custom BDS parameters or ADSL cost fields with explicit currency-year adjustment and inflation rules documented.
Common pitfalls in RWE-to-ADaM conversion
- Silent re-derivation between data cuts. If ADaM was built against one SDTM snapshot and tables are regenerated after a refresh, SRCSEQ pointers break. Lock the SDTM snapshot per submission package.
- Population-flag drift. ITTFL defined one way in ADSL but filtered differently in a TLF macro is the classic review finding. One definition, one place, referenced everywhere.
- Traceability gaps for custom variables. Any ADSL variable without an origin and algorithm in define.xml is undocumented data — reviewers flag it, and for RWD-derived variables (linkage flags, index-date rules) it is exactly the material they need to see.
- Mixing layers. Endpoint logic embedded in SDTM mapping (where it constrains all downstream analyses) or buried only in TLF macros (where no reviewer can see it) both fail; the correct home is ADaM derivations with define.xml documentation.
Pros, cons, and trade-offs
- vs analyzing straight from SDTM: SDTM holds tabulated observations; endpoint derivations done ad hoc in analysis code are invisible to the reviewer. ADaM makes every derivation explicit, versioned metadata.
- vs sponsor-internal analysis datasets: internal datasets are faster but cannot be submitted; conforming ADaM costs discipline up front and buys reviewer reconstructability.
- Trade-off: ADSL simplicity vs BDS/Occurrence structural complexity — choose per parameter type, not uniformly; over-normalizing simple subject-level facts into BDS multiplies rows without analytic gain.
When NOT to use
Not for descriptive or exploratory work product that will never be submitted — the traceability overhead buys nothing there. And never let ADaM absorb mappings that belong upstream in SDTM (source-code translation, unit harmonization); misplacing logic between layers breaks both.
Decision diagram
flowchart LR S[SDTM domains DM EX DS AE MH] --> A[ADSL one row per subject] S --> T[ADTTE one row per PARAMCD] S --> E[AEFF ADAE ADLB ADVS] A --> X[Define.xml v2.x + ADRG] T --> X E --> X X --> P[Pinnacle 21 conformance] P --> F[FDA submission XPT v5 package]
Worked example
Scenario
We need to build an ADTTE dataset for a real-world comparative effectiveness study of Drug A vs Drug B with outcome of time-to-first hospitalization. Each patient has an index date (treatment initiation from SDTM EX), follow-up until hospitalization (from SDTM DS plus HO) or censoring at end-of-enrollment or end-of-study. We build one row per analysis parameter (PARAMCD equals TTFHOSP), with AVAL equal to time in days from index to event or censor, CNSR the censor indicator, and SRCSEQ pointing back to the originating SDTM record.
Dataset
ADSL (subject-level) and ADTTE (time-to-event) for two subjects in the RWE cohort.
| dataset | USUBJID | ARM | ACTARM | ITTFL | PARAMCD | PARAM | AVAL | CNSR | SRCSEQ |
|---|---|---|---|---|---|---|---|---|---|
| ADSL - RWED-001 - DRUG A - DRUG A - Y - - - - - | |||||||||
| ADSL - RWED-002 - DRUG B - DRUG A - Y - - - - - | |||||||||
| ADTTE - RWED-001 - - - - TTFHOSP - Time to First Hospitalization (days) - 184 - 0 - 12 | |||||||||
| ADTTE - RWED-002 - - - - TTFHOSP - Time to First Hospitalization (days) - 365 - 1 - 7 |
Steps
Result
ADSL has two subjects with their arms and ITTFL set; ADTTE has the same two subjects with PARAMCD equal to TTFHOSP, AVAL equal to 184 (event) and 365 (censor), CNSR equal to 0 and 1, and SRCSEQ pointing back to the SDTM source records. A reviewer can reproduce the analysis from ADSL + ADTTE + define.xml + the ADRG.
Trade-offs
Runnable example
Toy ADTTE builder from claims-style inputs. Produces ADSL (one row per subject) and ADTTE (one row per PARAMCD=TTFHOSP) with AVAL, CNSR, SRCSEQ.
import pandas as pd
ADAMIG_VERSION = "1.3" # lock per submission
PARAMCD = "TTFHOSP"
def build_adsl(ex, enroll):
# In RWE ARM = protocol-defined intended exposure; ACTARM = actual observed first exposure.
adsl = ex.groupby("person_id").agg(
USUBJID=("person_id", lambda s: f"RWED-{s.iloc[0]:03d}"),
).reset_index(drop=True)
adsl["ARM"] = "DRUG A"
adsl["ACTARM"] = "DRUG A"
adsl["ITTFL"] = "Y" # operational rule: subject has >=1 day of follow-up
return adsl[["USUBJID", "ARM", "ACTARM", "ITTFL"]]
def build_adtte(adsl, ex, claims, enroll):
first_ex = ex.groupby("person_id")["ex_std_dc"].min().rename("index_date")
first_hosp = (claims[claims["claim_type"] == "HOSP"]
.groupby("person_id")["service_date"].min().rename("event_date"))
end_enr = enroll.set_index("person_id")["end_enroll_date"]
end_study = pd.Timestamp("2026-08-24") # LOCK per submission
base = (first_ex.to_frame()
.join(first_hosp, how="left")
.join(end_enr.rename("end_enroll_date")))
base["end_date"] = base["event_date"].fillna(base["end_enroll_date"]).fillna(end_study)
base["CNSR"] = base["event_date"].isna().astype(int) # 1=censored, 0=event
base["SRCSEQ"] = base["event_date"].notna().astype(int) # toy: 1 if event, 0 if censor
base["AVAL"] = (base["end_date"] - base["index_date"]).dt.days
base = base.reset_index().merge(adsl, left_on="person_id", right_on="USUBJID")
base["PARAMCD"] = PARAMCD
base["PARAM"] = "Time to First Hospitalization (days)"
base["PARAMTYP"] = "DERIVED"
return base[["USUBJID", "PARAMCD", "PARAM", "PARAMTYP", "AVAL", "CNSR", "SRCSEQ"]]
R and data.table version. Builds ADSL, then ADTTE mirroring the Python logic.
library(data.table)
PARAMCD <- "TTFHOSP"
END_STUDY <- as.Date("2026-08-24") # lock per submission
build_adsl <- function(ex, enroll) {
setDT(ex); setDT(enroll)
first_ex <- ex[, .(index_date = min(ex_std_dc)), by = person_id]
adsl <- first_ex[, .(USUBJID = sprintf("RWED-%03d", person_id),
ARM = "DRUG A",
ACTARM = "DRUG A",
ITTFL = "Y")]
adsl[]
}
build_adtte <- function(adsl, ex, claims, enroll) {
setDT(claims)
first_ex <- ex[, .(index_date = min(ex_std_dc)), by = person_id]
first_hosp <- claims[claim_type == "HOSP", .(event_date = min(service_date)), by = person_id]
end_enr <- enroll[, .(person_id, end_enroll_date)]
base <- merge(first_ex, first_hosp, by = "person_id", all.x = TRUE)
base <- merge(base, end_enr, by = "person_id", all.x = TRUE)
base[is.na(event_date) & is.na(end_enroll_date), end_date := END_STUDY]
base[!is.na(event_date), end_date := event_date]
base[is.na(event_date) & !is.na(end_enroll_date), end_date := end_enroll_date]
base[, CNSR := as.integer(is.na(event_date))]
base[, SRCSEQ := as.integer(!is.na(event_date))]
base[, AVAL := as.integer(end_date - index_date)]
base <- merge(base, adsl, by.x = "person_id", by.y = "USUBJID", all.x = TRUE)
base[, `:=`(PARAMCD = PARAMCD,
PARAM = "Time to First Hospitalization (days)",
PARAMTYP = "DERIVED")]
base[, .(USUBJID, PARAMCD, PARAM, PARAMTYP, AVAL, CNSR, SRCSEQ)]
}
SAS PROC SQL build. Produces ADTTE from claims-style inputs (work.claims, work.ex, work.enroll).
/* Constants: lock per submission package. */
%let paramcd = TTFHOSP;
%let end_study = '24AUG2026'd;
/* First exposure per subject = index date. */
proc sql;
create table work.first_ex as
select person_id, min(ex_std_dc) as index_date format=yymmdd10.
from work.ex
group by person_id;
quit;
/* First hospitalization per subject = event date (may be missing). */
proc sql;
create table work.first_hosp as
select person_id, min(service_date) as event_date format=yymmdd10.
from work.claims
where claim_type = 'HOSP'
group by person_id;
quit;
/* Build ADTTE — AVAL = end_date - index_date; CNSR = 1 if censored. */
proc sql;
create table work.adtte as
select
cats("RWED-", put(e.person_id, z3.)) as USUBJID length=10,
"¶mcd" as PARAMCD length=8,
"Time to First Hospitalization (days)" as PARAM length=40,
"DERIVED" as PARAMTYP length=8,
(coalesce(h.event_date, n.end_enroll_date, &end_study)
- e.index_date) as AVAL,
case when h.event_date is not null then 0 else 1 end as CNSR,
case when h.event_date is not null then 1 else 0 end as SRCSEQ
from work.first_ex e
left join work.first_hosp h on e.person_id = h.person_id
left join work.enroll n on e.person_id = n.person_id;
quit;
Citations
- [1]CDISC. Analysis Data Model (ADaM). CDISC foundational standard.
- [2]U.S. Food and Drug Administration. Providing Regulatory Submissions in Electronic Format — Standardized Study Data. FDA Guidance for Industry.
- [3]CDISC. ADaM Basic Data Structure (BDS) for Time-to-Event (TTE) Analyses v1.0. CDISC.