← Methods repository
CONCEPTADVANCEDPYTHON · R · SASlast reviewed 2026-08-24 · updated 2026-10-07 · 4 citations

ADaM for Real-World Data Submissions

The CDISC Analysis Data Model is the standardized structure for analysis datasets in an FDA submission; ADaM requires that every analysis value trace to the SDTM record it came from via ADSL and the define.xml PARAM/PARAMCD/PARAMTYP metadata, and that one variable per record carry the analysis-ready value, making the analytic dataset itself the audit-floor for a clinical- or real-world-data submission.

Data Standardcdiscadamsubmission-standardsregulatory-submissionfda-submissionrwe-submissiontraceabilityadtte
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

ADaM (Analysis Data Model) is the CDISC format for analysis-ready datasets in an FDA submission — the layer you actually run your statistical models on. Where SDTM holds raw observations normalized for review, ADaM holds one analysis value per row with explicit flags showing which records contribute to each analysis, where each value came from upstream, and how it was derived (a missing value filled in, an average over a window, etc.). Every value must trace back to an SDTM record through a documented chain so a reviewer can re-derive any result. For real-world data there is no randomization arm, so ADaM forces the sponsor to define and justify what counts as the treated and comparator group, the analysis populations, and the time zero.

When to use it
—Always build ADaM for FDA submission. Never analyze directly off SDTM.
—Use ADaM whenever the analysis will be reviewed (FDA, EMA, PMDA, HTA). For pure exploratory work, ad-hoc analytic datasets are fine but cannot be submitted.
—ADaM for FDA submissions and HTA evidence dossiers; OMOP for OHDSI network analyses, multi-database RWE, and rapid-cycle analytics. Many sponsors maintain parallel ETLs.
Watch out for
—ADaM is one more build layer over SDTM; if SDTM is missing or poorly mapped, ADaM's trace-back cannot be reconstructed.
—ADaM's strict traceability conventions add build time vs. a one-off analytic dataset. The discipline is worth it for any submission, exploratory or not.
—Parallel ETL pipelines (SDTM/ADaM and OMOP) require maintenance. Choose one as canonical source-of-truth and ETL to the other.

The Analysis Data Model (ADaM) is the CDISC standard for analysis-ready datasets submitted to the FDA. Where SDTM is the normalized, one-row-per-observation tabulation floor, ADaM is the one-row-per-subject-per-parameter (or per-record) analysis floor. The defining rule of ADaM is traceability — every analysis value in an ADaM dataset must trace back to the SDTM record (or upstream data source) that produced it, via the ADSL subject-level analysis dataset and the define.xml metadata describing how each variable was derived.

Why it matters for RWE

In a randomized trial, analysis populations and treatment groups are defined by the protocol. In RWE there is no randomization arm: ADaM forces the sponsor to make explicit — in ADSL population flags (ITTFL, PPSFL), ARM/ACTARM values, and time-zero (TRTSDT) derivations — exactly who counts as treated vs comparator, when follow-up starts, and how intercurrent events (discontinuation, switching, death) are handled. This is where the target-trial emulation's design decisions become visible to the reviewer: a grace period implemented inconsistently between SDTM EX records and the ADaM TRTSDT derivation is precisely the kind of defect define.xml and the ADRG exist to expose.

Core ADaM structures for RWE packages

  • ADSL (Subject-Level Analysis Dataset): one row per subject. Carries demographics, planned/actual treatment, study start/stop dates, population flags, and the denominators every other dataset checks against. In RWE builds, ADSL is where cohort membership, eligibility-window qualification, index-date definitions, and linked-data provenance flags live.
  • BDS (Basic Data Structure): one row per subject per parameter per timepoint — ADTTE (time-to-event with AVAL/CNSR/SRCSEQ), ADLB (labs from EHR LB domains), ADVS (vitals), ADEFF (efficacy). PARAM/PARAMCD identify the parameter; AVISIT/AVAL carry the analysis value; SRCSEQ/SRCVAR point back to SDTM.
  • OCCDS (Occurrence Data Structure): one row per occurrence event — adverse events (ADAE built from SDTM AE, which in RWE came from problem lists or incident diagnoses), concomitant medications summarized, medical history.

RWE-specific derivation discipline

  • Time zero and follow-up windows: TRTSDT must derive deterministically from SDTM EX (dispensing/administration) records under a pre-specified rule (first dispensing within eligibility window; grace period handling documented). Any imputation is a named algorithm in define.xml.
  • Censoring: ADTTE CNSRGx reasons must map to documented SDTM DS/censoring sources — administrative end of enrollment, plan switch, death — not to analyst judgment at run time.
  • Endpoint derivations from claims/EHR noise: deduplication windows, lookback windows for baseline comorbidity, and outcome-algorithm definitions belong in ADaM metadata (with the source algorithm in the ADRG), so the same endpoint is reproducible across data sources.
  • Cost and utilization analyses: claims payments become custom BDS parameters or ADSL cost fields with explicit currency-year adjustment and inflation rules documented.

Common pitfalls in RWE-to-ADaM conversion

  • Silent re-derivation between data cuts. If ADaM was built against one SDTM snapshot and tables are regenerated after a refresh, SRCSEQ pointers break. Lock the SDTM snapshot per submission package.
  • Population-flag drift. ITTFL defined one way in ADSL but filtered differently in a TLF macro is the classic review finding. One definition, one place, referenced everywhere.
  • Traceability gaps for custom variables. Any ADSL variable without an origin and algorithm in define.xml is undocumented data — reviewers flag it, and for RWD-derived variables (linkage flags, index-date rules) it is exactly the material they need to see.
  • Mixing layers. Endpoint logic embedded in SDTM mapping (where it constrains all downstream analyses) or buried only in TLF macros (where no reviewer can see it) both fail; the correct home is ADaM derivations with define.xml documentation.

Pros, cons, and trade-offs

  • vs analyzing straight from SDTM: SDTM holds tabulated observations; endpoint derivations done ad hoc in analysis code are invisible to the reviewer. ADaM makes every derivation explicit, versioned metadata.
  • vs sponsor-internal analysis datasets: internal datasets are faster but cannot be submitted; conforming ADaM costs discipline up front and buys reviewer reconstructability.
  • Trade-off: ADSL simplicity vs BDS/Occurrence structural complexity — choose per parameter type, not uniformly; over-normalizing simple subject-level facts into BDS multiplies rows without analytic gain.

When NOT to use

Not for descriptive or exploratory work product that will never be submitted — the traceability overhead buys nothing there. And never let ADaM absorb mappings that belong upstream in SDTM (source-code translation, unit harmonization); misplacing logic between layers breaks both.

Decision diagram

flowchart LR
  S[SDTM domains DM EX DS AE MH] --> A[ADSL one row per subject]
  S --> T[ADTTE one row per PARAMCD]
  S --> E[AEFF ADAE ADLB ADVS]
  A --> X[Define.xml v2.x + ADRG]
  T --> X
  E --> X
  X --> P[Pinnacle 21 conformance]
  P --> F[FDA submission XPT v5 package]
SDTM to ADaM to define.xml chain. ADaM AVALs trace back to SDTM records via SRCSEQ, and every derivation is documented in define.xml so the submission is reviewer-reproducible.

Worked example

Scenario

We need to build an ADTTE dataset for a real-world comparative effectiveness study of Drug A vs Drug B with outcome of time-to-first hospitalization. Each patient has an index date (treatment initiation from SDTM EX), follow-up until hospitalization (from SDTM DS plus HO) or censoring at end-of-enrollment or end-of-study. We build one row per analysis parameter (PARAMCD equals TTFHOSP), with AVAL equal to time in days from index to event or censor, CNSR the censor indicator, and SRCSEQ pointing back to the originating SDTM record.

Dataset

ADSL (subject-level) and ADTTE (time-to-event) for two subjects in the RWE cohort.

datasetUSUBJIDARMACTARMITTFLPARAMCDPARAMAVALCNSRSRCSEQ
ADSL - RWED-001 - DRUG A - DRUG A - Y - - - - -
ADSL - RWED-002 - DRUG B - DRUG A - Y - - - - -
ADTTE - RWED-001 - - - - TTFHOSP - Time to First Hospitalization (days) - 184 - 0 - 12
ADTTE - RWED-002 - - - - TTFHOSP - Time to First Hospitalization (days) - 365 - 1 - 7

Steps

1Build ADSL first — one row per USUBJID with ARM from the protocol-defined exposure, ACTARM from the actual observed first exposure, ITTFL set to Y for the pre-specified analysis population.
2Build ADTTE from SDTM — for each subject compute AVAL as (event or censor date) minus index_date; CNSR is 1 if censored (end-of-enrollment or end-of-study), 0 if event (hospitalization); SRCSEQ references the SDTM DS record whose DSDECOD matches HOSPITALIZATION.
3Lock the SDTMIG and ADaMIG versions and the WHODrug / MedDRA versions in define.xml; document partial-date imputation rules and population-definition rules in the ADRG.

Result

ADSL has two subjects with their arms and ITTFL set; ADTTE has the same two subjects with PARAMCD equal to TTFHOSP, AVAL equal to 184 (event) and 365 (censor), CNSR equal to 0 and 1, and SRCSEQ pointing back to the SDTM source records. A reviewer can reproduce the analysis from ADSL + ADTTE + define.xml + the ADRG.

Trade-offs

vs. SDTM
Pros of this
—ADaM is what the actual analysis runs on. Variables are analysis-ready (AVAL, CNSR, analysis flags), populated per parameter, with trace-back to SDTM. Define.xml algorithms document every derivation.
vs. Ad hoc analysis datasets
Pros of this
—ADaM enables reviewer replication — a reviewer can run your analysis from ADSL, the ADaM datasets, define.xml, and the ADRG without seeing your analysis code. This is what FDA's reproducibility standard actually requires.
vs. OMOP CDM analysis datasets
Pros of this
—ADaM is the FDA-mandated analysis format; OMOP is OHDSI's analysis format for distributed-network studies.

Runnable example

Toy ADTTE builder from claims-style inputs. Produces ADSL (one row per subject) and ADTTE (one row per PARAMCD=TTFHOSP) with AVAL, CNSR, SRCSEQ.

requires: pandas
import pandas as pd

ADAMIG_VERSION = "1.3"        # lock per submission
PARAMCD        = "TTFHOSP"

def build_adsl(ex, enroll):
    # In RWE ARM = protocol-defined intended exposure; ACTARM = actual observed first exposure.
    adsl = ex.groupby("person_id").agg(
        USUBJID=("person_id", lambda s: f"RWED-{s.iloc[0]:03d}"),
    ).reset_index(drop=True)
    adsl["ARM"]    = "DRUG A"
    adsl["ACTARM"] = "DRUG A"
    adsl["ITTFL"]  = "Y"  # operational rule: subject has >=1 day of follow-up
    return adsl[["USUBJID", "ARM", "ACTARM", "ITTFL"]]

def build_adtte(adsl, ex, claims, enroll):
    first_ex = ex.groupby("person_id")["ex_std_dc"].min().rename("index_date")
    first_hosp = (claims[claims["claim_type"] == "HOSP"]
                  .groupby("person_id")["service_date"].min().rename("event_date"))
    end_enr = enroll.set_index("person_id")["end_enroll_date"]
    end_study = pd.Timestamp("2026-08-24")  # LOCK per submission

    base = (first_ex.to_frame()
            .join(first_hosp, how="left")
            .join(end_enr.rename("end_enroll_date")))
    base["end_date"] = base["event_date"].fillna(base["end_enroll_date"]).fillna(end_study)
    base["CNSR"]   = base["event_date"].isna().astype(int)        # 1=censored, 0=event
    base["SRCSEQ"] = base["event_date"].notna().astype(int)       # toy: 1 if event, 0 if censor
    base["AVAL"]   = (base["end_date"] - base["index_date"]).dt.days
    base = base.reset_index().merge(adsl, left_on="person_id", right_on="USUBJID")
    base["PARAMCD"]  = PARAMCD
    base["PARAM"]    = "Time to First Hospitalization (days)"
    base["PARAMTYP"] = "DERIVED"
    return base[["USUBJID", "PARAMCD", "PARAM", "PARAMTYP", "AVAL", "CNSR", "SRCSEQ"]]

Citations

FOUNDATIONAL / METHODS
  1. [1]CDISC. Analysis Data Model (ADaM). CDISC foundational standard.
  2. [2]U.S. Food and Drug Administration. Providing Regulatory Submissions in Electronic Format — Standardized Study Data. FDA Guidance for Industry.
  3. [3]CDISC. ADaM Basic Data Structure (BDS) for Time-to-Event (TTE) Analyses v1.0. CDISC.
APPLIED EXAMPLES
  1. [4]Rizzoli S, Ori A, Mignani A, Ferri F, Simoni L. Implementation of Clinical Data Interchange Standard Consortium (CDISC) standards to Real-World Data: Challenges and Strategies in SDTM Development in the Setting of Observational Studies. Journal of the Society for Clinical Data Management. 2026;6(1).