← Methods repository
CONCEPTADVANCEDPYTHON · R · SASlast reviewed 2026-08-24 · updated 2026-08-25 · 4 citations

SDTM for Real-World Data Submissions

The CDISC Study Data Tabulation Model is the standardized structure and metadata format the FDA requires for study data packages; applying it to real-world data forces explicit modeling decisions about how to map claims dispensings, EHR encounters, and registry records into SDTM domains (DM, EX, AE, LB, etc.) and how to retain traceability from each tabulated value to the source record.

Data Standardcdiscsdtmsubmission-standardsregulatory-submissionfda-submissionrwe-submissiontraceabilityclaims-to-sdtm
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

SDTM (Study Data Tabulation Model) is the standardized, row-by-row data format the FDA requires for study data submissions. Each piece of information gets placed in a labeled "domain" — demographics here, exposure there, adverse events in their own slot — along with the date, a coded value, and a controlled vocabulary term so a reviewer can read the file and reconstruct what happened. Real- world data rarely starts in this shape: claims are payment records, EHR data is visit notes, and registries are disease-specific. Mapping them to SDTM is real modeling work — every code translation and every date assumption has to be documented in the define.xml metadata file and the Analysis Data Reviewer's Guide so an FDA reviewer can trace any submitted value back to the original record.

When to use it
—Always build SDTM for FDA submissions. Use ADaM for the actual analyses; ADaM records trace to SDTM via the ADSL and the define.xml PARAM/PARAMCD logic.
—Use SDTM for FDA submissions; use OMOP for OHDSI network analyses, multi-database RWE, and rapid-cycle analytics. Many sponsors maintain parallel ETLs from a common source-of-truth raw layer.
—Always for FDA submissions; rarely for exploratory analysis.
Watch out for
—SDTM is normalized and verbose; analytic datasets built directly from SDTM are inefficient and rarely used. ADaM datasets are the analytic substrate.
—Maintaining both is real work. ETL pipelines must be defensible in both directions.
—The SDTM build itself can be lossy if mappings are not carefully designed; preserve source values in SUPPQUAL or custom --TORG domains.

The Study Data Tabulation Model (SDTM) is the standardized structure for submitting study data to the FDA. Each observation is placed in a domain (a logical grouping such as DM for demographics, EX for exposure, AE for adverse events, LB for laboratory, CM for concomitant medications) along with identifier variables, timing variables, and a controlled-terminology-coded value. The format is intentionally normalized — one row per observation — so a reviewer can reconstruct the analytic dataset from the tabulations and the define.xml metadata.

Why it matters for RWE

SDTM was designed for clinical-trial data, where a CRF captures exactly the variables of interest on a fixed schedule. RWE sources look very different: claims are encounter- and payment-driven, EHR data is visit- and problem-list-driven, registries are disease- and assessment-driven. Mapping RWD into SDTM is therefore a modeling exercise, not a data dump. Every mapping decision — which NDC rolls up to which generic name, which ICD-10-CM code maps to which MedDRA preferred term for an AE, how to derive EX dates from dispensed-supply intervals — must be documented in the Analysis Data Reviewer's Guide (ADRG) and the define.xml so a reviewer can reconstruct it. This visibility is exactly what the FDA Data Standards Catalog and the Study Data Technical Conformance Guide require.

Core domain mapping in RWD

  • DM (Demographics): build from enrollment/insurance eligibility; populate ARM, ACTARM, SITEID, AGE, SEX, RACE, ETHNIC with the values available. In RWD there is no randomization arm — populate the planned and actual arms with the study-defined exposure (e.g., treatment vs comparator).
  • EX (Exposure): claims dispensings or EHR administrations become EX records, with EXDOSE, EXDOSU, EXDOSFRM, EXROUTE, EXSTDTC, EXENDTC from NDC + days supply (claims) or mar/admin records (EHR). Grace periods and carryover are operationalized before SDTM build, not after.
  • AE (Adverse Events): RWD usually does not have AEs in the CRF sense. Map EHR problem lists, diagnosis codes flagged as "incident" during follow-up, or spontaneous-report-system records to AE; translate source codes to MedDRA preferred terms (PT) using the MedDRA MSSO browser and document the version in define.xml.
  • CM (Concomitant Medications): claims dispensings for non-study medications, mapped to WHODrug B3 / C3 format with CMTRT, CMCAT, CMSTDTC, CMENDTC, and route.
  • DS (Disposition): derive from enrollment spans, plan switches, death records, and end-of- follow-up rules; populate DSCAT, DSDECOD, DSTERM, DSSTDTC.
  • LB (Laboratory): EHR only; LOINC-code the test, populate LBORRES, LBORRESU, LBNRIND, LBBLFL, LBTESTCD, LBTEST, LBCAT, LBFAST, LBSPEC, LBMETHOD.
  • MH (Medical History): baseline-window diagnoses from claims or EHR, with MHCAT and the appropriate MedDRA / ICD coding.

Common pitfalls in RWD-to-SDTM conversion

  • Date imputation and partial dates. SDTM ISO 8601 timestamps require precision. RWD often has only month or year. Apply pre-specified partial-date imputation rules (e.g., first-of-month if only month is known, midpoint if only year), document the rule in define.xml, and never silently leave blanks.
  • Coding version drift. MedDRA, WHODrug, LOINC, and ICD versions are time-stamped. The version used at SDTM build must be locked in the metadata and any re-coding (e.g., annual MedDRA re-coding) is a documented version change.
  • Lossy mappings. When many source codes collapse to one SDTM value, the audit chain breaks unless supplemental qualifiers (SUPP--.QVAL) or a custom domain (with --TORG prefix) preserve the source value. The FDA Technical Conformance Guide calls these out as needed for traceability.
  • Race and ethnicity handling. Race is federally standardized (OMB categories) but RWD often has missing or imputed values; document the source, the imputation (if any), and never aggregate away race categories without a stated rationale.
  • Unit standardization. Claims have no units. EHR LBs must be harmonized to UCUM (Unified Code for Units of Measure); document any unit conversions in the ADRG.

Operational depth for common RWD sources

  • Claims (MarketScan, Optum, Truveta, IQVIA, SNDS): rich dispensings (EX, CM), diagnoses (AE, MH when coded), but no labs (LB) and no vital signs (VS). Payments become custom cost domains or are absorbed into ADSL cost fields. Date precision is usually exact.
  • EHR (Epic, Cerner, HealthVerity): rich LB, VS, AE (when problem-list updated), EX (when administered), but encounter-level only. Date precision is exact but visit-driven; out-of-system care is invisible.
  • Linked claims–EHR: best of both, with explicit record-linkage provenance in a custom SUPPQUAL.
  • Registries (SEER-Medicare, CPRD, disease registries): disease-specific variables populate custom domains with --TORG prefixes; linked to claims via the linkage algorithm documented in the ADRG.

FDA submission posture

The FDA Data Standards Catalog lists the SDTM and ADaM versions required for the submission period; the Study Data Technical Conformance Guide specifies file formats (XPT v5), define.xml v2.x structure, the ADRG and SDRG expectations, and the conformance validator rules. A passing OpenCDISC/Pinnacle 21 validation is a prerequisite for FDA acceptance of the submission package; integration of the validator report is part of the ADRG. Pros, cons, and trade-offs.

  • vs leaving RWD in source format: SDTM imposes structure the FDA can review programmatically, at the cost of a genuine mapping exercise; unmapped source data is not acceptable for regulated submissions regardless.
  • vs mapping directly to ADaM: skipping SDTM collapses traceability — reviewers cannot reconstruct analysis values from tabulations, which is precisely what the Technical Conformance Guide exists to enforce.
  • Trade-off: fidelity vs normalization. Every lossy collapse (many NDCs to one generic name) must be preserved in SUPPQUAL or a custom domain or the audit chain breaks.

When NOT to use

Do not force RWD into SDTM for journal publications or internal exploratory analyses — the cost buys regulatory reviewability, not analytic insight; use analysis-ready source structures instead. And do not treat SDTM as an analytic layer: it is tabulation, not analysis — deriving endpoints belongs in ADaM.

Decision diagram

flowchart LR
  R[Raw RWD<br/>claims / EHR / registry] --> C[Curate & reconcile<br/>linkage, dedup, dates]
  C --> M[Map to SDTM domains<br/>DM, EX, AE, CM, MH, DS, LB...]
  M --> CT[Apply controlled terminology<br/>WHODrug, MedDRA, LOINC, UCUM]
  CT --> V[Pinnacle 21 conformance<br/>validator pass]
  V --> X{XPT v5<br/>package}
  X --> D[define.xml v2.x<br/>+ ADRG + SDRG]
  D --> F[FDA submission<br/>package]
RWD-to-SDTM pipeline — curate, map, code, validate, then wrap with define.xml and the reviewer guides before the FDA submission.

Worked example

Scenario

We need to convert a single patient's claims-derived dispensings into SDTM EX records so they can be included in an FDA submission. The patient is 58 years old, commercially insured, with three dispensings of 30-count 10mg atorvastatin over 90 days. We build one EX record per dispensing with the dose, unit, form, route, start, and end date populated, and one CM record for a concurrent 30-day amoxicillin course to show the convention.

Dataset

One patient's dispensings mapped into SDTM EX and CM records.

domainUSUBJIDEXSEQ/CMSEQEXTRT/CMTRTEXDOSE/CMDOSEEXDOSU/CMDOSUEXDOSFRM/CMDOSFRMEXROUTE/CMROUTEEXSTDTC/CMSTDTCEXENDTC/CMENDTC
EXRWED-0011ATORVASTATIN10mgTABLETORAL2025-01-152025-02-13
EXRWED-0012ATORVASTATIN10mgTABLETORAL2025-02-142025-03-14
EXRWED-0013ATORVASTATIN10mgTABLETORAL2025-03-152025-04-13
CMRWED-0011AMOXICILLIN500mgCAPSULEORAL2025-02-202025-03-21

Steps

1Identify the subject (RWED-001) and confirm subject-level DM is built first (age, sex, ARM, ACTARM).
2For each dispensing, emit one EX record with EXTRT from the generic name (not the NDC), EXDOSE in mg, EXDOSU from UCUM (mg), EXDOSFRM and EXROUTE populated, EXSTDTC = dispense date, EXENDTC = dispense date + days_supply - 1.
3For non-study dispensings (amoxicillin) emit CM records with CMCAT = "CONCOMITANT MEDICATION" and the source code-to-WHODrug mapping version documented in define.xml.
4Record the SDTM version (e.g., SDTMIG v3.4) and the controlled-terminology package version in define.xml so a reviewer can re-create the file exactly.

Result

Three EX records (ATORVASTATIN 10 mg TABLET ORAL, 30-day supplies) and one CM record (AMOXICILLIN 500 mg CAPSULE ORAL) under subject RWED-001, with all dates ISO 8601, all values controlled-terminology coded, and the SDTMIG / WHODrug / MedDRA versions locked in the ADRG and define.xml.

Trade-offs

vs. ADaM analysis datasets
Pros of this
—SDTM is the audit-floor for an FDA submission — every value in ADaM must trace back to an SDTM record. Building SDTM rigorously forces the trace from raw source to analytic estimate.
vs. OMOP CDM
Pros of this
—SDTM is what FDA accepts; OMOP CDM is the OHDSI network's analysis substrate. SDTM is regulatory-mandated; OMOP is research-oriented and powers distributed network studies.
vs. Raw or minimally transformed source datasets
Pros of this
—Standardized domains and controlled terminology make reviews faster and let the FDA's conformance validator catch structural issues automatically.

Runnable example

Toy SDTM EX builder from claims-style long-form input. Required inputs: claims : subject_id, ndc, generic_name, rx_date (date), days_supply, dose_mg, form, route dm : subject_id, armcd, age, sex, race Returns EX and DM dataframes with ISO-8601 dates, controlled-terminology upper-case strings, and a SUPP--EX data...

requires: pandas
import pandas as pd

SDTMIG_VERSION   = "3.4"        # update per submission lock
WHODRUG_VERSION  = "C3-2024"    # record in define.xml
CONTROLLED_TERM  = "2024-12-19" # CDISC controlled-terminology package version

def to_iso(d: pd.Timestamp) -> str:
    # SDTM uses ISO 8601 dates; partial dates use SDTM imputation rules (documented in define.xml).
    return d.strftime("%Y-%m-%d")

def build_sdtm_ex(claims: pd.DataFrame, dm: pd.DataFrame) -> tuple[pd.DataFrame, pd.DataFrame]:
    dm = dm.copy()
    dm.columns = [c.upper() for c in dm.columns]
    dm["DOMAIN"] = "DM"
    dm["USUBJID"] = dm["SUBJECT_ID"].astype(str)
    dm["SUBJID"]  = dm["USUBJID"]
    dm["SITEID"]  = "RWED"
    for col in ("ARM", "ACTARM"):
        if col not in dm.columns:
            dm[col] = dm.get("ARMCD", "NOTASSIGNED")
    dm_keep = ["STUDYID", "DOMAIN", "USUBJID", "SUBJID", "SITEID",
               "AGE", "SEX", "RACE", "ARM", "ACTARM"]
    dm_out = dm[[c for c in dm_keep if c in dm.columns]]

    ex = claims.copy()
    ex["DOMAIN"]    = "EX"
    ex["USUBJID"]   = ex["SUBJECT_ID"].astype(str)
    ex["EXSEQ"]     = ex.groupby("USUBJID").cumcount() + 1
    ex["EXTRT"]     = ex["GENERIC_NAME"].str.upper()
    ex["EXDOSE"]    = ex["DOSE_MG"]
    ex["EXDOSU"]    = "mg"                          # UCUM
    ex["EXDOSFRM"]  = ex["FORM"].str.upper()
    ex["EXROUTE"]   = ex["ROUTE"].str.upper()
    ex["EXSTDTC"]   = ex["RX_DATE"].apply(to_iso)
    # EXENDTC is DERIVED from days_supply; flag this in define.xml algorithm.
    ex["EXENDTC"]   = (ex["RX_DATE"] + pd.to_timedelta(ex["DAYS_SUPPLY"].fillna(1).astype(int) - 1, unit="D")).apply(to_iso)
    ex["EXOCCUR"]   = "Y"
    ex_out = ex[["DOMAIN", "USUBJID", "EXSEQ", "EXTRT", "EXDOSE", "EXDOSU",
                 "EXDOSFRM", "EXROUTE", "EXSTDTC", "EXENDTC", "EXOCCUR"]]
    supp = pd.DataFrame({
        "STUDYID": "RWED-001",
        "RDOMAIN": "EX",
        "USUBJID": ex["USUBJID"],
        "IDVAR":   "EXSEQ",
        "IDVARVAL": ex["EXSEQ"].astype(str),
        "QNAM":    "NDC",
        "QLABEL":  "Source NDC",
        "QVAL":    claims["NDC"].astype(str),
        "QORIG":   "SDTM",
        "QEVAL":   "",
    })
    return ex_out, dm_out, supp

Citations

FOUNDATIONAL / METHODS
  1. [1]Rizzoli S, Ori A, Mignani A, Ferri F, Simoni L. Implementation of Clinical Data Interchange Standard Consortium (CDISC) standards to Real-World Data: Challenges and Strategies in SDTM Development in the Setting of Observational Studies. Journal of the Society for Clinical Data Management. 2026;6(1).
  2. [2]U.S. Food and Drug Administration. Providing Regulatory Submissions in Electronic Format — Standardized Study Data. FDA Guidance for Industry.
  3. [3]CDISC. Study Data Tabulation Model (SDTM). CDISC foundational standard.
APPLIED EXAMPLES
  1. [4]U.S. Food and Drug Administration. Study Data Technical Conformance Guide — Technical Specifications Document.