SDTM for Real-World Data Submissions
The CDISC Study Data Tabulation Model is the standardized structure and metadata format the FDA requires for study data packages; applying it to real-world data forces explicit modeling decisions about how to map claims dispensings, EHR encounters, and registry records into SDTM domains (DM, EX, AE, LB, etc.) and how to retain traceability from each tabulated value to the source record.
On this page
SDTM (Study Data Tabulation Model) is the standardized, row-by-row data format the FDA requires for study data submissions. Each piece of information gets placed in a labeled "domain" — demographics here, exposure there, adverse events in their own slot — along with the date, a coded value, and a controlled vocabulary term so a reviewer can read the file and reconstruct what happened. Real- world data rarely starts in this shape: claims are payment records, EHR data is visit notes, and registries are disease-specific. Mapping them to SDTM is real modeling work — every code translation and every date assumption has to be documented in the define.xml metadata file and the Analysis Data Reviewer's Guide so an FDA reviewer can trace any submitted value back to the original record.
The Study Data Tabulation Model (SDTM) is the standardized structure for submitting study data to the FDA. Each observation is placed in a domain (a logical grouping such as DM for demographics, EX for exposure, AE for adverse events, LB for laboratory, CM for concomitant medications) along with identifier variables, timing variables, and a controlled-terminology-coded value. The format is intentionally normalized — one row per observation — so a reviewer can reconstruct the analytic dataset from the tabulations and the define.xml metadata.
Why it matters for RWE
SDTM was designed for clinical-trial data, where a CRF captures exactly the variables of interest on a fixed schedule. RWE sources look very different: claims are encounter- and payment-driven, EHR data is visit- and problem-list-driven, registries are disease- and assessment-driven. Mapping RWD into SDTM is therefore a modeling exercise, not a data dump. Every mapping decision — which NDC rolls up to which generic name, which ICD-10-CM code maps to which MedDRA preferred term for an AE, how to derive EX dates from dispensed-supply intervals — must be documented in the Analysis Data Reviewer's Guide (ADRG) and the define.xml so a reviewer can reconstruct it. This visibility is exactly what the FDA Data Standards Catalog and the Study Data Technical Conformance Guide require.
Core domain mapping in RWD
- DM (Demographics): build from enrollment/insurance eligibility; populate ARM, ACTARM, SITEID, AGE, SEX, RACE, ETHNIC with the values available. In RWD there is no randomization arm — populate the planned and actual arms with the study-defined exposure (e.g., treatment vs comparator).
- EX (Exposure): claims dispensings or EHR administrations become EX records, with EXDOSE, EXDOSU, EXDOSFRM, EXROUTE, EXSTDTC, EXENDTC from NDC + days supply (claims) or mar/admin records (EHR). Grace periods and carryover are operationalized before SDTM build, not after.
- AE (Adverse Events): RWD usually does not have AEs in the CRF sense. Map EHR problem lists, diagnosis codes flagged as "incident" during follow-up, or spontaneous-report-system records to AE; translate source codes to MedDRA preferred terms (PT) using the MedDRA MSSO browser and document the version in define.xml.
- CM (Concomitant Medications): claims dispensings for non-study medications, mapped to WHODrug B3 / C3 format with CMTRT, CMCAT, CMSTDTC, CMENDTC, and route.
- DS (Disposition): derive from enrollment spans, plan switches, death records, and end-of- follow-up rules; populate DSCAT, DSDECOD, DSTERM, DSSTDTC.
- LB (Laboratory): EHR only; LOINC-code the test, populate LBORRES, LBORRESU, LBNRIND, LBBLFL, LBTESTCD, LBTEST, LBCAT, LBFAST, LBSPEC, LBMETHOD.
- MH (Medical History): baseline-window diagnoses from claims or EHR, with MHCAT and the appropriate MedDRA / ICD coding.
Common pitfalls in RWD-to-SDTM conversion
- Date imputation and partial dates. SDTM ISO 8601 timestamps require precision. RWD often has only month or year. Apply pre-specified partial-date imputation rules (e.g., first-of-month if only month is known, midpoint if only year), document the rule in define.xml, and never silently leave blanks.
- Coding version drift. MedDRA, WHODrug, LOINC, and ICD versions are time-stamped. The version used at SDTM build must be locked in the metadata and any re-coding (e.g., annual MedDRA re-coding) is a documented version change.
- Lossy mappings. When many source codes collapse to one SDTM value, the audit chain breaks unless supplemental qualifiers (SUPP--.QVAL) or a custom domain (with --TORG prefix) preserve the source value. The FDA Technical Conformance Guide calls these out as needed for traceability.
- Race and ethnicity handling. Race is federally standardized (OMB categories) but RWD often has missing or imputed values; document the source, the imputation (if any), and never aggregate away race categories without a stated rationale.
- Unit standardization. Claims have no units. EHR LBs must be harmonized to UCUM (Unified Code for Units of Measure); document any unit conversions in the ADRG.
Operational depth for common RWD sources
- Claims (MarketScan, Optum, Truveta, IQVIA, SNDS): rich dispensings (EX, CM), diagnoses (AE, MH when coded), but no labs (LB) and no vital signs (VS). Payments become custom cost domains or are absorbed into ADSL cost fields. Date precision is usually exact.
- EHR (Epic, Cerner, HealthVerity): rich LB, VS, AE (when problem-list updated), EX (when administered), but encounter-level only. Date precision is exact but visit-driven; out-of-system care is invisible.
- Linked claims–EHR: best of both, with explicit record-linkage provenance in a custom SUPPQUAL.
- Registries (SEER-Medicare, CPRD, disease registries): disease-specific variables populate custom domains with --TORG prefixes; linked to claims via the linkage algorithm documented in the ADRG.
FDA submission posture
The FDA Data Standards Catalog lists the SDTM and ADaM versions required for the submission period; the Study Data Technical Conformance Guide specifies file formats (XPT v5), define.xml v2.x structure, the ADRG and SDRG expectations, and the conformance validator rules. A passing OpenCDISC/Pinnacle 21 validation is a prerequisite for FDA acceptance of the submission package; integration of the validator report is part of the ADRG. Pros, cons, and trade-offs.
- vs leaving RWD in source format: SDTM imposes structure the FDA can review programmatically, at the cost of a genuine mapping exercise; unmapped source data is not acceptable for regulated submissions regardless.
- vs mapping directly to ADaM: skipping SDTM collapses traceability — reviewers cannot reconstruct analysis values from tabulations, which is precisely what the Technical Conformance Guide exists to enforce.
- Trade-off: fidelity vs normalization. Every lossy collapse (many NDCs to one generic name) must be preserved in SUPPQUAL or a custom domain or the audit chain breaks.
When NOT to use
Do not force RWD into SDTM for journal publications or internal exploratory analyses — the cost buys regulatory reviewability, not analytic insight; use analysis-ready source structures instead. And do not treat SDTM as an analytic layer: it is tabulation, not analysis — deriving endpoints belongs in ADaM.
Decision diagram
flowchart LR
R[Raw RWD<br/>claims / EHR / registry] --> C[Curate & reconcile<br/>linkage, dedup, dates]
C --> M[Map to SDTM domains<br/>DM, EX, AE, CM, MH, DS, LB...]
M --> CT[Apply controlled terminology<br/>WHODrug, MedDRA, LOINC, UCUM]
CT --> V[Pinnacle 21 conformance<br/>validator pass]
V --> X{XPT v5<br/>package}
X --> D[define.xml v2.x<br/>+ ADRG + SDRG]
D --> F[FDA submission<br/>package]Worked example
Scenario
We need to convert a single patient's claims-derived dispensings into SDTM EX records so they can be included in an FDA submission. The patient is 58 years old, commercially insured, with three dispensings of 30-count 10mg atorvastatin over 90 days. We build one EX record per dispensing with the dose, unit, form, route, start, and end date populated, and one CM record for a concurrent 30-day amoxicillin course to show the convention.
Dataset
One patient's dispensings mapped into SDTM EX and CM records.
| domain | USUBJID | EXSEQ/CMSEQ | EXTRT/CMTRT | EXDOSE/CMDOSE | EXDOSU/CMDOSU | EXDOSFRM/CMDOSFRM | EXROUTE/CMROUTE | EXSTDTC/CMSTDTC | EXENDTC/CMENDTC |
|---|---|---|---|---|---|---|---|---|---|
| EX | RWED-001 | 1 | ATORVASTATIN | 10 | mg | TABLET | ORAL | 2025-01-15 | 2025-02-13 |
| EX | RWED-001 | 2 | ATORVASTATIN | 10 | mg | TABLET | ORAL | 2025-02-14 | 2025-03-14 |
| EX | RWED-001 | 3 | ATORVASTATIN | 10 | mg | TABLET | ORAL | 2025-03-15 | 2025-04-13 |
| CM | RWED-001 | 1 | AMOXICILLIN | 500 | mg | CAPSULE | ORAL | 2025-02-20 | 2025-03-21 |
Steps
Result
Three EX records (ATORVASTATIN 10 mg TABLET ORAL, 30-day supplies) and one CM record (AMOXICILLIN 500 mg CAPSULE ORAL) under subject RWED-001, with all dates ISO 8601, all values controlled-terminology coded, and the SDTMIG / WHODrug / MedDRA versions locked in the ADRG and define.xml.
Trade-offs
Runnable example
Toy SDTM EX builder from claims-style long-form input. Required inputs: claims : subject_id, ndc, generic_name, rx_date (date), days_supply, dose_mg, form, route dm : subject_id, armcd, age, sex, race Returns EX and DM dataframes with ISO-8601 dates, controlled-terminology upper-case strings, and a SUPP--EX data...
import pandas as pd
SDTMIG_VERSION = "3.4" # update per submission lock
WHODRUG_VERSION = "C3-2024" # record in define.xml
CONTROLLED_TERM = "2024-12-19" # CDISC controlled-terminology package version
def to_iso(d: pd.Timestamp) -> str:
# SDTM uses ISO 8601 dates; partial dates use SDTM imputation rules (documented in define.xml).
return d.strftime("%Y-%m-%d")
def build_sdtm_ex(claims: pd.DataFrame, dm: pd.DataFrame) -> tuple[pd.DataFrame, pd.DataFrame]:
dm = dm.copy()
dm.columns = [c.upper() for c in dm.columns]
dm["DOMAIN"] = "DM"
dm["USUBJID"] = dm["SUBJECT_ID"].astype(str)
dm["SUBJID"] = dm["USUBJID"]
dm["SITEID"] = "RWED"
for col in ("ARM", "ACTARM"):
if col not in dm.columns:
dm[col] = dm.get("ARMCD", "NOTASSIGNED")
dm_keep = ["STUDYID", "DOMAIN", "USUBJID", "SUBJID", "SITEID",
"AGE", "SEX", "RACE", "ARM", "ACTARM"]
dm_out = dm[[c for c in dm_keep if c in dm.columns]]
ex = claims.copy()
ex["DOMAIN"] = "EX"
ex["USUBJID"] = ex["SUBJECT_ID"].astype(str)
ex["EXSEQ"] = ex.groupby("USUBJID").cumcount() + 1
ex["EXTRT"] = ex["GENERIC_NAME"].str.upper()
ex["EXDOSE"] = ex["DOSE_MG"]
ex["EXDOSU"] = "mg" # UCUM
ex["EXDOSFRM"] = ex["FORM"].str.upper()
ex["EXROUTE"] = ex["ROUTE"].str.upper()
ex["EXSTDTC"] = ex["RX_DATE"].apply(to_iso)
# EXENDTC is DERIVED from days_supply; flag this in define.xml algorithm.
ex["EXENDTC"] = (ex["RX_DATE"] + pd.to_timedelta(ex["DAYS_SUPPLY"].fillna(1).astype(int) - 1, unit="D")).apply(to_iso)
ex["EXOCCUR"] = "Y"
ex_out = ex[["DOMAIN", "USUBJID", "EXSEQ", "EXTRT", "EXDOSE", "EXDOSU",
"EXDOSFRM", "EXROUTE", "EXSTDTC", "EXENDTC", "EXOCCUR"]]
supp = pd.DataFrame({
"STUDYID": "RWED-001",
"RDOMAIN": "EX",
"USUBJID": ex["USUBJID"],
"IDVAR": "EXSEQ",
"IDVARVAL": ex["EXSEQ"].astype(str),
"QNAM": "NDC",
"QLABEL": "Source NDC",
"QVAL": claims["NDC"].astype(str),
"QORIG": "SDTM",
"QEVAL": "",
})
return ex_out, dm_out, suppR version using data.table. Mirrors the Python EX builder: ingests long-form claim records and a demographics table, emits SDTM-format EX and DM data frames with EXENDTC derived from days_supply, plus a SUPP--EX data frame carrying the source NDC for traceability.
library(data.table)
SDTMIG_VERSION <- "3.4" # lock per submission
WHODRUG_VERSION <- "C3-2024" # record in define.xml
build_sdtm_ex <- function(claims, dm) {
setDT(claims); setDT(dm)
dm[, DOMAIN := "DM"]
dm[, USUBJID := as.character(SUBJECT_ID)]
dm[, SUBJID := USUBJID]
if (!"SITEID" %in% names(dm)) dm[, SITEID := "RWED"]
if (!"ARM" %in% names(dm) && "ARMCD" %in% names(dm)) dm[, ARM := ARMCD]
if (!"ACTARM" %in% names(dm) && "ARMCD" %in% names(dm)) dm[, ACTARM := ARMCD]
ex <- copy(claims)
ex[, DOMAIN := "EX"]
ex[, USUBJID := as.character(SUBJECT_ID)]
ex[, EXSEQ := seq_len(.N), by = USUBJID]
ex[, EXTRT := toupper(GENERIC_NAME)]
ex[, EXDOSE := DOSE_MG]
ex[, EXDOSU := "mg"]
ex[, EXDOSFRM := toupper(FORM)]
ex[, EXROUTE := toupper(ROUTE)]
ex[, EXSTDTC := format(as.Date(RX_DATE), "%Y-%m-%d")]
ex[, EXENDTC := format(as.Date(RX_DATE) + pmax(as.integer(DAYS_SUPPLY) - 1L, 0L), "%Y-%m-%d")]
ex[, EXOCCUR := "Y"]
ex_out <- ex[, .(DOMAIN, USUBJID, EXSEQ, EXTRT, EXDOSE, EXDOSU,
EXDOSFRM, EXROUTE, EXSTDTC, EXENDTC, EXOCCUR)]
supp <- data.table(
STUDYID = "RWED-001", RDOMAIN = "EX", USUBJID = ex$USUBJID,
IDVAR = "EXSEQ", IDVARVAL = as.character(ex$EXSEQ),
QNAM = "NDC", QLABEL = "Source NDC", QVAL = as.character(claims$NDC),
QORIG = "SDTM", QEVAL = ""
)
list(ex = ex_out, dm = dm, supp = supp)
}SAS build using PROC SQL for the SDTM EX domain. Maps claims-style long-form input (work.claims) and a demographics table (work.dm) into SDTM EX, plus a SUPP--EX data set for NDC traceability. Dates are ISO 8601; EXENDTC is derived as RX_DATE + DAYS_SUPPLY - 1 (flag in define.xml).
/* Constants: lock per submission package. */
%let sdtmig_version = 3.4;
%let whodrug_version = C3-2024;
/* EX domain from long-form claims. */
proc sql;
create table work.ex as
select
"EX" as DOMAIN,
put(subject_id, best.) as USUBJID,
monotonic() as EXSEQ, /* sequence within subject */
upcase(generic_name) as EXTRT,
dose_mg as EXDOSE,
"mg" as EXDOSU,
upcase(form) as EXDOSFRM,
upcase(route) as EXROUTE,
put(rx_date, yymmdd10.) as EXSTDTC,
put(rx_date + days_supply - 1, yymmdd10.) as EXENDTC,
"Y" as EXOCCUR
from work.claims
group by subject_id, rx_date
order by subject_id, rx_date;
quit;
/* SUPP--EX to preserve source NDC. */
data work.supp_ex;
length STUDYID $8 RDOMAIN $8 IDVAR $8 QNAM $8 QLABEL $40 QORIG $6 QEVAL $1;
STUDYID = "RWED-001"; RDOMAIN = "EX"; IDVAR = "EXSEQ";
set work.claims;
IDVARVAL = put(_N_, best.);
QNAM = "NDC"; QLABEL = "Source NDC"; QVAL = put(ndc, $20.); QORIG = "SDTM"; QEVAL = "";
keep STUDYID RDOMAIN USUBJID IDVAR IDVARVAL QNAM QLABEL QVAL QORIG QEVAL;
run;Citations
- [1]Rizzoli S, Ori A, Mignani A, Ferri F, Simoni L. Implementation of Clinical Data Interchange Standard Consortium (CDISC) standards to Real-World Data: Challenges and Strategies in SDTM Development in the Setting of Observational Studies. Journal of the Society for Clinical Data Management. 2026;6(1).
- [2]U.S. Food and Drug Administration. Providing Regulatory Submissions in Electronic Format — Standardized Study Data. FDA Guidance for Industry.
- [3]CDISC. Study Data Tabulation Model (SDTM). CDISC foundational standard.