← Methods repository
CONCEPTADVANCEDPYTHON · R · SASlast reviewed 2026-08-24 · updated 2026-10-07 · 4 citations

Tables, Listings, and Figures (TLFs) for RWE Submissions

Tables, Listings, and Figures (TLFs) are the reviewer-visible evidence artifacts in an FDA submission — generated from ADaM analysis datasets — that summarize baseline characteristics, disposition, efficacy, safety, and exploratory analyses; for RWE the TLF generation pipeline must reproduce the same displays from the locked ADaM package, with each TLF carrying its ARM metadata so a reviewer can re-derive any cell from the supporting ADaM records.

Data Standardtlftables-listings-figuressubmission-standardsarmanalysis-results-metadatafda-submissionrwe-submissionr4csr
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

Tables, Listings, and Figures are the headline evidence artifacts in any FDA submission — the demographic table, the primary endpoint result, the safety summary, the Kaplan-Meier curve. They are generated from ADaM analysis datasets, not built by hand. For RWE submissions the same ADaM-to-TLF pipeline applies, but with an extra requirement: each TLF must carry an ARM (Analysis Results Metadata) reference back to the ADaM record so a reviewer can re-derive any cell.

When to use it
—Always for an FDA submission. Hand-built tables cannot be reproduced or audited.
—When the sponsor is R-native and willing to validate a smaller number of OSS packages.
Watch out for
—Requires a programming function for every table; a 'one-off' table is rarely worth the upfront investment.
—Some packages require R 4.x; some RTF style requirements (e.g., specific fonts, page numbering) are sponsor-specific.

Tables, Listings, and Figures (TLFs)

are the headline evidence artifacts of an FDA submission: tables summarize baseline characteristics, primary and secondary endpoints, and safety; listings dump subject-level records showing every observation behind table cells; figures display Kaplan-Meier curves, forest plots, swimmer plots, and CONSORT-style flow diagrams. Together they turn the ADaM analysis package into the clinical study report (CSR) body and the regulatory evidence package a review team actually reads.

Why they matter for RWE

In a trial, TLF shells follow a protocol-specified statistical analysis plan. In an RWE package, the TLFs must additionally make the design legible: the cohort-attrition flow figure (how many patients survived each eligibility criterion — the observational analogue of CONSORT), baseline tables that reveal covariate imbalance driving any weighting, weighted vs unweighted sensitivity columns, and KM figures whose numbers-at-risk tables expose censoring patterns. Reviewers of external-control arguments scrutinize TLFs for exactly the things observational designs can hide: immortal-time artifacts in time-zero alignment, differential depletion, positivity violations visible as extreme weights or empty strata.

Production discipline

  • Shells first. Mockup every table aligned to the SAP/population definitions before programming; late shell changes cascade through validation.
  • Double programming. The primary programmer produces each output; an independent validator re-produces it from the specifications; outputs reconcile cell-by-cell. Pinnacle 21 and ADRG cross-checks supplement but do not replace this.
  • Macro standardization. Population filters come from ADSL flags only — never hard-coded subsets inside table macros, which is how population drift sneaks in.
  • Numbering and cross-referencing. Consistent numbering (Table 14.x.x conventions), footers stating population, cut date, and analysis dataset versions.

RWE-specific content requirements

  • Cohort disposition figure: eligibility funnel from source population to analysis cohorts, with counts and reasons at each step.
  • Baseline tables by cohort and by weighting scheme: unweighted, IPTW-weighted (with standardized mean differences), and any matched sample side-by-side.
  • Follow-up and censoring displays: KM curves with numbers at risk, censoring marks, and median follow-up by arm.
  • Sensitivity tiers: pre-specified alternative definitions (on-treatment vs ITT-analogue, different censoring rules, outcome windows) presented consistently.

Pros, cons, and trade-offs

  • vs ad hoc summary tables: TLFs carry SAP alignment, double-programming verification, and reconciliation pedigree; ad hoc outputs do not survive regulatory scrutiny.
  • vs figures alone: curves communicate but hide cell counts; tables carry the denominators reviewers check. Use both, cross-referenced.
  • Trade-off: completeness vs readability — every output must serve the CSR narrative; redundant cuts multiply validation cost without evidentiary gain.

When NOT to use

Exploratory sensitivity runs stay in the statistical appendix or off-submission environment. Never present a TLF whose population, censoring rules, or windows differ from the locked SAP without a documented amendment.

Decision diagram

flowchart LR
  ADaM[ADaM analysis datasets<br/>ADSL, ADTTE, ADEFF, ADAE] --> T1[Table 1: Baseline characteristics]
  ADaM --> Disp[Disposition table]
  ADaM --> PE[Primary endpoint table]
  ADaM --> Sec[Secondary endpoint tables]
  ADaM --> Saf[Safety summary tables]
  ADaM --> List[Subject listings]
  ADaM --> Fig[Figures: KM curves, forest plots]
  ADaM --> ARM[ARM XML extension to define.xml]
  T1 --> CSR[Clinical study report]
  Disp --> CSR
  PE --> CSR
  Sec --> CSR
  Saf --> CSR
  List --> CSR
  Fig --> CSR
  ARM --> CSR
ADaM-to-TLF pipeline. Every TLF is generated from a locked ADaM dataset; ARM links each TLF to its source records; the bundle goes into the CSR submission package.

Worked example

Scenario

We need a Table 14.2.1 (Primary Endpoint) and a Table 14.1.1 (Baseline Characteristics) for an RWE submission comparing Drug A vs Drug B with a time-to-first-hospitalization outcome. Both come from the locked ADaM package; the ARM extension in define.xml carries the Population filter (ITTFL='Y'), the Analysis (ADTTE PARAMCD=TTFHOSP), and the Programming Statement that produced each cell.

Dataset

Headline TLF summaries from the locked ADaM ADTTE.

armneventsperson_yearsevents_per_100py
DRUG A - 840 - 61 - 1124 - 5.4
DRUG B - 799 - 83 - 1081 - 7.7

Steps

1Lock the ADaM package (datasets + define.xml + ARM) before generating TLFs.
2Generate Table 14.1.1 from ADSL; the ARM Output references the ADSL ITTFL filter and the variables AGE, SEX, RACE.
3Generate Table 14.2.1 from ADTTE; the ARM Output references PARAMCD=TTFHOSP, ITTFL filter, and the Cox PH programming statement.
4Run the validator's independent re-implementation and reconcile cell-by-cell.
5Package the RTF outputs, the XPT files, define.xml with ARM, the ADRG, and the SAP into the m5 submission section.

Result

Two RTF tables delivered: Table 14.1.1 (baseline characteristics by ACTARM) and Table 14.2.1 (primary endpoint summary with events and person-time by ACTARM). The ARM XML in define.xml carries the Output definitions, Analysis references, Filters, and Programming Statements. A reviewer opens the RTF, clicks the ARM reference, and re-derives any cell from the ADaM records.

Trade-offs

vs. Hand built tables in Excel
Pros of this
—Code-driven TLF generation is reproducible, auditable, and version-controlled; every cell traces back to a row in ADaM.
vs. OSS R packages (r2rtf, tern, gtreg)
Pros of this
—OSS R packages produce RTF outputs that match sponsor standards (PharmaSUG, ICH); ARM metadata can be generated alongside.

Runnable example

Toy pandas functions for Table 1 and the primary endpoint table from ADaM inputs; production pipelines use validated, locked libraries.

requires: pandas
# Toy example: ADaM ADSL + ADTTE -> Table 1 baseline summary + primary endpoint table.
# In a real submission these are validated functions with locked source; the toy shows the
# cell-level trace-back discipline.
import pandas as pd

def table_1(adsl: pd.DataFrame, value_vars=("AGE", "SEX", "RACE")) -> pd.DataFrame:
    # Baseline characteristics summary, stratified by ACTARM. Each cell links to a row.
    rows = []
    for v in value_vars:
        summary = adsl.groupby(["ACTARM", v]).size().unstack(fill_value=0)
        for arm in summary.index:
            for val, n in summary.loc[arm].items():
                rows.append({"Variable": v, "Arm": arm, "Value": val, "N": int(n)})
    return pd.DataFrame(rows)

def primary_endpoint_table(adtte: pd.DataFrame) -> pd.DataFrame:
    # Hazard ratio + 95% CI for the first PARAMCD; in production use lifelines or statsmodels.
    # Toy: just count events and person-time by ACTARM.
    out = (adtte.groupby("ACTARM")
                 .agg(events=("CNSR", lambda s: int((s == 0).sum())),
                      censored=("CNSR", lambda s: int((s == 1).sum())),
                      person_days=("AVAL", "sum"))
                 .reset_index())
    out["person_years"] = (out["person_days"] / 365.25).round(1)
    return out

Citations

FOUNDATIONAL / METHODS
  1. [1]R Consortium R for Clinical Study Reports and Submission Working Group. TLF Overview. r4csr.org.
  2. [2]CDISC. Analysis Results Metadata (ARM) v1.0 for Define-XML v2.0. CDISC standards.
  3. [3]U.S. Food and Drug Administration. Study Data Technical Conformance Guide — Technical Specifications Document.
APPLIED EXAMPLES
  1. [4]Yan F, Zhang Y, Tian Y. Evidence Behind the Automation of Clinical Trial Statistical Programming: A Scoping Review of Technology Adoption, Validation Frameworks, and AI/ML Integration (2020-2025). medRxiv. 2025.