Flatiron Health Clinico-Genomic Database
An oncology-specific US real-world database combining curated electronic health records from hundreds of community cancer clinics (the Flatiron Network) with linked genomic sequencing data via the Foundation Medicine clinico-genomic linkage - providing tumor-level clinical depth plus molecular profiling unavailable in claims or generic EHR sources.
On this page
Flatiron built the leading US oncology RWD asset by combining structured-and-abstracted EHR data from community oncology practices (diagnoses, regimens, lines of therapy, progression, labs, vitals, mortality) with Foundation Medicine genomic profiles linked at the patient level. It is the reference source for biomarker-defined treatment-pattern and outcomes research in advanced cancers.
Flatiron Health's Clinico-Genomic Database (CGDB)
merges two oncology data layers: (1) curated EHR data from the Flatiron Network of ~280+ US community cancer clinics, where technology-assisted abstraction adds what raw EHR lacks — confirmed diagnoses, regimen definitions, lines of therapy, progression events, and mortality follow-up; and (2) Foundation Medicine comprehensive genomic profiling linked at the patient level, supplying mutation, copy-number, and fusion status plus tumor mutational burden and microsatellite instability.
Why it matters for RWE
Oncology exposes the weakness of generic sources: progression is never billed, regimens are hard to infer from claims J-codes alone, and biomarker status determines eligibility for modern therapy. Flatiron solves all three by construction — human-in-the-loop abstraction validates regimen starts/stops and progression, and CG linkage makes molecularly defined cohorts possible. Published comparisons show Flatiron populations resemble treated oncology populations at academic/community mix better than SEER-Medicare resembles all comers (who include untreated patients).
Operational characteristics
- Curated variables: line-of-therapy numbering, response/progression dates, ECOG-derived performance indicators, mortality via enhanced multi-source follow-up.
- Genomic layer: CGP reports (hundreds of genes), TMB, MSI, and selected PD-L1 results; testing availability is itself selection-conditioned (tested patients skew toward advanced/metastatic settings).
- Population: predominantly advanced/metastatic disease in community settings; under-represents early-stage and purely academic-managed patients.
- Regulatory standing: FDA has accepted Flatiron-based external controls in oncology submissions; the platform's curation processes are documented in published validation studies.
Common pitfalls
- Testing-conditional selection: genomic analyses describe tested patients only; test indication patterns create spectrum bias.
- Progression imprecision: abstracted progression improves on claims but remains scan-cycle-bound and reviewer-dependent.
- Line-of-therapy conventions: definitions (counting maintenance, peri-operative, rechallenge) must match study objectives explicitly.
- Community-clinic geography: coverage tracks the Flatiron Network footprint; not nationally representative by design.
Pros, cons, and trade-offs
- vs SEER-Medicare: clinical/genomic depth and all-age reach vs population-based incidence capture and elderly cost data.
- vs claims sources: progression/regimen fidelity vs complete utilization/cost capture.
- Trade-off: treated-cohort realism vs population denominators — Flatiron describes treated cancer patients superbly but cannot speak to undiagnosed/untreated populations.
When NOT to use
Incidence/prevalence estimation; cost-of-care requiring payer-paid amounts; early-stage/screening populations; questions about untested patients.
Decision diagram
flowchart LR FN[Flatiron Network clinics - 280+] --> TA[Tech-assisted abstraction\nlines - progression - death] FM[Foundation Medicine CGP reports] --> LK[Patient-level linkage] TA --> CGDB[Clinico-Genomic Database] LK --> CGDB CGDB --> O[Biomarker-defined oncology RWE\nexternal controls]
Worked example
Scenario
Describe overall survival by EGFR mutation class among first-line osimertinib-treated metastatic NSCLC patients.
Dataset
OS by EGFR mutation subclass in first-line osimertinib users.
| mutation_class | n | median_os_mo | hr_vs_common |
|---|---|---|---|
| ex19del/L858R - 1240 - 38.6 - 1.00 | |||
| ex20ins - 118 - 13.1 - 3.05 | |||
| G719X/S768I - 84 - 29.8 - 1.31 |
Steps
Result
Ex20ins subgroup showed markedly inferior OS (13.1 vs 38.6 months); results robust to landmark and censoring-at-last-visit choices.
Trade-offs
Runnable example
First-line cohort definition with OS estimation skeleton from Flatiron-style frames.
\
import pandas as pd
from lifelines import KaplanMeierFitter
def first_line_cohort(regimens, dx, min_stage="IIIB"):
r1 = regimens[regimens.line_number == 1].sort_values(["patient_id","start_date"])
idx = r1.groupby("patient_id").first().reset_index()
return idx.merge(dx[["patient_id","stage","histology"]], on="patient_id")
def os_curve(cohort, cutoff):
kmf = KaplanMeierFitter()
t = (cohort.death_date.fillna(cutoff) - cohort.start_date).dt.days.clip(lower=0)
e = cohort.death_date.notna().astype(int)
kmf.fit(t, e)
return kmf.median_survival_time_, kmf.confidence_interval_
R version using survfit for OS by mutation class.
\
library(survival)
os_by_class <- function(cohort, cutoff) {
cohort$time <- pmin(as.numeric(cohort$death_date - cohort$start_date, units = "days"),
as.numeric(cutoff - cohort$start_date, units = "days"))
cohort$event <- as.integer(!is.na(cohort$death_date))
survfit(Surv(time, event) ~ mutation_class, data = cohort) |>
surv_median()
}
SAS PROC LIFETEST version.
\
proc lifetest data=coh plots=survival(cb);
time os_days * death(0);
strata mutation_class / test=logrank adjust=sidak;
run;
/* Median OS with CI */
ods output quartiles=q;
proc lifetest data=coh;
time os_days * death(0);
strata mutation_class;
run;
Citations
- [1]Snow TS, et al. Comparison of Population Characteristics in Real-World Clinical Oncology Databases in the US: Flatiron Health-Foundation Medicine Clinico-Genomic Databases, Flatiron Health Research Databases, and the National Cancer Database and SEER. medRxiv. 2023.
- [2]Ma X, et al. Comparison of Population Characteristics in Real-World Clinical Oncology Databases in the US: Flatiron Health, National Cancer Database, and SEER. medRxiv. 2020.
- [3]Flatiron Health. Flatiron Data and Research documentation.