← Methods repository
CONCEPTINTERMEDIATEPYTHON · R · SASlast reviewed 2026-08-25 · updated 2026-08-25 · 4 citations

Truveta

A US multi-health-system real-world data platform combining full-fidelity EHR data from dozens of participating provider organizations with linked claims and mortality data, normalized to a common data model and refreshed continuously — distinctive for retaining source-level clinical detail (vitals, labs, notes-derived variables) that claims-only sources lack.

Data Sourcetruvetadata-sourceehrmulti-health-systemclaims-linkageus-rwd
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

Truveta aggregates de-identified electronic health records from a large consortium of US health systems (including providers such as Providence and Trinity Health) into one research-accessible platform. Unlike claims databases it captures clinical richness — lab values, vitals, imaging indications, clinician notes — and unlike single-system EHR extracts it spans many regions and payer types. Data is cleaned to a common model daily, can be linked to insurance claims and Social Security Death Master File mortality, and is queried in an analytic workbench rather than shipped as flat files. Its main limitations are the EHR-native ones: care outside participating systems is invisible, coding follows billing/problem-list incentives, and completeness varies by how each system implements its EHR.

When to use it
—Endpoint requires clinical confirmation or biomarkers; population includes pediatrics/commercially insured.
—Multi-site generalizable inference; avoid when one system's curated registry already answers the question.
—Truveta when clinical fidelity + governance matter; OMOP networks when breadth across dozens of databases matters.
Watch out for
—Out-of-system care invisible; shorter continuous observation per patient than mature claims enrollment spans.
—Cross-system harmonization noise; no single site's deep local completeness.
—Closed workbench vs open federated analytics across many more databases.

Truveta

is a real-world data platform founded in 2020 that aggregates de-identified, full-fidelity EHR data from a consortium of US health systems (30+ systems spanning thousands of care sites), together with linked insurance claims and mortality (Social Security Death Master File) data. Where legacy claims databases capture billing events and legacy single-system EHR extracts lack generalizability, Truveta's differentiator is multi-system clinical fidelity at national scale: laboratory results with units and reference ranges, vital signs, immunizations, clinician documentation, and structured problem lists — retained close to source fidelity and mapped to a common data model that is re-processed continuously as upstream EHR data updates.

Why it matters for RWE

Truveta occupies the middle of the claims-vs-EHR trade-off: richer than claims for clinical phenotyping (lab-confirmed outcomes, disease severity, pregnancy gestational age), broader than any single health system for generalizability across geography, age, and payer mix. The continuous-refresh architecture supports rapid-cycle evidence (its original public-health use case was near-real-time COVID-19 monitoring) but also introduces versioning discipline requirements: analyses must record the data cut date because point estimates move as back-filled records arrive.

Operational characteristics

  • Population: tens of millions of patients, skewed toward systems' catchment areas (strong West/Midwest/South representation depending on participating systems); all ages including pediatric, unlike Medicare-restricted sources.
  • Encounter-driven capture: visits generate data; out-of-system care (filling a prescription elsewhere, hospitalization at a non-member system) is invisible unless captured in problem lists or reconciliation fields.
  • Linkage: within-platform patient identity resolution across member systems plus linkage to claims and death data; researchers cannot bring their own identifiers (privacy-preserving tokenization inside the platform).
  • Access model: cloud workbench analysis (code-in, results-out); patient-level export is generally not part of the standard model — plan analytic pipelines accordingly.

Common pitfalls

  • Treating EHR presence as coverage. A patient with sparse encounters has missing data, not absence of events; distinguish 'no visit' from 'no event'.
  • Ignoring cut-date versioning. Back-filled labs and late-arriving encounter records change results between cuts; lock and report the data snapshot.
  • Outcome ascertainment drift. Lab-based outcomes (e.g., HbA1c-defined control) are strong; diagnosis-code-only outcomes inherit problem-list noise.
  • Denominator confusion. The enrolled-ish population is 'people with ≥1 encounter', not a stable insurance denominator — incidence rates need careful person-time definitions.

Pros, cons, and trade-offs

  • vs commercial claims (MarketScan, Optum): clinical depth (labs, vitals, notes) vs longitudinal continuity and complete out-of-system capture; claims see every filled script regardless of where care happens, EHR sees only in-system activity.
  • vs single health-system EHR: generalizability and scale vs within-system record completeness; single-system data has fewer cross-site gaps for patients who stay local.
  • Trade-off: the workbench access model controls privacy leakage but constrains tooling; methods requiring patient-level export or bespoke packages may not be portable.

When NOT to use

Not ideal for long pre-period lookback beyond the platform's accumulation window per system, for outcomes almost exclusively captured out-of-system (e.g., retail pharmacy fills without EHR reconciliation), or for studies needing patient-level data export to local infrastructure under the standard access model.

Decision diagram

flowchart LR
  HS1[Health System A EHR] --> T
  HS2[Health System B EHR] --> T
  HS3[... 30+ systems] --> T
  CL[Claims feed] --> L[Identity tokenization & linkage] --> T[Truveta common data model\ncontinuous refresh]
  D[Social Security mortality] --> L
  T --> W[Cloud workbench\ncode-in / results-out]
  W --> S[RWE study\ndata-cut stamped]
Truveta architecture: multi-system EHR plus claims/mortality feeds, tokenized linkage, common-model refresh, workbench access.

Worked example

Scenario

Compare time-to-HbA1c<7% after initiating two GLP-1 receptor agonists using Truveta EHR labs, with 1-year follow-up.

Dataset

Lab-based endpoint construction for two treatment cohorts.

cohortn_initiatorsmedian_months_to_goalhazard_ratio_adj
drug_A - 4120 - 7.8 - 1.00 (ref)
drug_B - 3865 - 9.1 - 0.86

Steps

1Define initiator cohort: first dispensing/order of each agent with 12-month baseline encounters and baseline HbA1c 7.5-11%.
2Build HbA1c series per patient from LOINC-coded lab results; define goal attainment as first result <7% after index+30 days.
3Fit Cox models adjusted for baseline HbA1c, comorbidity burden, age, sex; censor at last encounter (inverse-probability-of-censoring sensitivity).
4Report the data cut date and repeat on the next quarterly cut to quantify refresh drift.

Result

Adjusted HR 0.86 (95% CI 0.80-0.92) favoring drug A on the locked March cut; the June cut shifted the estimate by <2%, supporting robustness.

Trade-offs

vs. Commercial claims databases (MarketScan, Optum)
Pros of this
—Clinical granularity — labs, vitals, notes-derived variables — enabling phenotype precision claims cannot support.
vs. Single health system EHR
Pros of this
—Multi-region scale and heterogeneity improve generalizability and subgroup power.
vs. OMOP network studies
Pros of this
—Single harmonized model with governed provenance and consistent refresh, vs heterogeneous site-specific mappings.

Runnable example

Cohort construction and lab-series endpoint building against an extracted Truveta-style EHR frame (person, encounter, lab_result, drug_exposure tables).

requires: pandas
import pandas as pd

def hba1c_goal_cohort(person, encounter, lab, drug, index_drugs=("A","B"),
                      goal=7.0, baseline_lo=7.5, baseline_hi=11.0):
    # First qualifying dispensing per person
    dx = drug[drug.drug_name.isin(index_drugs)].sort_values(["person_id","drug_date"])
    idx = dx.groupby("person_id").first().reset_index()
    idx = idx.rename(columns={"drug_date":"index_date","drug_name":"cohort"})

    # Baseline window: 180 days pre-index
    base = lab.merge(idx[["person_id","index_date"]], on="person_id")
    base = base[(base.result_date < base.index_date) &
                (base.result_date >= base.index_date - pd.Timedelta(days=180)) &
                (base.loinc == "4548-4")]
    bval = base.groupby("person_id").result_value.median().rename("baseline_a1c")
    idx = idx.merge(bval, on="person_id")
    idx = idx[idx.baseline_a1c.between(baseline_lo, baseline_hi)]

    # Follow-up: first post-index result below goal (>=30 days post-index)
    fu = lab.merge(idx[["person_id","index_date"]], on="person_id")
    fu = fu[(fu.loinc=="4548-4") & (fu.result_date >= fu.index_date + pd.Timedelta(days=30))]
    first_goal = (fu[fu.result_value < goal]
                  .groupby("person_id").result_date.min().rename("goal_date"))
    return idx.merge(first_goal, left_on="person_id", right_index=True, how="left")

# Person-time and KM estimation proceed with standard survival libraries;
# always attach the data-cut date to every persisted artifact.

Citations

FOUNDATIONAL / METHODS
  1. [1]Burkhardt C, et al. From fragmented records to living evidence: health system-governed, artificial intelligence-driven, continuous learning. JAMIA Open. 2026.
  2. [2]Dahlen AD, et al. Evaluating the generalizability of commercial healthcare claims data. American Journal of Epidemiology. 2025.
  3. [3]Truveta. Truveta: Health system-governed real world data platform.
REPORTING & GUIDANCE
  1. [4]Snow TS, et al. Comparison of Population Characteristics in Real-World Clinical Oncology Databases in the US. medRxiv. 2023.