Truveta
A US multi-health-system real-world data platform combining full-fidelity EHR data from dozens of participating provider organizations with linked claims and mortality data, normalized to a common data model and refreshed continuously — distinctive for retaining source-level clinical detail (vitals, labs, notes-derived variables) that claims-only sources lack.
On this page
Truveta aggregates de-identified electronic health records from a large consortium of US health systems (including providers such as Providence and Trinity Health) into one research-accessible platform. Unlike claims databases it captures clinical richness — lab values, vitals, imaging indications, clinician notes — and unlike single-system EHR extracts it spans many regions and payer types. Data is cleaned to a common model daily, can be linked to insurance claims and Social Security Death Master File mortality, and is queried in an analytic workbench rather than shipped as flat files. Its main limitations are the EHR-native ones: care outside participating systems is invisible, coding follows billing/problem-list incentives, and completeness varies by how each system implements its EHR.
Truveta
is a real-world data platform founded in 2020 that aggregates de-identified, full-fidelity EHR data from a consortium of US health systems (30+ systems spanning thousands of care sites), together with linked insurance claims and mortality (Social Security Death Master File) data. Where legacy claims databases capture billing events and legacy single-system EHR extracts lack generalizability, Truveta's differentiator is multi-system clinical fidelity at national scale: laboratory results with units and reference ranges, vital signs, immunizations, clinician documentation, and structured problem lists — retained close to source fidelity and mapped to a common data model that is re-processed continuously as upstream EHR data updates.
Why it matters for RWE
Truveta occupies the middle of the claims-vs-EHR trade-off: richer than claims for clinical phenotyping (lab-confirmed outcomes, disease severity, pregnancy gestational age), broader than any single health system for generalizability across geography, age, and payer mix. The continuous-refresh architecture supports rapid-cycle evidence (its original public-health use case was near-real-time COVID-19 monitoring) but also introduces versioning discipline requirements: analyses must record the data cut date because point estimates move as back-filled records arrive.
Operational characteristics
- Population: tens of millions of patients, skewed toward systems' catchment areas (strong West/Midwest/South representation depending on participating systems); all ages including pediatric, unlike Medicare-restricted sources.
- Encounter-driven capture: visits generate data; out-of-system care (filling a prescription elsewhere, hospitalization at a non-member system) is invisible unless captured in problem lists or reconciliation fields.
- Linkage: within-platform patient identity resolution across member systems plus linkage to claims and death data; researchers cannot bring their own identifiers (privacy-preserving tokenization inside the platform).
- Access model: cloud workbench analysis (code-in, results-out); patient-level export is generally not part of the standard model — plan analytic pipelines accordingly.
Common pitfalls
- Treating EHR presence as coverage. A patient with sparse encounters has missing data, not absence of events; distinguish 'no visit' from 'no event'.
- Ignoring cut-date versioning. Back-filled labs and late-arriving encounter records change results between cuts; lock and report the data snapshot.
- Outcome ascertainment drift. Lab-based outcomes (e.g., HbA1c-defined control) are strong; diagnosis-code-only outcomes inherit problem-list noise.
- Denominator confusion. The enrolled-ish population is 'people with ≥1 encounter', not a stable insurance denominator — incidence rates need careful person-time definitions.
Pros, cons, and trade-offs
- vs commercial claims (MarketScan, Optum): clinical depth (labs, vitals, notes) vs longitudinal continuity and complete out-of-system capture; claims see every filled script regardless of where care happens, EHR sees only in-system activity.
- vs single health-system EHR: generalizability and scale vs within-system record completeness; single-system data has fewer cross-site gaps for patients who stay local.
- Trade-off: the workbench access model controls privacy leakage but constrains tooling; methods requiring patient-level export or bespoke packages may not be portable.
When NOT to use
Not ideal for long pre-period lookback beyond the platform's accumulation window per system, for outcomes almost exclusively captured out-of-system (e.g., retail pharmacy fills without EHR reconciliation), or for studies needing patient-level data export to local infrastructure under the standard access model.
Decision diagram
flowchart LR HS1[Health System A EHR] --> T HS2[Health System B EHR] --> T HS3[... 30+ systems] --> T CL[Claims feed] --> L[Identity tokenization & linkage] --> T[Truveta common data model\ncontinuous refresh] D[Social Security mortality] --> L T --> W[Cloud workbench\ncode-in / results-out] W --> S[RWE study\ndata-cut stamped]
Worked example
Scenario
Compare time-to-HbA1c<7% after initiating two GLP-1 receptor agonists using Truveta EHR labs, with 1-year follow-up.
Dataset
Lab-based endpoint construction for two treatment cohorts.
| cohort | n_initiators | median_months_to_goal | hazard_ratio_adj |
|---|---|---|---|
| drug_A - 4120 - 7.8 - 1.00 (ref) | |||
| drug_B - 3865 - 9.1 - 0.86 |
Steps
Result
Adjusted HR 0.86 (95% CI 0.80-0.92) favoring drug A on the locked March cut; the June cut shifted the estimate by <2%, supporting robustness.
Trade-offs
Runnable example
Cohort construction and lab-series endpoint building against an extracted Truveta-style EHR frame (person, encounter, lab_result, drug_exposure tables).
import pandas as pd
def hba1c_goal_cohort(person, encounter, lab, drug, index_drugs=("A","B"),
goal=7.0, baseline_lo=7.5, baseline_hi=11.0):
# First qualifying dispensing per person
dx = drug[drug.drug_name.isin(index_drugs)].sort_values(["person_id","drug_date"])
idx = dx.groupby("person_id").first().reset_index()
idx = idx.rename(columns={"drug_date":"index_date","drug_name":"cohort"})
# Baseline window: 180 days pre-index
base = lab.merge(idx[["person_id","index_date"]], on="person_id")
base = base[(base.result_date < base.index_date) &
(base.result_date >= base.index_date - pd.Timedelta(days=180)) &
(base.loinc == "4548-4")]
bval = base.groupby("person_id").result_value.median().rename("baseline_a1c")
idx = idx.merge(bval, on="person_id")
idx = idx[idx.baseline_a1c.between(baseline_lo, baseline_hi)]
# Follow-up: first post-index result below goal (>=30 days post-index)
fu = lab.merge(idx[["person_id","index_date"]], on="person_id")
fu = fu[(fu.loinc=="4548-4") & (fu.result_date >= fu.index_date + pd.Timedelta(days=30))]
first_goal = (fu[fu.result_value < goal]
.groupby("person_id").result_date.min().rename("goal_date"))
return idx.merge(first_goal, left_on="person_id", right_index=True, how="left")
# Person-time and KM estimation proceed with standard survival libraries;
# always attach the data-cut date to every persisted artifact.
Same lab-based endpoint construction using dplyr semantics.
library(dplyr); library(survival)
hba1c_goal_cohort <- function(person, encounter, lab, drug,
goal = 7.0, base_lo = 7.5, base_hi = 11.0) {
idx <- drug %>%
filter(drug_name %in% c("A","B")) %>%
arrange(person_id, drug_date) %>%
group_by(person_id) %>% slice_head(n = 1) %>%
ungroup() %>% rename(index_date = drug_date, cohort = drug_name)
base <- lab %>% inner_join(idx %>% select(person_id, index_date), by = "person_id") %>%
filter(loinc == "4548-4",
result_date >= index_date - 180, result_date < index_date) %>%
group_by(person_id) %>% summarise(baseline_a1c = median(result_value))
idx <- idx %>% inner_join(base, by = "person_id") %>%
filter(between(baseline_a1c, base_lo, base_hi))
fu <- lab %>% inner_join(idx %>% select(person_id, index_date), by = "person_id") %>%
filter(loinc == "4548-4", result_date >= index_date + 30)
goal_dates <- fu %>% filter(result_value < goal) %>%
group_by(person_id) %>% summarise(goal_date = min(result_date))
left_join(idx, goal_dates, by = "person_id")
}
SAS version: baseline-window lab summarization and endpoint flagging with PROC SQL.
/* Baseline HbA1c: median over 180-day pre-index window */
proc sql;
create table base as
select l.person_id, median(l.result_value) as baseline_a1c
from lab l, idx i
where l.person_id = i.person_id
and l.loinc = '4548-4'
and i.baseline_lo <= calculated baseline_a1c <= i.baseline_hi
group by l.person_id;
/* Endpoint: first post-index result below goal */
proc sql;
create table endpoint as
select p.person_id, min(l.result_date) as goal_date format yymmdd10.
from person p
left join lab l
on p.person_id = l.person_id and l.loinc = '4548-4'
and l.result_value < 7.0
and l.result_date >= p.index_date + 30
group by p.person_id;
quit;
Citations
- [1]Burkhardt C, et al. From fragmented records to living evidence: health system-governed, artificial intelligence-driven, continuous learning. JAMIA Open. 2026.
- [2]Dahlen AD, et al. Evaluating the generalizability of commercial healthcare claims data. American Journal of Epidemiology. 2025.
- [3]Truveta. Truveta: Health system-governed real world data platform.