Optum Clinformatics Data Mart
A large de-identified US claims database from UnitedHealth Group spanning commercial, Medicare Advantage, and Medicaid lines - tens of millions of members with medical, pharmacy, enrollment, and cost records linkable at the person level, plus an optional structured outpatient laboratory feed.
On this page
Optum's Clinformatics Data Mart is a workhorse US claims database for pharmacoepidemiology and health economics: longitudinal person-level claims with generous continuous-enrollment windows, direct paid-amount cost fields, strong Medicare Advantage representation, and an attachable lab feed that partially closes the classic claims gap of no biomarker confirmation. Its limits are claims limits: clinical detail only as good as billing, seniors dominated by MA utilization-management artifacts, and no visibility outside covered plans.
Optum's Clinformatics Data Mart
is a de-identified longitudinal US claims database covering commercial, Medicare Advantage, and Medicaid lines - roughly 60-70 million unique members historically with 15-20 million in any year. Medical claims (ICD-10-CM/CPT/HCPCS/revenue codes), pharmacy dispensings (NDC), eligibility spans, and paid amounts link at the person level; a structured outpatient laboratory feed attaches labs to a subset of members.
Why it matters for RWE
Long continuous-enrollment windows support deep baseline lookback; cost fields support economic endpoints without external price benchmarks; MA coverage reaches seniors that FFS-only sources miss. The lab feed enables biomarker-confirmed intermediate outcomes (HbA1c, LDL-C, creatinine) inside a claims architecture.
Operational characteristics
- Enrollment-driven denominators: person-time derives from eligibility spans; distinguish medical-only vs medical+pharmacy enrollment.
- MA dominance among seniors: prior authorization and network effects shift utilization vs FFS Medicare; do not pool casually.
- Cost detail: line-level paid amounts enable cost-of-care and budget-impact analyses.
- Access: licensed flat-file extracts analyzed in researcher environments.
Common pitfalls
- Claims absence is not clinical absence - unmeasured, not normal, where the lab feed does not reach.
- Lab-feed membership is non-random; quantify selection before generalizing biomarker endpoints.
- Employer-plan churn creates artificial turnover unrelated to health status; apply censoring diagnostics.
- Annual refreshes re-number members; lock dataset versions and document cut dates.
Pros, cons, and trade-offs
- vs MarketScan: integrated lab feed and single-payer consistency vs broader multi-employer pooling.
- vs Truveta-style EHR: complete within-plan capture and continuity vs clinical depth from notes/vitals.
- Trade-off: MA-heavy senior reach inherits managed-care artifacts absent from FFS.
When NOT to use
Outcomes needing nuance beyond codes when labs do not cover them; uninsured or non-covered-service questions; sole regulatory evidence without clinical corroboration.
Decision diagram
flowchart LR UHG[UnitedHealth ecosystem] --> CM[Clinformatics Data Mart] UHG --> LF[Outpatient lab feed] --> LK[Person-level linkage] --> CM CM --> A[Cohort build - enrollment-gated] A --> O[Safety / effectiveness / cost RWE]
Worked example
Scenario
Estimate 1-year LDL-C goal attainment after initiating evolocumab using claims plus the attached lab feed.
Dataset
Lab-confirmed lipid endpoint among claims-defined initiators.
| cohort | n_with_baseline_lab | pct_goal_lt70 | median_followup_d |
|---|---|---|---|
| evolocumab - 1841 - 0.42 - 540 | |||
| ezetimibe_addon - 2207 - 0.21 - 512 |
Steps
Result
42 percent of evaluable evolocumab initiators reached LDL-C under 70 mg/dL vs 21 percent on ezetimibe add-on; lab-feed members skewed urban-commercial, documented as selection limitation.
Trade-offs
Runnable example
Cohort definition with continuous-enrollment gating and lab-based endpoint.
\
import pandas as pd
def enroll_gate(elig, pre_days=365):
idx = elig["index_date"]
return elig[(elig.enroll_start <= idx - pd.Timedelta(days=pre_days)) &
(elig.enroll_end >= idx + pd.Timedelta(days=30))]
def ldl_attainment(labs, cohort, goal=70.0):
m = labs.merge(cohort[["member_id","index_date"]], on="member_id")
m = m[(m.panel=="LDLC") &
(m.result_date >= m.index_date + pd.Timedelta(days=90)) &
(m.result_date <= m.index_date + pd.Timedelta(days=365))]
best = m.groupby("member_id").result_value.min().rename("min_ldlc")
out = cohort.merge(best, left_on="member_id", right_index=True, how="left")
out["attained"] = (out.min_ldlc < goal).astype("Int64")
return out
R/dplyr equivalent.
\
library(dplyr)
enroll_gate <- function(elig, pre_days = 365) {
elig %>% filter(enroll_start <= index_date - days(pre_days),
enroll_end >= index_date + days(30))
}
ldl_attainment <- function(labs, cohort, goal = 70) {
labs %>% inner_join(cohort %>% select(member_id, index_date), by = "member_id") %>%
filter(panel == "LDLC", result_date >= index_date + days(90),
result_date <= index_date + days(365)) %>%
group_by(member_id) %>% summarise(min_ldlc = min(result_value)) %>%
right_join(cohort, by = "member_id") %>%
mutate(attained = as.integer(min_ldlc < goal))
}
SAS PROC SQL version.
\
proc sql;
create table gated as
select e.*, i.index_date
from elig e join idx i on e.member_id = i.member_id
where e.enroll_start <= i.index_date - 365
and e.enroll_end >= i.index_date + 30;
create table attain as
select g.member_id,
min(l.result_value) as min_ldlc,
(calculated min_ldlc < 70) as attained
from gated g left join labs l
on l.member_id = g.member_id and l.panel='LDLC'
and l.result_date between g.index_date+90 and g.index_date+365
group by g.member_id;
quit;
Citations
- [1]Dahlen AD, et al. Benchmarking commercial healthcare claims data. medRxiv. 2024.
- [2]Dahlen AD, et al. Evaluating the generalizability of commercial healthcare claims data. American Journal of Epidemiology. 2025.
- [3]Optum. Optum products: de-identified data and analytics (Clinformatics).