← Methods repository
CONCEPTINTERMEDIATEPYTHON · R · SASlast reviewed 2026-08-25 · updated 2026-08-25 · 4 citations

Optum Clinformatics Data Mart

A large de-identified US claims database from UnitedHealth Group spanning commercial, Medicare Advantage, and Medicaid lines - tens of millions of members with medical, pharmacy, enrollment, and cost records linkable at the person level, plus an optional structured outpatient laboratory feed.

Data Sourceoptumclinformaticsdata-sourceclaimsmedicare-advantagelab-linkageus-rwd
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

Optum's Clinformatics Data Mart is a workhorse US claims database for pharmacoepidemiology and health economics: longitudinal person-level claims with generous continuous-enrollment windows, direct paid-amount cost fields, strong Medicare Advantage representation, and an attachable lab feed that partially closes the classic claims gap of no biomarker confirmation. Its limits are claims limits: clinical detail only as good as billing, seniors dominated by MA utilization-management artifacts, and no visibility outside covered plans.

When to use it
—Longitudinal drug-safety and cost questions; EHR sources when phenotype needs clinical detail.
—MA-majority senior questions; resdac FFS when fee-for-service fidelity is required.
—Either works for most claims questions; choose by lab need and license.
Watch out for
—No notes/vitals depth; clinical picture limited to codes plus optional labs.
—MA utilization-management artifacts; no automatic Part D comparability.
—Smaller commercial pool than pooled MarketScan extracts.

Optum's Clinformatics Data Mart

is a de-identified longitudinal US claims database covering commercial, Medicare Advantage, and Medicaid lines - roughly 60-70 million unique members historically with 15-20 million in any year. Medical claims (ICD-10-CM/CPT/HCPCS/revenue codes), pharmacy dispensings (NDC), eligibility spans, and paid amounts link at the person level; a structured outpatient laboratory feed attaches labs to a subset of members.

Why it matters for RWE

Long continuous-enrollment windows support deep baseline lookback; cost fields support economic endpoints without external price benchmarks; MA coverage reaches seniors that FFS-only sources miss. The lab feed enables biomarker-confirmed intermediate outcomes (HbA1c, LDL-C, creatinine) inside a claims architecture.

Operational characteristics

  • Enrollment-driven denominators: person-time derives from eligibility spans; distinguish medical-only vs medical+pharmacy enrollment.
  • MA dominance among seniors: prior authorization and network effects shift utilization vs FFS Medicare; do not pool casually.
  • Cost detail: line-level paid amounts enable cost-of-care and budget-impact analyses.
  • Access: licensed flat-file extracts analyzed in researcher environments.

Common pitfalls

  • Claims absence is not clinical absence - unmeasured, not normal, where the lab feed does not reach.
  • Lab-feed membership is non-random; quantify selection before generalizing biomarker endpoints.
  • Employer-plan churn creates artificial turnover unrelated to health status; apply censoring diagnostics.
  • Annual refreshes re-number members; lock dataset versions and document cut dates.

Pros, cons, and trade-offs

  • vs MarketScan: integrated lab feed and single-payer consistency vs broader multi-employer pooling.
  • vs Truveta-style EHR: complete within-plan capture and continuity vs clinical depth from notes/vitals.
  • Trade-off: MA-heavy senior reach inherits managed-care artifacts absent from FFS.

When NOT to use

Outcomes needing nuance beyond codes when labs do not cover them; uninsured or non-covered-service questions; sole regulatory evidence without clinical corroboration.

Decision diagram

flowchart LR
  UHG[UnitedHealth ecosystem] --> CM[Clinformatics Data Mart]
  UHG --> LF[Outpatient lab feed] --> LK[Person-level linkage] --> CM
  CM --> A[Cohort build - enrollment-gated]
  A --> O[Safety / effectiveness / cost RWE]
Optum Clinformatics pipeline: claims core with attachable lab feed and enrollment-gated cohorts.

Worked example

Scenario

Estimate 1-year LDL-C goal attainment after initiating evolocumab using claims plus the attached lab feed.

Dataset

Lab-confirmed lipid endpoint among claims-defined initiators.

cohortn_with_baseline_labpct_goal_lt70median_followup_d
evolocumab - 1841 - 0.42 - 540
ezetimibe_addon - 2207 - 0.21 - 512

Steps

1Define initiators via NDC with 12-month continuous enrollment pre-index.
2Restrict to members present in the lab feed; compare included vs excluded covariates for selection bias.
3Take lowest LDL-C in days 90-365 post-index as attainment value.
4Report linkage-rate diagnostics and rerun excluding labs as claims-only sensitivity.

Result

42 percent of evaluable evolocumab initiators reached LDL-C under 70 mg/dL vs 21 percent on ezetimibe add-on; lab-feed members skewed urban-commercial, documented as selection limitation.

Trade-offs

vs. Truveta style multi system EHR
Pros of this
—Complete within-plan capture including out-of-network fills; multi-year continuity; direct cost fields.
vs. Medicare FFS (resdac)
Pros of this
—Includes MA majority plus younger commercial members in one identifier space.
vs. MarketScan
Pros of this
—Integrated lab feed; single-payer ecosystem consistency.

Runnable example

Cohort definition with continuous-enrollment gating and lab-based endpoint.

requires: pandas
\
import pandas as pd

def enroll_gate(elig, pre_days=365):
    idx = elig["index_date"]
    return elig[(elig.enroll_start <= idx - pd.Timedelta(days=pre_days)) &
                (elig.enroll_end   >= idx + pd.Timedelta(days=30))]

def ldl_attainment(labs, cohort, goal=70.0):
    m = labs.merge(cohort[["member_id","index_date"]], on="member_id")
    m = m[(m.panel=="LDLC") &
          (m.result_date >= m.index_date + pd.Timedelta(days=90)) &
          (m.result_date <= m.index_date + pd.Timedelta(days=365))]
    best = m.groupby("member_id").result_value.min().rename("min_ldlc")
    out = cohort.merge(best, left_on="member_id", right_index=True, how="left")
    out["attained"] = (out.min_ldlc < goal).astype("Int64")
    return out

Citations

FOUNDATIONAL / METHODS
  1. [1]Dahlen AD, et al. Benchmarking commercial healthcare claims data. medRxiv. 2024.
  2. [2]Dahlen AD, et al. Evaluating the generalizability of commercial healthcare claims data. American Journal of Epidemiology. 2025.
  3. [3]Optum. Optum products: de-identified data and analytics (Clinformatics).
APPLIED EXAMPLES
  1. [4]Strom BL. Data validity issues in using claims data. Pharmacoepidemiology and Drug Safety. 2001.