← Methods repository
CONCEPTINTERMEDIATEPYTHON · R · SASlast reviewed 2026-08-25 · updated 2026-08-25 · 4 citations

MarketScan (Truven/Merative)

A family of de-identified US commercial claims databases (Commercial Claims and Encounters, Medicare Supplemental, Medicaid Multi-State) pooling employer and insurer records for tens of millions of enrollees annually - the most widely used US claims source in pharmacoepidemiology and health-services research.

Data Sourcemarketscantrivenmerativedata-sourceclaimscommercial-insuranceus-rwd
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

MarketScan pools de-identified medical, pharmacy, and enrollment records from hundreds of US employers and insurers into person-level files spanning commercially insured lives, Medicare-eligible supplementally insured retirees, and Medicaid beneficiaries across selected states. Its decades-long history, large pool, and breadth of published validation studies make it the default choice for many drug-safety and comparative-effectiveness questions in working-age populations.

When to use it
—Rare events, long follow-up, benchmarking against prior MarketScan literature.
—Under-65 questions; resdac for elderly FFS-specific policy work.
—Claims-native questions; EHR sources when clinical phenotype precision matters.
Watch out for
—No integrated lab feed; pool composition shifts over time.
—Senior file is supplemental, not FFS; no DME/part-D-level detail parity.
—Zero laboratory/vitals content; severity limited to codes.

MarketScan

(originally Truven Health Analytics, now Merative) is a family of de-identified US claims databases: Commercial Claims and Encounters (employer- and health-plan-sourced commercial insurance), Medicare Supplemental and Coordination of Benefits (retirees with Medicare plus supplemental commercial coverage), and Medicaid Multi-State (selected state programs). Person-level linkage of inpatient, outpatient, and pharmacy claims with enrollment timelines supports cohort construction with defined lookback windows.

Why it matters for RWE

MarketScan has been the dominant US commercial claims source for two decades; most standard pharmacoepidemiologic methods (new-user designs, propensity-score matching, self-controlled designs) were routinely applied to it, producing an unusually rich validation literature. Its scale enables rare-outcome safety studies; its multi-payer pooling reduces single-payer idiosyncrasy.

Operational characteristics

  • Enrollment-driven: continuous-enrollment gating defines cohorts; typical designs require 6-12 months pre-index lookback.
  • Family linkage: enrollment files identify family units, enabling household and pregnancy-partner analyses.
  • No results data: no labs or vitals; biomarker endpoints are impossible without external linkage.
  • Medicare Supplemental caveat: covers Medicare-eligible retirees with supplemental commercial coverage, not FFS Medicare itself; dual coverage complicates claim attribution.

Common pitfalls

  • Employer-group turnover creates non-health-related censoring; treat disenrollment carefully in time-to-event analyses.
  • Pool composition shifts year to year as employers join/leave; secular trends can reflect the pool rather than practice.
  • Out-of-network care may generate claims only via out-of-pocket reimbursement submissions; completeness varies.
  • No clinical results: severity adjustment is limited to coded comorbidities and utilization proxies.

Pros, cons, and trade-offs

  • vs Optum Clinformatics: larger pooled multi-payer population vs integrated labs and single-payer consistency.
  • vs SEER-Medicare: working-age generalizability vs cancer-registry clinical depth in seniors.
  • Trade-off: scale vs depth — MarketScan maximizes sample size and follow-up length but carries pure claims granularity.

When NOT to use

Elderly-focused questions better served by Medicare FFS; biomarker-defined endpoints without external lab linkage; outcomes dominated by care outside covered networks.

Decision diagram

flowchart LR
  E[Employers] --> P[Payers / insurers]
  P --> MS[MarketScan pooling - de-identified]
  MS --> CCAE[Commercial CCAE]
  MS --> MCR[Medicare Supplemental]
  MS --> MCD[Medicaid Multi-State]
  CCAE --> S[Cohort RWE studies]
  MCR --> S
  MCD --> S
MarketScan pooling structure: employer and payer feeds de-identified into three research files.

Worked example

Scenario

New-user cohort study comparing cardiovascular hospitalization risk between two antihypertensive classes in commercially insured adults.

Dataset

New-user design summary statistics.

cohortn_new_usersps_matched_nhhosp_1yr_pctadj_hr
ACEi - 148220 - 96400 - 0.021 - 1.00
ARB - 112870 - 96400 - 0.018 - 0.87

Steps

1Identify new users: index dispensing with 12-month baseline free of either class.
2Trim to continuous enrollment 12m pre / 365d post.
3Propensity-score match on demographics, comorbidities, comedications, utilization.
4Estimate hazard ratios with robust SEs; run negative-control outcome diagnostics.

Result

PS-matched HR 0.87 (95% CI 0.79-0.96) favoring ARBs; negative controls balanced, supporting residual-confounding control.

Trade-offs

vs. Optum Clinformatics
Pros of this
—Larger multi-payer pool; longest historical archive; richest methods-validation literature.
vs. Medicare FFS (resdac)
Pros of this
—Includes under-65 commercial lives unavailable to Medicare-only sources.
vs. EHR based sources
Pros of this
—Complete within-plan capture regardless of provider site; long spans.

Runnable example

New-user cohort construction with PS matching skeleton.

requires: pandas · scikit-learn
\
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import NearestNeighbors

def new_user_cohort(disp, elig):
    first = disp.sort_values(["member_id","rx_date"]).groupby("member_id").first()
    idx = first.reset_index().rename(columns={"rx_date":"index_date","drug_class":"cohort"})
    # 12-month continuous enrollment gate
    ok = []
    for _, r in idx.iterrows():
        e = elig[elig.member_id == r.member_id]
        ok.append(((e.enroll_start <= r.index_date - pd.Timedelta(days=365)) &
                   (e.enroll_end   >= r.index_date + pd.Timedelta(days=365))).any())
    return idx[pd.Series(ok, index=idx.index)]

def ps_match(cohort, covariates):
    X = cohort[covariates]
    ps = LogisticRegression(max_iter=1000).fit(X, cohort.cohort).predict_proba(X)[:,1]
    cohort = cohort.assign(ps=ps)
    treated = cohort[cohort.cohort=="ARB"]; control = cohort[cohort.cohort=="ACEi"]
    nn = NearestNeighbors(n_neighbors=1).fit(control[["ps"]])
    _, ix = nn.kneighbors(treated[["ps"]])
    return treated.merge(control.iloc[ix.flatten()], on=None, how="left",
                         left_index=True, right_index=True, suffixes=("_t","_c"))

Citations

FOUNDATIONAL / METHODS
  1. [1]Schaefer EW, et al. Truven Health Analytics MarketScan Databases for Clinical Research in Colon and Rectal Surgery. Clinics in Colon and Rectal Surgery. 2019.
  2. [2]Dahlen AD, et al. Evaluating the generalizability of commercial healthcare claims data. American Journal of Epidemiology. 2025.
  3. [3]Merative (formerly IBM Watson Health / Truven Health Analytics). MarketScan Research Databases.
REPORTING & GUIDANCE
  1. [4]Sada A, et al. National trends in multimodality therapy for locally advanced gastric cancer. Journal of Surgical Research. 2019.