Health Equity-Stratified Analysis in RWE
The deliberate design and analysis of real-world evidence studies to measure how treatments, outcomes, and care processes differ across socially defined populations - race, ethnicity, socioeconomic status, geography, language, insurance type - so that evidence products quantify rather than average away inequities.
On this page
Most RWE studies adjust for demographics as confounders and move on. Equity-stratified analysis treats population subgroups as findings in their own right: reporting effect estimates within racial, ethnic, socioeconomic, and geographic groups, testing interaction honestly, measuring differential access and follow-up, and interpreting differences through structural mechanisms rather than biology by default. Regulators (FDORA action plans), HTA bodies (equity lenses), and journals (reporting checklists) increasingly expect it.
Health equity-stratified analysis
designs RWE so that differences across socially defined groups are measured and interpreted, not merely adjusted away. Where conventional practice treats race/ethnicity and socioeconomic position as nuisance covariates, equity analysis asks: does the treatment work equally well, is it equally accessible, and is the evidence itself equally representative — across these groups?
Why it matters for RWE
Three forces make this a first-class design concern rather than an optional sensitivity analysis. First, regulation: FDORA (2022) directed FDA to publish diversity action plans for trials, and external-control RWE inherits the expectation that its populations reflect intended treatment populations. Second, validity: differential outcome ascertainment (fewer labs in marginalized groups), differential censoring (insurance churn concentrated in low-income enrollees), and differential treatment allocation all create effect-measure modification and bias that marginal estimates conceal. Third, policy use: HTA bodies and payers making coverage decisions need to know whether evidence generalizes to the populations they cover.
Design elements
- Pre-specified strata: define equity-relevant subgroups (race/ethnicity, SES measures like area deprivation index or dual-eligibility, geography/rurality, language, insurance type) in the protocol with power justification — not post hoc fishing.
- Interaction testing with humility: test effect-measure modification on additive and multiplicative scales; additive-scale differences matter most for policy because they express absolute benefit gaps.
- Differential-bias diagnostics: evaluate missingness, censoring, and misclassification by subgroup; a validated algorithm overall can be badly calibrated within groups (e.g., pulse oximetry, eGFR race corrections).
- Representativeness accounting: compare study population composition against the target population and report who was excluded and why.
Measurement cautions
- Race is social, not biological: model it as exposure to structural forces (racism, segregation, access), never as intrinsic risk absent that framing; avoid race-adjusted clinical algorithms without scrutiny.
- SES measurement choice matters: individual income is rarely available; area-level indices (ADI, SVI) introduce ecological misclassification — state which and why.
- Small-stratum instability: rare groups yield unstable estimates; report precision limits honestly instead of collapsing categories silently.
Common pitfalls
- Adjusting for mediators (e.g., adjusting away insurance status when studying access interventions) erases the mechanism of interest.
- Post-hoc subgroup trawling presented as pre-specified equity analysis.
- Treating 'no significant interaction' as absence of inequity when power was inadequate.
- Ignoring differential follow-up: shorter observation in marginalized groups masquerades as lower event rates.
Pros, cons, and trade-offs
- vs pooled marginal estimation: reveals distributional effects and policy-relevant gaps vs simpler models and tighter precision.
- vs dedicated equity studies: embeds equity in existing evidence workflows vs purpose-built cohorts with richer social variables.
- Trade-off: multiple-strata inference multiplies Type I error surface — pre-specification and hierarchical testing discipline are the price of credibility.
When NOT to use
Do not perform underpowered subgroup theater for compliance optics; if the data cannot support stratified inference, say so and scope a fit-for-purpose study. And never present race-stratified associations without a structural interpretation framework.
Decision diagram
flowchart LR P[Protocol] --> SP[Pre-specified equity strata\n+ power + multiplicity] SP --> A[IPTW analysis overall] SP --> B[Differential-bias audit\nascertainment - censoring - missingness by stratum] A --> C[Stratified effects:\nmultiplicative + additive scale] B --> C C --> I[Interpretation through\nstructural mechanisms]
Worked example
Scenario
Evaluate whether a heart-failure SGLT2i effectiveness signal differs by area-deprivation quintile in a claims cohort.
Dataset
Effectiveness by ADI quintile with interaction tests.
| adi_quintile | n | hhosp_or | risk_diff_pp | interaction_p_additive |
|---|---|---|---|---|
| Q1_least_deprived - 18400 - 0.78 - -3.1 - ref | ||||
| Q5_most_deprived - 9100 - 0.84 - -1.8 - 0.04 |
Steps
Result
Treatment reduced hospitalization in all strata but absolute benefit was smaller in Q5 (-1.8 vs -3.1 pp); additive-scale interaction p=0.04, driven partly by differential follow-up length corrected by censoring weights.
Trade-offs
Runnable example
Stratified IPTW effectiveness with additive-scale interaction test skeleton.
\
import pandas as pd
import statsmodels.formula.api as smf
def equity_stratified(df, treatment, outcome, strata="adi_q", ps_covs=None):
# IPTW via logistic PS
import statsmodels.api as sm
X = sm.add_constant(df[ps_covs])
df = df.assign(ps=sm.Logit(treatment, X).fit().predict(X))
df = df.assign(w=df[treatment]/df.ps + (1-df[treatment])/(1-df.ps))
# Outcome model with treatment x strata interaction (additive via GLM identity)
m = smf.glm(f"{outcome} ~ {treatment} * C({strata})", data=df,
family=sm.families.Gaussian(), freq_weights=df.w).fit(cov_type="HC1")
return m # interaction terms give stratum-specific absolute RDs
R version with survey-weighted stratified models and interaction tests.
\
library(survey)
equity_stratified <- function(df, treatment, outcome, strata = "adi_q") {
des <- svydesign(ids = ~1, weights = ~w, data = df)
m <- svyglm(as.formula(paste(outcome, "~", treatment, "*", strata)),
design = des)
# Additive-scale interaction: linear model coefficients directly give RD contrasts
regTermTest(m, strata) # joint interaction p-value
summary(m)$coefficients
}
SAS GENMOD with weighted stratum-interaction model.
\
proc genmod data=coh;
class adi_q(ref='Q1') treatment(ref='0') / param=reference;
weight w;
model y = treatment adi_q treatment*adi_q / dist=normal link=identity;
lsmeans treatment*adi_q / diff oddsratio cl;
/* Interaction LRT */
run;
/* Differential censoring check */
proc lifetest data=coh notable;
time followup_days*dropout(0);
strata adi_q / testcmd=logrank;
run;
Citations
- [1]U.S. Food and Drug Administration. FDORA Diversity Action Plans guidance activities.
- [2]ISPOR/NPC Health Equity initiatives: good practices for incorporating health equity considerations in RWE and HEOR.
- [3]NCI SEER-Medicare: patterns-of-care and disparities research applications.