Subgroup Analysis and Heterogeneity of Treatment Effect
The practice of estimating treatment effects separately within patient subgroups — defined by age, sex, comorbidity, genomic marker, or other pre-specified baseline characteristics — to test whether the effect is larger or smaller in certain groups (heterogeneity of treatment effect, or HTE); the principal pathologies are testing within-subgroup p-values instead of a formal interaction test (the cardinal sin), multiplicity inflation across many subgroups, insufficient power for interaction (interaction tests need roughly four times the sample of a main-effect test), and scale dependence (an effect can appear heterogeneous on one scale and homogeneous on another); credibility is assessed with the ICEMAN framework, while the PATH statement and causal-forest methods offer modern risk-based and data-adaptive alternatives.
On this page
A subgroup analysis asks whether a drug works better or worse in specific groups of patients, such as men versus women or younger versus older patients. The most common mistake is concluding the drug "works in group A but not group B" just because the result is statistically significant in one group and not the other — that comparison is invalid and requires a different statistical test called an interaction test. These interaction tests need roughly four times as many patients as the main analysis to have enough statistical power, meaning most published subgroup findings lack the data needed to confirm genuine differences. Researchers have developed credibility checklists (ICEMAN) and risk-based alternatives (PATH statement) to distinguish reliable subgroup findings from statistical noise.
What subgroup analysis is and what question it actually answers
A subgroup analysis estimates the treatment effect within a subset of the study population defined by a baseline characteristic — sex, age group, geographic region, baseline biomarker level, comorbidity, genomic marker, or line of therapy. The core question is whether the effect is heterogeneous across the groups, meaning it is meaningfully larger in one group than another. Subgroup analyses are among the most read and most misread sections of any clinical study report.
Payers and clinicians naturally want to know "does this drug work for patients like mine?" RWE teams are asked to reproduce and extend trial subgroup findings in much larger databases. The result is an enormous literature of subgroup analyses, most of which are statistically underpowered to detect the heterogeneity they claim to report.
The key precondition for any valid subgroup claim is that the subgroup variable and the expected direction of modification are defined in the protocol or statistical analysis plan before the database is queried.
The cardinal sin: within-subgroup p-values instead of interaction tests
The most pervasive and consequential error in subgroup analysis is the following pattern:
- compute a treatment effect estimate and p-value within subgroup A (e.g., men);
- compute a treatment effect estimate and p-value within subgroup B (e.g., women);
- observe that the result is "statistically significant in men (p = 0.01) but not in women (p = 0.16)"; and
- conclude that the effect is heterogeneous.
This reasoning is invalid. The two p-values answer two different null hypotheses in two different populations — they do not test whether the effects in men and women differ from each other. Wang et al. (2007) documented that the majority of published subgroup analyses in NEJM trials relied on exactly this invalid inference.
A formal test of heterogeneity requires an interaction test: fit a model with the treatment variable, the subgroup variable, and their product (the interaction term), and evaluate the coefficient on the product term. The ratio of the subgroup-specific treatment effects (on the ratio scale) or the difference in effects (on the difference scale) is the heterogeneity estimate; its confidence interval and p-value are the only valid measures of whether the effects differ.
The worked example below illustrates this exactly: two subgroup-specific effects that are "significant in men only" but whose interaction estimate (ratio of risk ratios = 9/7, approximately 1.29) carries a wide confidence interval crossing 1.0, providing no evidence of genuine heterogeneity.
Multiplicity: fishing across many subgroups inflates false discoveries
A study that tests 10 pre-specified subgroups will, under the global null, produce at least one false-positive finding roughly 40% of the time at alpha = 0.05 per test. When subgroups are selected after viewing the data, the false-positive rate is far higher and cannot be calculated.
The standard remedies are:
- pre-specification of subgroups in the protocol or statistical analysis plan before database lock
- a pre-specified hierarchy (one primary subgroup, the rest exploratory)
- and multiplicity- adjusted p-values (Bonferroni, Holm) or reporting of all subgroup results with honest labeling of their exploratory status.
The concept on multiplicity and multiple comparisons covers adjustment methods in full; that entry and this one are complementary — subgroup analysis is the domain, multiple-comparisons methods are the remedy when many subgroups are tested.
Power: interaction tests need roughly four times the sample size
An interaction test is fundamentally a test of a difference-of-differences: it asks whether (effect in A) minus (effect in B) differs from zero. Detecting a difference between two effect estimates is harder than detecting either estimate individually. The quantitative rule of thumb is that an interaction test requires approximately four times the total sample size of the main-effect test for the same power — which means that most RWE studies and virtually all trials are severely underpowered to detect modest interactions.
A trial designed with 80% power to detect an overall hazard ratio of 0.75 has roughly 20% power to detect a genuine interaction of the same magnitude. The practical consequence is that a null interaction test is almost always uninformative: it cannot be distinguished from genuine absence of heterogeneity and extreme underpowering. Credible subgroup claims must document the interaction power.
Scale dependence: additive versus multiplicative interaction — RERI
Effect modification is scale-dependent: a treatment can show no interaction on the multiplicative (ratio) scale and strong interaction on the additive (risk-difference) scale, or vice versa. When two variables each have a positive main effect, they cannot be additive on both scales simultaneously — a mathematical consequence of the relationship between ratios and differences.
For public-health and HTA decisions — who to prioritize for treatment, absolute benefit, number-needed-to-treat — the additive scale is the relevant one, because absolute risk reductions determine how many patients benefit. The standard additive-interaction summary measure is the RERI (relative excess risk due to interaction): RERI = RR11 minus RR10 minus RR01 plus 1, where the subscripts indicate joint presence of the two factors.
RERI = 0 indicates no additive interaction; RERI greater than 0 indicates positive synergy (supra- additive effect); RERI less than 0 indicates sub-additivity. A multiplicative interaction that is null on the ratio scale is compatible with a non-zero RERI.
The causal-mediation and effect-modification entry covers RERI computation in detail; the core instruction here is that any subgroup claim should report both scales.
Credibility assessment: the ICEMAN criteria
Schandelmaier et al. (2020) developed the ICEMAN tool (Instrument to assess the Credibility of Effect Modification Analyses) for evaluating the trustworthiness of subgroup claims.
ICEMAN asks eight questions covering:
- pre-specification of the subgroup variable and the direction of the expected modification;
- whether the claim is based on a formal interaction test rather than within-subgroup p-values;
- consistency of the finding across related endpoints;
- biological plausibility;
- whether the finding is supported by external evidence from prior trials or mechanistic studies;
- the number of subgroup tests performed, providing the multiple-testing context;
- sample sizes within subgroups relative to the power needed for the interaction test; and
- whether the finding is pre-specified and confirmatory or post-hoc and exploratory.
A subgroup claim with low ICEMAN credibility — post-hoc, no formal interaction test, small n, implausible mechanism — should not drive clinical or reimbursement decisions. ICEMAN is the current standard for peer review and HTA evaluation of subgroup claims.
Modern HTE: risk-stratified subgroup analysis (PATH statement)
Kent et al. (2020) proposed a principled alternative to single-variable subgroup analysis in the Predictive Approaches to Treatment effect Heterogeneity (PATH) Statement. The core insight is that treatment effects in observational data and trials tend to vary more with a patient's overall baseline risk than with any single covariate. A patient at high baseline risk of the outcome has more absolute room to benefit even if the relative risk reduction is uniform across the spectrum.
The PATH approach groups patients by a validated baseline risk score — preferably derived from control-arm data only, to avoid confounding the risk score with treatment assignment — computes the treatment effect within each risk quintile or tertile, and asks whether the absolute risk reduction is larger in the higher-risk groups.
This approach often reveals more stable and interpretable heterogeneity than single-covariate subgroups, is less vulnerable to multiple-testing inflation (one continuous risk dimension rather than many binary splits), and aligns directly with clinical decision-making because it identifies patients with the most to gain from treatment.
It is now the recommended primary approach for HTE analysis in comparative effectiveness research.
Causal forests and ML-based HTE
Causal forests (Wager and Athey 2018, implemented in the R package grf) and related meta-learners (X-learner, R-learner) estimate the conditional average treatment effect (CATE) as a flexible function of many baseline covariates simultaneously, without pre-specifying which covariates modify the effect. They use sample-splitting and honesty constraints to provide valid confidence intervals for individual-level CATE estimates. In large, high-dimensional RWE datasets they can reveal heterogeneity patterns that parametric interaction terms would miss.
Three warnings are essential:
- overfitting is the dominant risk without careful train-test splits and calibration;
- a high estimated CATE in a region of covariate space does not identify the causal modifier — it describes correlation with the CATE surface, which may simply reflect baseline risk variation rather than biological modification;
- causal forests inherit all confounding from the underlying data, so they estimate a confounded CATE unless the outcome model is properly adjusted.
Use ML-HTE for hypothesis generation and exploration, not for confirmatory claims without pre-registration and independent replication in a separate dataset.
RWE-specific: confounding is inherited differentially across subgroups
In confounded observational data, subgroups do not inherit the same confounding structure as the overall cohort. Propensity-score or covariate-adjustment methods calibrated on the whole population may not produce adequate confounding control within small subgroups, even when the main analysis is well-adjusted.
Within-subgroup covariate distributions can differ substantially from the pooled distribution, so a pooled propensity score does not guarantee exchangeability within strata. A subgroup that is, for example, defined by age below 65 may have very different prescribing patterns and indication severity compared to the full cohort, creating differential confounding that the main-analysis propensity score was not designed to remove.
The practical requirement is to assess balance within each primary subgroup separately — using standardized mean differences on baseline covariates — before reporting subgroup-specific effect estimates. When balance is poor within a subgroup, fit a separate propensity model within the subgroup or include explicit subgroup-by- confounder interaction terms in the outcome model.
Pros, cons, and trade-offs
Pre-specified, interaction-tested subgroup analysis:
- Pros: can identify populations who derive greater or lesser benefit, informing labeling, clinical guidelines, and HTA reimbursement restrictions; supports individualized treatment decisions; required by regulators when safety or efficacy may differ across demographic groups; ICEMAN-credible findings with formal interaction tests and pre-specification are defensible at regulatory review.
- Cons: severely underpowered for interaction in most studies; multiplicity inflation without correction manufactures false findings; scale dependence means a null multiplicative interaction is compatible with a clinically important additive interaction; subgroup variable definitions in RWE (especially in claims) are often proxies for the true clinical variable, introducing misclassification; confounding within subgroups may differ from confounding in the full cohort.
vs reporting only the main (pooled) effect: - The pooled treatment effect is more precisely estimated and more credible; a spurious subgroup finding can damage scientific credibility and mislead practice. Use the pooled result as the primary claim; subgroups are clearly exploratory or hypothesis-generating unless pre-specified with adequate power.
vs causal-ML HTE (causal forests, meta-learners): - Parametric subgroup analysis with pre-specified interaction terms is simple, transparent, and directly testable; ML-HTE provides richer, high-dimensional heterogeneity estimates but is less interpretable and more vulnerable to overfitting.
vs risk-stratified HTE (PATH statement): - Risk-stratified analysis uses one pre-built risk score instead of many binary splits, reducing multiple-testing inflation and aligning with clinical targeting; it is the preferred approach when no single covariate has strong prior evidence for interaction and the goal is identifying who benefits most absolutely.
When NOT to use
- Do not report "significant in subgroup A but not B" as evidence of heterogeneity. This is the cardinal sin. The only valid evidence is the interaction estimate with its own confidence interval and p-value. The astrological-sign subgroup result in the ISIS-2 aspirin trial is the canonical cautionary example of what multiple within-group tests without a formal interaction test produce.
- Do not mine subgroups after observing the data. Data-driven subgroup discovery without pre-specification and multiplicity correction reliably generates false positives. Post-hoc subgroup findings require explicit labeling as exploratory and independent replication before influencing practice or coverage.
- Do not interpret a non-significant interaction test as confirmation of no heterogeneity. Most subgroup analyses are severely underpowered for interaction. A null interaction test result in an underpowered study is uninformative, not evidence of homogeneity.
- Do not report only the multiplicative scale. A non-significant ratio-of-ratios is compatible with a clinically important additive interaction. Always report the RERI and within-subgroup absolute risk differences alongside the multiplicative interaction estimate.
- Do not apply a pooled propensity score to confounded RWE subgroup estimates without verifying within-subgroup balance. Pooled-cohort adjustment does not guarantee confounding control within strata where covariate distributions deviate from the pooled average.
Interpreting the output
In the worked example below, the drug produces a risk ratio of 0.70 in men and 0.90 in women. The within-sex p-values are approximately 0.01 (men, significant) and 0.16 (women, not significant). The interaction estimate — the ratio of risk ratios — equals 0.90/0.70 = 9/7 (approximately 1.29), with a wide 95% confidence interval that easily crosses 1.0.
(1) Formal interpretation. The ratio of subgroup-specific risk ratios is 0.90/0.70 (approximately 1.29), meaning the effect in women is approximately 29% weaker than in men on the multiplicative scale. However, the 95% confidence interval for this ratio (approximately 0.70 to 2.30) includes 1.0, the null value for multiplicative interaction. The interaction p-value exceeds 0.20.
The study is severely underpowered for the interaction test (men n = 1000 total, women n = 400 total, far below the roughly four-fold excess needed relative to the main-effect sample size to detect an interaction of this magnitude with 80% power). There is no statistical evidence that the treatment effect differs by sex, despite the pattern of significant-in-men-only when within-subgroup p-values are compared.
(2) Practical interpretation. The "significant in men, not in women" finding is misleading: it arises from the smaller women sample (400 total vs 1000 total for men) and from different within-stratum baseline risks rather than from a detected biological difference in drug mechanism. A payer or clinician should not restrict treatment to men based on this pattern. The correct scientific conclusion is that evidence for sex-based HTE is absent — the interaction test is null — and a specifically powered follow-up study is warranted before any differential coverage decision.
Decision diagram
flowchart TD
Q[Subgroup finding in the data] --> Pre{Pre-specified in\nprotocol or SAP?}
Pre -->|No — post-hoc| PostHoc[Label as exploratory;\napply multiplicity adjustment;\nICEMAN credibility score]
Pre -->|Yes| Int{Based on a formal\ninteraction test?}
Int -->|No — within-group p-values only| Cardinal["CARDINAL SIN\nWithin-subgroup p-values do NOT\ntest whether effects differ;\nalways use the interaction term"]:::bad
Int -->|Yes| Power{Study powered\nfor the interaction?\n(need ~4x main n)}
Power -->|No — underpowered| Uninf["NULL INTERACTION IS UNINFORMATIVE\nLabel exploratory; report CI;\ndo not conclude homogeneity"]:::warn
Power -->|Yes| Scale{Report BOTH scales?}
Scale -->|Multiplicative only| MissAdd["Missing the additive scale —\ncompute RERI and within-subgroup\nabsolute risk differences"]:::warn
Scale -->|Both multiplicative and additive| ICEMAN{ICEMAN credibility\nassessment}
ICEMAN -->|Low credibility| Replicate[Hypothesis-generating;\nrequires replication]
ICEMAN -->|High credibility| Confirm[Potentially confirmatory;\ndocument and communicate\nuncertainty honestly]:::good
classDef bad fill:#F6F6F5,stroke:#9E3B14;
classDef warn fill:#fef9c3,stroke:#ca8a04;
classDef good fill:#dcfce7,stroke:#15803d;Worked example
Scenario
A health outcomes researcher analyzes a 1400-patient retrospective cohort study comparing a new cardiovascular drug versus an active comparator on the one-year risk of a major adverse cardiovascular event (MACE). The team pre-specified sex as a subgroup. The drug is statistically significant in men (p = 0.01) but not in women (p = 0.16). The PI wants to report "the drug works in men but not in women." Before making this claim, the analyst applies the correct interaction test to see whether there is genuine evidence of effect modification by sex.
Dataset
Event counts from the 1400-patient cohort stratified by sex and treatment arm. Event rates (risk) equal events divided by patients. The drug is significant in men only when within-subgroup p-values are compared, but the interaction test will determine whether the two rates differ significantly from each other.
| subgroup | arm | n_patients | n_events | event_rate |
|---|---|---|---|---|
| men | drug | 500 | 105 | 0.21 |
| men | comparator | 500 | 150 | 0.3 |
| women | drug | 200 | 54 | 0.27 |
| women | comparator | 200 | 60 | 0.3 |
Steps
Result
Risk ratio in men = 105/150 = 0.70 (95% CI approximately 0.54 to 0.91, p ≈ 0.01). Risk ratio in women = 54/60 = 0.90 (95% CI approximately 0.65 to 1.24, p ≈ 0.16). Ratio of risk ratios (interaction estimate) = 0.90/0.70 = 9/7 (approximately 1.29), 95% CI approximately 0.70 to 2.30, interaction p > 0.20. No demonstrated heterogeneity of treatment effect by sex. The "significant in men only" pattern reflects lower sample size and lower baseline risk in the women stratum, not a detected biological difference in drug response.
Trade-offs
Runnable example
Subgroup interaction test and stratified risk ratios using statsmodels Poisson regression (robust SE), plus additive RERI on the risk scale. Required input DataFrame `df` (one row per patient, cohort fully built): arm : 1 = study drug, 0 = active comparator (assigned at index_date) sex : 1 = men, 0 = women (binary...
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
import statsmodels.api as sm
# ── 1. Formal interaction test (multiplicative scale) ──────────────────────────
# Poisson GLM with robust HC3 standard errors; coefficient on arm:sex is the
# log-ratio-of-rate-ratios (log of the multiplicative interaction estimate).
m_int = smf.glm(
"outcome ~ arm * sex + age + cci",
data=df,
family=sm.families.Poisson()
).fit(cov_type='HC3')
b = m_int.params
ci = m_int.conf_int()
rr_interaction = np.exp(b["arm:sex"])
ci_lo = np.exp(ci.loc["arm:sex", 0])
ci_hi = np.exp(ci.loc["arm:sex", 1])
p_int = m_int.pvalues["arm:sex"]
print(f"Ratio of RRs (women / men): {rr_interaction:.3f}")
print(f"95% CI: {ci_lo:.3f} to {ci_hi:.3f}")
print(f"Interaction p-value: {p_int:.4f}")
print("IMPORTANT: a non-significant interaction result does NOT confirm homogeneity")
print(" in an underpowered study — it is uninformative, not evidence of no HTE.")
# ── 2. Stratified estimates (sex-specific RR with confounders) ──────────────────
# These are useful AFTER establishing whether the interaction test is credible;
# do NOT compare these within-group p-values to assess heterogeneity.
for sex_label, sex_val in [("women (sex=0)", 0), ("men (sex=1)", 1)]:
sub = df[df["sex"] == sex_val].copy()
m_sub = smf.glm(
"outcome ~ arm + age + cci",
data=sub,
family=sm.families.Poisson()
).fit(cov_type='HC3')
rr = np.exp(m_sub.params["arm"])
lo, hi = np.exp(m_sub.conf_int().loc["arm"])
p = m_sub.pvalues["arm"]
print(f"RR in {sex_label}: {rr:.2f} (95% CI {lo:.2f}–{hi:.2f}), p={p:.4f}")
# ── 3. Additive interaction (RERI) from the Poisson interaction model ──────────
# RERI = RR11 - RR10 - RR01 + 1 (Knol & VanderWeele 2012 formula).
# RR10 = effect of arm among women (sex=0); RR01 = effect of sex among unexposed (arm=0).
RR10 = np.exp(b["arm"]) # drug effect in reference sex (women)
RR01 = np.exp(b["sex"]) # sex effect when arm=0
RR11 = np.exp(b["arm"] + b["sex"] + b["arm:sex"]) # joint effect (men on drug)
reri = RR11 - RR10 - RR01 + 1
print(f"\nRERI (additive interaction) = {reri:.3f}")
print(" RERI=0: no additive interaction; RERI>0: supra-additive; RERI<0: sub-additive")
print(" Bootstrap RERI CI recommended for valid inference (delta-method approximation shown here).")Subgroup interaction test, stratified risk ratios, RERI, and a causal-forest HTE sketch using statsmodels-equivalent tools in R. Input data.frame `df` mirrors the Python version (arm, sex, outcome, age, cci). The sandwich::vcovHC call provides robust standard errors for the Poisson model.
library(sandwich)
library(lmtest)
library(grf)
## ── 1. Formal multiplicative interaction test ──────────────────────────────────
fit_int <- glm(outcome ~ arm * sex + age + cci, family = poisson, data = df)
# Robust standard errors via HC3 sandwich estimator
ct <- coeftest(fit_int, vcov = vcovHC(fit_int, type = "HC3"))
print(ct)
# Ratio of risk ratios (the interaction estimate): exponentiate the arm:sex coefficient
rr_int <- exp(coef(fit_int)["arm:sex"])
ci_int <- exp(confint.default(fit_int, vcov. = vcovHC(fit_int, type = "HC3"))["arm:sex",])
p_int <- ct["arm:sex", "Pr(>|z|)"]
cat(sprintf("Ratio of RRs: %.3f (95%% CI %.3f to %.3f), p = %.4f\n",
rr_int, ci_int[1], ci_int[2], p_int))
cat("A non-significant p does NOT confirm homogeneity in an underpowered study.\n")
## ── 2. Stratified effect estimates ────────────────────────────────────────────
## Compute sex-specific RRs from the same Poisson model for each stratum.
## Do NOT compare within-stratum p-values to judge heterogeneity.
for (sx in c(0, 1)) {
sub <- subset(df, sex == sx)
m <- glm(outcome ~ arm + age + cci, family = poisson, data = sub)
rr <- exp(coef(m)["arm"])
ci <- exp(confint.default(m, vcov. = vcovHC(m, type = "HC3"))["arm",])
lab <- if (sx == 0) "women" else "men"
cat(sprintf("RR in %s: %.2f (95%% CI %.2f to %.2f)\n", lab, rr, ci[1], ci[2]))
}
## ── 3. Additive interaction (RERI) ────────────────────────────────────────────
b <- coef(fit_int)
RR10 <- exp(b["arm"]) # drug effect in women (reference)
RR01 <- exp(b["sex"]) # sex effect when unexposed
RR11 <- exp(b["arm"] + b["sex"] + b["arm:sex"]) # joint effect (men on drug)
reri <- RR11 - RR10 - RR01 + 1
cat(sprintf("RERI = %.3f (0 = no additive interaction; positive = supra-additive)\n", reri))
## ── 4. Causal forest for data-adaptive HTE (HYPOTHESIS-GENERATING ONLY) ──────
## Replace X columns with your full baseline covariate matrix.
## WARNING: results describe covariate-CATE correlation, not confirmed causal modifiers.
X <- as.matrix(df[, c("sex", "age", "cci")]) # replace with your full covariate matrix
set.seed(2024)
cf <- causal_forest(
X = X,
Y = df$outcome,
W = df$arm,
num.trees = 2000
)
## Is there detectable heterogeneity at all?
print(test_calibration(cf)) # p < 0.05 suggests heterogeneity exists somewhere
## Is sex associated with the estimated CATE surface?
blp <- best_linear_projection(cf, A = df[, "sex", drop = FALSE])
print(blp) # a significant coefficient means sex correlates with estimated CATE
cat("NOTE: 'best_linear_projection' significance does NOT certify sex as a causal modifier.\n")
cat("This output is hypothesis-generating and requires independent replication.\n")Subgroup interaction test and stratum-specific hazard ratios in PROC PHREG, plus RERI from PROC GENMOD (log-binomial). Required input dataset work.cohort (one row per patient; cohort fully built before this block): arm : 1 = study drug, 0 = active comparator sex : 1 = men, 0 = women (binary subgroup;
/* ── 1. PROC PHREG: interaction test + stratum-specific HRs ────────────────── */
proc phreg data=work.cohort;
class arm(ref='0') sex(ref='0') / param=ref;
model time_*event(0) = arm sex arm*sex age cci / ties=efron rl;
/* Coefficient on arm*sex is the log-ratio-of-hazard-ratios (the interaction estimate). */
/* IMPORTANT: a non-significant arm*sex p-value does NOT confirm homogeneity if the study */
/* is underpowered for the interaction test (interaction tests need ~4x the main n). */
/* Stratum-specific hazard ratios (arm effect within each sex level): */
hazardratio arm / at(sex=0) diff=ref; /* HR of arm=1 vs arm=0 among women (sex=0) */
hazardratio arm / at(sex=1) diff=ref; /* HR of arm=1 vs arm=0 among men (sex=1) */
/* These stratum-specific HRs are NOT an interaction test — they are supplementary. */
/* Compare the ratio of these two HRs to the arm*sex interaction coefficient for the */
/* multiplicative HTE estimate; both must be interpreted together. */
run;
/* ── 2. PROC GENMOD: additive interaction (RERI) via log-binomial ──────────── */
/* RERI = RR11 - RR10 - RR01 + 1 (Knol & VanderWeele 2012). */
/* Log-binomial gives risk ratios so that ESTIMATE builds RERI from coefficients. */
proc genmod data=work.cohort;
class arm(ref='0') sex(ref='0') / param=ref;
model event(event='1') = arm sex arm*sex age cci / dist=binomial link=log;
/* RR10 = exp(arm coefficient); RR01 = exp(sex coefficient) */
/* RR11 = exp(arm + sex + arm*sex) */
/* RERI = RR11 - RR10 - RR01 + 1 — assembled from three ESTIMATEs below. */
estimate 'log RR10' arm 1 sex 0 arm*sex 0;
estimate 'log RR01' arm 0 sex 1 arm*sex 0;
estimate 'log RR11' arm 1 sex 1 arm*sex 1;
/* Exponentiate each estimate (exp column in output) to get RR10, RR01, RR11; */
/* then compute RERI = RR11 - RR10 - RR01 + 1 in a subsequent DATA step. */
run;
/* ── 3. Within-stratum balance check (required in RWE before reporting subgroup HRs) ── */
/* Assess standardized mean differences within each subgroup stratum separately. */
/* A pooled propensity score does NOT guarantee balance within subgroups. */
proc means data=work.cohort mean std;
by sex arm;
var age cci; /* add all baseline covariates here */
output out=work.balance_check mean=m_ std=sd_ / autoname;
run;
/* Compute SMD = (m_arm1 - m_arm0) / pooled_SD within each sex level; */
/* flag any SMD > 0.10 as requiring subgroup-specific PS re-estimation or */
/* additional outcome-model adjustment for that covariate within the subgroup. */Citations
- [1]Wang R, Lagakos SW, Ware JH, Hunter DJ, Drazen JM. Statistics in Medicine — Reporting of Subgroup Analyses in Clinical Trials. New England Journal of Medicine. 2007;357(21):2189-2194.
- [2]Kent DM, Steyerberg E, van Klaveren D. Personalized evidence based medicine: predictive approaches to heterogeneous treatment effects. BMJ. 2018;363:k4245. [PATH Statement] Kent DM, van Klaveren D, Paulus JK, et al. The Predictive Approaches to Treatment effect Heterogeneity (PATH) Statement. Annals of Internal Medicine. 2020;172(1):35-45.