← Methods repository
CONCEPTADVANCEDPYTHON · R · SAS4 citations

Relative and Net Survival

A method for estimating cancer-attributable survival from registry data where cause of death is unreliable or missing: relative survival divides the observed all-cause survival of a cancer cohort by the expected survival of a matched general-population group from population life tables; the result estimates net survival -- the probability of surviving the cancer in a hypothetical world where no other cause of death can occur.

Inferential Statisticsrelative-survivalnet-survivalcancer-registryexcess-hazardpopulation-life-tablespohar-permeage-standardizationICSS-weights
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

Relative survival answers the question "how much does a cancer diagnosis reduce a patient's chances of being alive at five years?" without needing to know the exact cause of death -- a piece of information that is often wrong or missing on death certificates. The method compares how many cancer patients actually survived against how many would have been expected to survive based on their age and sex from general-population life tables; the ratio of those two numbers gives net survival, the survival attributable to the cancer alone. A net survival of 0.50 means cancer patients survived at half the rate of comparable people in the general population, with the gap blamed on the cancer and its treatment rather than background diseases like heart attacks or strokes.

When to use it
Population-based registry studies; any setting where cause-of-death coding is unreliable, missing, or heterogeneous across the comparison groups; cross-registry benchmarking and international comparisons.
Whenever the scientific or policy question is about cancer-attributable survival rather than raw all-cause survival, and whenever comparing across time periods or registries with different background life expectancies.
All new analyses should use Pohar Perme; use Ederer II only when explicitly reproducing or extending older results and when small-sample variance inflation would be material.
Watch out for
Cannot isolate cancer deaths from treatment-toxicity deaths (both are excess mortality); depends on the independence assumption; requires a well-matched life table;
More complex to compute (requires life-table linkage); interpretation requires understanding the net-survival estimand, which is hypothetical by construction.
Slightly higher variance than Ederer II in small samples; cannot directly reproduce older published literature that used Ederer II without re-estimating.

The cancer-registry problem that motivates relative survival

In most cancer registries, the death certificate records a cause of death -- but those entries are notoriously unreliable. A patient with stage-IV colon cancer who dies of myocardial infarction may be coded as a cancer death; a patient dying of uncontrolled metastatic disease may be certified as dying of sepsis. Misclassification is differential by cancer site, age, race/ethnicity, and data period.

In Medicare claims and other administrative sources the problem is worse still: the only cause-of-death signal is the ICD code on the discharge claim or the enrollment death-file entry, not a reviewed death certificate. The practical consequence is that cause-specific survival analysis requires knowing why a patient died, and that knowledge is often unavailable, partial, or systematically wrong in population-based registry data.

Relative survival bypasses the problem entirely: it never asks why a patient died, only whether the rate at which they died exceeded what was expected from the general population of the same age, sex, and calendar year.

The estimand: net survival

Net survival is the survival probability in a conceptual world where the cancer is the only possible cause of death -- all background mortality has been removed by assumption. It is not the observed all-cause survival (which includes deaths from heart disease, stroke, and accidents unrelated to the cancer), and it is not cause-specific survival (which requires reliable cause-of-death coding).

Net survival is the quantity that SEER data products, CONCORD, and EUROCARE comparisons report, and it is the international standard for population-based cancer surveillance. The Pohar Perme estimator (2012) is the nonparametric estimator of net survival that is unbiased under the independence assumption: a patient's expected general-population survival is independent of their actual cancer prognosis given age, sex, and calendar year. This is the current recommendation of the International Agency for Research on Cancer.

The relative survival ratio: observed divided by expected

Relative survival is computed as:

relative_survival = observed_survival / expected_survival

where observed_survival is the all-cause Kaplan-Meier survival probability from the cancer cohort, and expected_survival is the survival probability that the cohort would have experienced if they had the same all-cause mortality as the general population matched on age, sex, and calendar year -- obtained from population life tables. A relative survival of 0.50 at five years means the cancer cohort survived at 50% of the rate that a comparable general-population group would have survived.

Note: the 0.50 is a ratio, not a percentage-point deficit — the absolute deficit in this example is 40 percentage points (observed 0.40 minus expected 0.80 = −0.40). Relative survival above 1.0 is possible (a healthy-worker effect or a population selected for low baseline mortality) but requires careful interpretation.

Statistical cure in the net survival framework is indicated when the excess hazard returns to zero and the relative survival curve plateaus — it stops declining and stabilises at a fixed value that equals the proportion of the cohort effectively cured. This plateau is generally below 1.0 unless all patients are eventually cured; relative survival does not in general approach 1.0 as follow-up lengthens.

Pohar Perme versus Ederer I, Ederer II, and Hakulinen: why the estimator matters

Older methods -- Ederer I (1961), Ederer II (1961), and Hakulinen (1982) -- estimated relative survival using different constructions of expected survival in the denominator. All three are biased when the age distribution within the cohort is informative -- which it almost always is in cancer data, because older patients both have worse cancer prognosis and worse expected general-population survival.

The bias is upward: older patients accumulate deaths quickly and exit the risk set early, so later survival estimates come from the younger (better-prognosis) survivors; the expected-survival denominator is incorrectly handled, and net survival is over-estimated.

Pohar Perme (2012) corrected this by inverse probability weighting: at each event time, each patient is up-weighted by the inverse of their expected general-population survival, amplifying the contribution of older patients whose expected mortality is high and whose net survival would otherwise be under-counted.

In practice the Pohar Perme estimate is lower than Ederer II for older, higher-mortality cancer cohorts -- which is the correct direction; the earlier estimators were optimistically biased. Use Pohar Perme for all new analyses; revisit Ederer II only when replicating historical literature.

Life-table choice: national, regional, SES-specific, and the US claims-data caveat

The expected-survival denominator comes from a population life table stratified by age, sex, and calendar year. In most countries the national life table is the default comparator, and it is required for international comparisons. However:

  • Regional or socioeconomic-stratum-specific tables give more accurate comparators when the cancer cohort is not representative of the national population. A cohort of rural patients matched to an urban-dominant national life table will under-estimate expected mortality (rural mortality is higher than urban-dominated national rates), making expected survival too high and thereby deflating relative survival — the cancer appears more lethal than it is for that rural population.
  • In the United States, commercially insured populations have substantially lower all-cause mortality than the national population at the same age and sex -- because people who are healthy enough to be employed and insured are selected from a lower-mortality subgroup. Using the national life table in a commercial-claims analysis will over-state expected mortality rates, which lowers the expected survival in the denominator and thereby inflates relative survival — making the cancer appear less lethal than it actually is for that insured population. When possible, use an insured-population life table as the primary or at least as a sensitivity analysis. This is one of the most consequential yet under-discussed limitations of applying relative-survival methods to US administrative data.
  • Medicare populations are closer to the national elderly life table, but the dual-eligible (Medicaid and Medicare) subgroup has higher background mortality than national tables predict. Pre-specify the life table in the protocol, document its source, and always carry a sensitivity analysis with an alternative table.

Age-standardization for comparisons across registries and time periods

Comparing relative survival across cancer registries, time periods, or countries requires age-standardization, because different populations diagnose cancer patients at different age distributions. The International Cancer Survival Standard (ICSS) provides three weight sets for different cancer sites. The age-standardized relative survival is a weighted average of age-stratum-specific relative survival estimates using the ICSS weights (Corazziari et al. 2004) -- one line of arithmetic applied after stratum-specific estimation.

Excess hazard regression

When covariates must be modeled, the relative-survival framework uses excess hazard regression. The total observed hazard in the cancer cohort is decomposed as:

total_hazard = excess_hazard + expected_hazard

where expected_hazard is read from the life table at each patient's age/sex/year, and excess_hazard is the cancer-attributable component to be modeled. The excess hazard is modeled as a function of covariates (cancer stage, grade, treatment, socioeconomic status) using Poisson regression with an offset for the expected hazard (Dickman et al. 2004).

Regression coefficients are excess hazard ratios: the multiplicative change in cancer-attributable mortality associated with a one-unit change in the covariate. Excess hazard regression is the relative-survival analogue of Cox regression for cause-specific survival.

Route to relative survival cure models

When the excess hazard returns to zero at finite follow-up -- the cohort's survival converges to the matched general-population life table -- there is statistical evidence of cure. Relative survival cure models partition the cohort into a cured fraction (those whose excess hazard has reached zero) and an uncured fraction still experiencing excess mortality.

The cured fraction's survival thereafter follows the general-population life table. This is one of the primary input pathways for mixture-cure models in cancer registry settings; see the cure-models-mixture-cure entry for the full cure-model framework. Relative survival provides the necessary input when cause-of-death data are unavailable.

Pros, cons, and trade-offs

Relative / net survival versus cause-specific survival (when cause-of-death data exist):

  • Pros: Requires no cause-of-death coding at all; immune to misclassification of cause; directly comparable across registries with varying certification quality; the established standard for population-based cancer surveillance; estimable from any population with a matching life table.
  • Cons: Requires a matching population life table that accurately represents the background mortality of the study cohort; depends on the independence assumption (discussed above); cannot separate deaths caused by the cancer from deaths caused by cancer treatment toxicity (both are "excess" mortality); relative survival can exceed 1.0 in selected populations, which is numerically unintuitive; the Pohar Perme estimator has slightly higher variance than the simpler Ederer II at small sample sizes.
  • When to prefer: Population-based registry studies, SEER, EUROCARE, national cancer-plan evaluations, and any setting where cause-of-death coding is unreliable, missing, or differentially misclassified across comparison groups.

National versus population-specific life tables:

  • Pros of national tables: Universally available; standard for international comparisons; simple to implement.
  • Cons: Can over- or under-state expected mortality in non-representative cohorts (insured, rural, SES extreme). For US commercial-claims analyses, the national table overstates expected mortality rates, which deflates the expected survival denominator and thereby inflates the relative survival estimate, making the cancer appear less lethal than it is in that insured population.
  • When to prefer: National tables for SEER and registry work; population-specific or insured-cohort life tables for claims-based analyses when available, or at minimum as a sensitivity analysis.

When NOT to use

  • Non-fatal outcomes: Relative survival is a survival-only framework. For outcomes such as disease progression, response rate, readmission, or healthcare utilization there is no general-population expected comparator and the method does not apply; use standard time-to-event or frequency methods instead.
  • When reliable cause-of-death data are available and a cause-specific estimand is the target: If death certificates are adjudicated, if cancer-specific survival is the primary endpoint defined in the study protocol, or if the scientific question explicitly concerns cause-specific mortality (e.g., cardiovascular deaths in a cancer survivor cohort), prefer competing-risks methods with a validated mortality source hierarchy. Net survival and cause-specific survival answer subtly different questions and can diverge materially when non-cancer mortality is high (elderly patients, indolent cancers, heavily comorbid cohorts).
  • When the independence assumption is materially violated: If the cancer shares strong risk factors with non-cancer mortality (e.g., smoking-related lung cancer where the cohort would have had higher-than-national cardiovascular mortality even without the cancer), the excess hazard is overestimated and net survival is underestimated. Flag this as a structural limitation and carry cause-specific survival as a sensitivity analysis.
  • Do not apply to US commercial-claims data without a population-appropriate life table: Using the US national life table for a commercially insured cohort will produce upwardly biased expected survival and downwardly biased relative survival, systematically overstating the cancer's lethal burden in that selected population.

Interpreting the output

From the worked example: a colon cancer cohort of 5 patients followed 5 years. Observed all-cause survival at 5 years = 0.40 (2 of 5 patients alive at 5 years). Mean expected 5-year survival from matched life tables (age/sex/calendar year) = 0.80. Net (relative) survival = 0.40 / 0.80 = 0.50.

(1) Formal interpretation. The relative survival of 0.50 is an estimate of net survival computed here as a simple Ederer-style observed/mean-expected ratio — the worked example uses this transparent arithmetic to illustrate the concept. The Pohar Perme estimator applies inverse-probability weighting at each event time (upweighting older patients whose background mortality is higher) and gives a different, unbiased estimate; it should be used in any real analysis.

Under the independence assumption that each patient's expected general-population survival is independent of their actual cancer prognosis conditional on age, sex, and calendar year, the Pohar Perme estimator is consistent for net survival. Relative survival is a ratio of two survival probabilities, not itself a bounded probability; it can exceed 1.0 in populations with below-national-average background mortality. A value of 0.50 means the cancer cohort's observed all-cause survival is 50% of what the same-age, same-sex general population would have experienced -- with the remaining deficit attributed to the excess mortality from colon cancer and its treatment.

(2) Practical interpretation. The relative survival of 0.50 is a ratio: the cohort survived at half the rate of the matched general population (0.40 vs. 0.80). The absolute deficit is 0.80 − 0.40 = 0.40 — approximately 40 fewer would be alive per 100 newly diagnosed patients at 5 years compared with similarly aged people in the general population who did not receive a cancer diagnosis. The 0.50 itself is not a 50-percentage-point deficit; confusing the ratio with the absolute gap is a common communication error. This is the quantity reported in international cancer-survival comparisons (CONCORD, EUROCARE) and in SEER publications; it is interpretable across time periods and registries regardless of differences in cause-of-death certification quality.

For clinical communication, emphasize that it represents the survival experience attributable to the cancer diagnosis itself, net of the background mortality that those patients would have experienced regardless of the cancer.

Decision diagram

flowchart TD
  REG[Cancer registry: all-cause death date<br/>No reliable cause-of-death coding] -->|Observed survival| KM[Kaplan-Meier<br/>all-cause survival curve]
  LT[Population life table<br/>matched on age / sex / calendar year] -->|Expected survival| EXP[Expected survival<br/>for matched general population]
  KM --> RATIO[Relative survival<br/>= observed / expected]
  EXP --> RATIO
  RATIO -->|Equals| NET[Net survival estimate<br/>Cancer-attributable survival probability<br/>Pohar Perme weighting]
  NET --> Q{Analysis goal?}
  Q -->|Descriptive: survival curve| PP[Report net survival with 95% CI<br/>relsurv rs.surv Pohar Perme]
  Q -->|Covariate adjustment| EH[Excess hazard regression<br/>Poisson model, Dickman et al. 2004]
  Q -->|Long-term: cure fraction| CM[Relative survival cure model<br/>see cure-models-mixture-cure]
  style REG fill:#F6F6F5,stroke:#d97706
  style NET fill:#EDEDEB,stroke:#059669
Data flow from cancer registry and population life table through observed and expected survival to the relative / net survival estimate, then to the three downstream analysis goals -- descriptive net survival curve, excess hazard regression with covariates, and relative survival cure modeling.
flowchart TD
  Q1{Is cause-of-death<br/>reliably coded?} -- yes --> CMP[Competing-risks framework<br/>cause-specific / Fine-Gray<br/>with mortality source hierarchy]
  Q1 -- no or unreliable --> Q2{Is the estimand<br/>cancer-attributable survival?}
  Q2 -- yes --> Q3{Is a matched<br/>population life table<br/>available?}
  Q2 -- no --> OTHER[Non-fatal endpoint<br/>or all-cause analysis<br/>standard survival methods]
  Q3 -- yes --> PP2[Pohar Perme net survival<br/>relsurv rs.surv]
  Q3 -- no --> WARN[Cannot estimate relative survival<br/>use cause-specific or all-cause]
  PP2 --> COVA{Covariate adjustment<br/>needed?}
  COVA -- yes --> EXH[Excess hazard regression<br/>Dickman Poisson model]
  COVA -- no --> AGE[Age-standardize with ICSS weights<br/>for cross-registry comparison]
Decision logic for choosing relative survival versus cause-specific survival. Route to the Pohar Perme estimator when cause-of-death is unreliable and a matched life table exists; route to competing-risks methods when cause-of-death is well-coded.

Worked example

Scenario

A state cancer registry tracks 5 patients newly diagnosed with colon cancer in 2018, all followed for exactly 5 years. For each patient, the registry also records their expected 5-year survival probability, looked up from the US national life table matched on their age at diagnosis, sex, and calendar year. We want to estimate the 5-year net survival for this cancer -- the survival attributable to the colon cancer itself, independent of background mortality. Two patients are still alive at 5 years; three died during follow-up. Because death certificates in this registry are not reliably coded for cancer vs non-cancer cause, we use relative survival rather than cause-specific survival.

Dataset

One row per patient: follow-up outcome at 5 years and life-table expected 5-year survival. Expected survival is the probability the general population (matched on age/sex/year) would survive 5 years, read from the US national life table.

person_idvital_status_5yrdays_followedexpected_5yr_survival
1001alive18250.85
1002dead7300.78
1003alive18250.82
1004dead3650.76
1005dead10950.79
FIG. 1 — DESIGN TIMELINE
Relative survival: 5 colon cancer patients over 5-year (1825-day) follow-up
Relative survival: 5 colon cancer patients over 5-year (1825-day) follow-up

Steps

1Step 1 -- Count patients alive at 5 years: patients 1001 and 1003 survived all 1825 days. Observed 5-year survival: 2 / 5 = 0.40.
2Step 2 -- Compute mean expected 5-year survival from life tables: sum of expected_5yr = 0.85 + 0.78 + 0.82 + 0.76 + 0.79 = 4.00; mean expected survival = 4.00 / 5 = 0.80. This is the survival the age/sex/year-matched general population would have experienced.
3Step 3 -- Compute relative (net) survival: 0.40 / 0.80 = 0.5. This is the net survival estimate: for every person in the general population who survived 5 years, only 0.50 of the cancer patients survived, with the deficit attributed to the colon cancer.
4Step 4 -- Interpretation: a net survival of 0.5 means the colon cancer cohort had 50% of the 5-year survival probability that their age/sex/year-matched general-population counterparts had. The 40% observed survival minus the 80% expected survival yields the 40-percentage-point gap attributable to the cancer and its treatment.

Result

Observed 5-year survival = 2 / 5 = 0.40. Mean expected 5-year survival from life tables = 4.00 / 5 = 0.80. Relative (net) 5-year survival = 0.40 / 0.80 = 0.50. Net survival of 0.50 means this cancer cohort survived at 50% the rate of the age/sex-matched general population; the 50% deficit is attributed to the colon cancer and its treatment, net of background mortality.

Trade-offs

vs. Cause specific survival with competing risks framework
Pros of this
Does not require cause-of-death coding; eliminates misclassification; internationally comparable across registries with varying certification quality; applicable whenever a population life table exists; the established standard for cancer surveillance.
vs. Kaplan Meier all cause survival (observed survival, no life table adjustment)
Pros of this
Separates the cancer's contribution to mortality from the background mortality the cohort would have experienced regardless; enables time-trend comparisons by removing the confounding effect of secular improvements in background life expectancy.
vs. Ederer II or Hakulinen estimators (older relative survival methods)
Pros of this
Pohar Perme is unbiased for net survival; removes the systematic optimistic bias of older estimators in aged and high-mortality cohorts; recommended by IARC; implemented in current R packages.

Runnable example

Conceptual relative survival calculation in Python using lifelines for the Kaplan-Meier observed survival and a manual life-table division to produce relative survival. IMPORTANT: lifelines does not implement the Pohar Perme net survival estimator. For a production analysis use R (relsurv or popEpi) or SAS.

requires: pandas · lifelines
import pandas as pd
from lifelines import KaplanMeierFitter

# ── Worked-example cohort (5 colon cancer patients) ──
data = pd.DataFrame({
    "person_id":    [1001,  1002,  1003,  1004,  1005],
    "fu_days":      [1825,   730,  1825,   365,  1095],
    "event":        [   0,     1,     0,     1,     1],   # 0=alive, 1=dead (all-cause)
    "expected_5yr": [0.85,  0.78,  0.82,  0.76,  0.79],  # from national life table
})

HORIZON_DAYS = 1825  # 5 years

# ── 1. Observed all-cause survival at 5 years (Kaplan-Meier) ──
kmf = KaplanMeierFitter()
kmf.fit(data["fu_days"], event_observed=data["event"])
observed_surv = float(kmf.predict(HORIZON_DAYS))
print(f"Observed 5-year survival (KM):      {observed_surv:.4f}")

# ── 2. Mean expected 5-year survival from the life table ──
expected_surv = data["expected_5yr"].mean()
print(f"Mean expected 5-year survival (LT): {expected_surv:.4f}")

# ── 3. Relative (net) survival = observed / expected ──
# NOTE: This simple ratio is the intuitive definition. The Pohar Perme estimator applies
# inverse-probability weights at each event time to remove the age-structure bias; the
# ratio below agrees with the worked example (0.40 / 0.80 = 0.50) but does not implement
# the full weighting. Use relsurv in R for production Pohar Perme estimation.
relative_surv = observed_surv / expected_surv
print(f"Relative (net) 5-year survival:     {relative_surv:.4f}")
print()
print("Verify: 2 alive / 5 total = 0.40 observed; "
      "sum expected = 4.00, mean = 0.80; ratio = 0.40 / 0.80 = 0.50")

# ── 4. Excess hazard (approximate, for illustration) ──
# For a single time horizon, excess cumulative hazard = -ln(rel_surv).
import math
excess_cum_haz = -math.log(relative_surv) if relative_surv > 0 else float("inf")
print(f"Approximate 5-year cumulative excess hazard: {excess_cum_haz:.4f}")
print("For full excess hazard regression, use relsurv::excess() or popEpi::relpoisreg in R.")

Citations

FOUNDATIONAL / METHODS
  1. [1]Perme MP, Stare J, Esteve J. On estimation in relative survival. Biometrics. 2012;68(1):113-120.
  2. [2]Dickman PW, Adami HO. Interpreting trends in cancer patient survival. Journal of Internal Medicine. 2006;260(2):103-117.
APPLIED EXAMPLES
  1. [3]Dickman PW, Sloggett A, Hills M, Hakulinen T. Regression models for relative survival. Statistics in Medicine. 2004;23(1):51-64.
REPORTING & GUIDANCE
  1. [4]Corazziari I, Quinn M, Capocaccia R. Standard cancer patient population for age standardising survival ratios. European Journal of Cancer. 2004;40(15):2307-2316.