← Methods repository
CONCEPTINTERMEDIATEPYTHON · R4 citations

Multi-Criteria Decision Analysis (MCDA)

A structured family of methods that makes a healthcare decision's value criteria explicit (efficacy, safety, unmet need, equity, disease burden, cost), elicits weights for how much each criterion matters (swing weighting, AHP pairwise comparisons, or DCE-derived weights), scores each alternative on each criterion against declared worst-best anchors, and aggregates - most often as a weighted additive value model - into a transparent total value per alternative, used to structure HTA deliberation, portfolio prioritization, and quantitative benefit-risk rather than to replace them.

Economic Evaluationmcdamulti-criteria-decision-analysisswing-weightingahpadditive-value-modelbenefit-riskhta-deliberationportfolio-prioritization
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

Multi-criteria decision analysis (MCDA) is a way to compare treatment options when several things matter at once - how well a drug works, how safe it is, how badly patients need something new. A committee first agrees on the list of criteria, scores each option on each criterion from 0 to 100, and assigns weights that say how much each criterion matters; each option's weighted scores are then added up into one total for comparison. The total makes the committee's trade-offs visible and checkable, but it is only as good as the weights - different people, methods, or framings can produce different weights and sometimes a different winner, which is why the ranking is meant to structure the discussion, not replace it.

When to use it
Prefer MCDA when the decision genuinely turns on criteria outside the QALY or no threshold logic exists (portfolio triage, severity frameworks, benefit-risk);
Prefer MCDA when multiple non-commensurable outcomes must be traded off explicitly; prefer cost-effectiveness when one clinically accepted outcome dominates the decision and cost per unit of it is the question.
They compose rather than compete - run a preference study (DCE/conjoint) when the weight source must be patients or the public at scale, and feed those weights into the MCDA;
Watch out for
The total value score has no opportunity-cost interpretation - 72 points buys no statement about health displaced elsewhere in the budget - and loses the QALY's cross-appraisal comparability and threshold decision rule.
Requires weight elicitation machinery (panels, swing exercises, sensitivity analysis) that a CEA does not, and its outputs are method- and panel-dependent where a CEA's are data-dependent.
Committee-elicited MCDA weights come from a handful of people in a room; a well-designed DCE brings a defensible patient or population sample and confidence intervals around every weight.

Most health decisions are multi-attribute whether we admit it or not: an HTA committee weighing a drug is trading off survival gain against toxicity against unmet need against budget; a payer prioritizing a formulary is doing the same across products; a regulator's benefit-risk call balances effect sizes against harms

MCDA makes that implicit weighing explicit

The canonical process (ISPOR MCDA Emerging Good Practices Task Force, Thokala et al. 2016; Marsh et al. 2016) is:

  1. structure the problem - define the decision, the alternatives, and a criteria set that is complete, non-redundant, and preferentially independent;
  2. measure performance - build a performance matrix of each alternative on each criterion from trials, RWE, and elicitation;
  3. score - convert natural-unit performance to a common 0-100 partial value scale against declared worst and best anchors (linear or elicited value functions);
  4. weight - elicit how much a swing from worst to best on each criterion matters relative to the others;
  5. aggregate - usually the additive value model V(a) = sum over criteria of w_k x s_k(a), with weights normalized to sum to 1;
  6. test - sensitivity analysis on weights and scores;
  7. deliberate - the numbers structure the discussion, they do not end it.

Value measurement methods

Swing weighting asks the committee to imagine the worst hypothetical alternative and rank-then-rate which single criterion swing (worst to best) they would fix first; the top swing gets 100 points, others are rated relative to it, and points are normalized to weights It is the method most consistent with the additive model because weights are anchored to the actual criterion ranges - a weight only means anything relative to the swing it covers

AHP (Analytic Hierarchy Process) derives weights from pairwise comparisons on a 1-9 verbal scale via the principal eigenvector, with a consistency ratio to flag incoherent judgments; it is easy to field but its verbal scale and rank-reversal behavior are theoretically contested DCE-derived weights estimate them from choice experiments over attribute profiles, importing preference-study machinery (and its sample, framing, and attribute-range dependence) into the weight set

Outranking methods (ELECTRE, PROMETHEE) avoid full aggregation by pairwise-comparing alternatives with concordance thresholds - useful when trade-offs are contested, but harder to explain to a deliberative committee.

The additive model's independence assumptions are the load-bearing wall

A weighted sum is only valid when criteria are mutually preferentially independent: how much you value a swing on safety must not depend on the level of efficacy When criteria interact (a toxicity matters more when survival gain is small), the additive form misstates value and you need multiplicative or other non-additive forms - or a re-structured criteria set

Double counting is the everyday violation: putting "QALY gain" and "quality of life" and "severity" in one criteria set counts the same value twice; cost criteria alongside health criteria quietly re-derive a cost-effectiveness threshold the committee never agreed to.

Pros, cons, and trade-offs

(specific and comparative).

  • vs cost-utility analysis (CUA): CUA collapses value to one metric (the QALY) and one decision rule (the ICER vs a threshold), which buys comparability across appraisals and decades of methods guidance - but it cannot natively carry unmet need, severity, equity, or innovation except as ad hoc modifiers. MCDA carries them explicitly with committee-owned weights, at the price of losing the QALY's interpersonal comparability and opportunity-cost logic: an MCDA total value of 72 has no exchange rate against the health forgone elsewhere in the budget. Prefer CUA for reimbursement decisions inside a budget-constrained system with an established threshold; prefer MCDA when the decision explicitly trades off criteria the QALY cannot hold, or where no threshold logic exists (portfolio triage, early pipeline, orphan/severity frameworks).
  • vs deliberation alone (unaided committee judgment): Unaided deliberation is flexible and cheap but opaque - weights live in members' heads, anchoring and loudest-voice effects go unmeasured, and consistency across meetings is unauditable. MCDA forces the value judgments into the open where they can be challenged and reused. The cost is real: elicitation burden, false precision risk, and the temptation to treat the score as the decision.
  • vs preference studies (DCE/conjoint) as the weight source: Committee swing weighting is fast and produces weights owned by the actual decision makers, but from a handful of people. DCE-derived weights bring a defensible sample (patients, public) and statistical machinery, but the weights inherit the experiment's attribute ranges and framing, and the committee may not feel bound by preferences it did not express.

When NOT to use - and when it is actively misleading

  • Do not use MCDA as a substitute for economic evaluation in a budget-constrained reimbursement decision. Total value scores carry no opportunity-cost information; ranking by MCDA score and funding down the list can displace more health than it buys. If cost enters at all, keep it outside the value score (value-for-money displays) rather than as a weighted criterion.
  • Double counting. Overlapping criteria (efficacy + QoL + severity that is itself defined by efficacy shortfall) silently multiply one attribute's weight. Criteria sets must be tested for redundancy before any weight is elicited.
  • Weight elicitation fragility. Weights move with the elicitation method (swing vs AHP vs DCE on the same problem yield different weights), with attribute ranges (halve the efficacy range and its swing weight should roughly halve - committees routinely fail this range sensitivity check), with framing, and with who is in the room. Reporting a single weighted total without weight sensitivity analysis is misleading precision.
  • Independence violations. If criterion values interact, the additive sum is the wrong functional form - check preferential independence explicitly during problem structuring, not after the ranking is computed.
  • Score laundering. When the performance matrix cells come from weak or heterogeneous evidence (a registry rate next to an RCT effect next to an expert guess), the tidy 0-100 scores hide the uncertainty gradient; carry evidence uncertainty into the sensitivity analysis or display it alongside the scores.

Interpreting the output

Consider the worked example: Drug A scores 72 and Drug B scores 67 under the committee's weights (efficacy 0.5, safety 0.3, unmet need 0.2), placing Drug A first by 5 points.

Formal interpretation: The weighted total of 72 is a composite index constructed by multiplying each criterion's 0–100 rescaled performance score by a normalized weight and summing The 5-point gap is only as meaningful as the weights that produced it: weight elicitation is known to be sensitive to the method used (swing weighting, AHP, DCE), the attribute ranges presented (halving the efficacy anchor range should roughly halve efficacy's swing weight — committees routinely fail this range-sensitivity check), and the composition of the eliciting group

The 0–100 rescaling assumes linearity within each criterion's range; if the committee's preferences are non-linear (e.g., diminishing marginal value of additional OS gain beyond 8 months), the additive model mis-ranks alternatives Preferential independence — required for the additive model to be valid — must be checked during problem structuring, not inferred from a clean-looking output.

Practical interpretation: Report the weighted total alongside the underlying performance matrix and the weights, not as a standalone score. Show the weight-sensitivity threshold at which the ranking flips — in this case, if the efficacy weight falls below the flip point, Drug B wins — so deliberation focuses on whether the committee is confident enough in the efficacy weight to sustain the Drug A ranking. Do not use the total score for cost-effectiveness inference: MCDA scores carry no opportunity-cost information and cannot substitute for an incremental cost- effectiveness ratio.

Decision diagram

flowchart TD
  Q[Decision problem<br/>alternatives + decision makers] --> C[Select criteria<br/>complete, non-redundant,<br/>preferentially independent]
  C --> PM[Performance matrix<br/>natural units per criterion<br/>from trials + RWE + elicitation]
  PM --> S[Score: partial value 0-100<br/>against declared worst-best anchors]
  C --> W[Weight: swing weighting / AHP /<br/>DCE-derived, normalized to sum to 1]
  S --> V[Aggregate: additive value model<br/>V = sum of weight x score]
  W --> V
  V --> SA{Sensitivity analysis:<br/>do plausible weight changes<br/>flip the ranking?}
  SA -- Ranking stable --> D[Deliberate with the numbers<br/>MCDA structures, committee decides]
  SA -- Ranking flips --> R[Report the flip point -<br/>the decision turns on that weight]
The value-measurement MCDA pipeline - structure the problem, build the performance matrix from evidence, score against anchors, elicit and normalize weights, aggregate additively, and stress-test the ranking before the committee deliberates with (not under) the numbers.
gantt
  title One HTA committee MCDA cycle from scoping to deliberation
  dateFormat YYYY-MM-DD
  axisFormat %d %b
  section Model build
  Scoping workshop - criteria set agreed :crit, e1, 2025-09-01, 14d
  Swing-weight elicitation panel :crit, e2, 2025-09-15, 14d
  Performance matrix + partial value scoring :crit, e3, 2025-09-29, 21d
  section Decision
  Deliberation + weight sensitivity :done, e4, 2025-10-20, 5d
A realistic eight-week MCDA cycle - two weeks of problem structuring, two of weight elicitation, three of evidence scoring, then a deliberation week where the weighted totals (Drug A 72 vs Drug B 67) and their sensitivity structure the committee discussion.

Worked example

Scenario

An HTA committee compares two drugs for the same disease using three criteria - efficacy (overall-survival gain in months), safety (serious adverse events per 100 patients), and unmet need addressed (a committee score from 0 to 100). Anchors were agreed in advance: survival gain runs from 0 (worst) to 10 months (best), the adverse-event rate from 20 per 100 (worst) to 0 (best), and unmet need is already on a 0-100 scale. In the swing-weighting session the committee put the efficacy swing first at 100 points, the safety swing at 60, and the unmet-need swing at 40. We rescale each drug's performance to 0-100 scores, normalize the weights, and add up the weighted scores.

Dataset

The performance matrix the committee sees - each drug's measured performance per criterion in natural units (swing points elicited separately - efficacy 100, safety 60, unmet need 40).

alternativeos_gain_monthssae_rate_per100unmet_need_score
Drug A8870
Drug B6250
FIG. 1 — DESIGN TIMELINE
Timeline bars for the scoping workshop, swing-weight elicitation, performance scoring, and deliberation phases of one MCDA cycle, with the model-build and decision-week spans shaded beneath.
Generated timeline of the eight-week MCDA cycle - scoping workshop, swing-weight elicitation (points 100/60/40), performance-matrix scoring, and the deliberation week where the weighted totals (Drug A 72 vs Drug B 67) structure the committee discussion.

Steps

1Normalize the swing points to weights that sum to 1. Total points = 100 + 60 + 40 = 200, so efficacy weight = 100/200 = 0.5, safety weight = 60/200 = 0.3, unmet-need weight = 40/200 = 0.2.
2Rescale efficacy to a 0-100 score between the anchors (0 worst, 10 best). Drug A = (8-0)/(10-0)*100 = 80; Drug B = (6-0)/(10-0)*100 = 60.
3Rescale safety the same way, remembering lower is better (20 worst, 0 best). Drug A = (20-8)/(20-0)*100 = 60; Drug B = (20-2)/(20-0)*100 = 90.
4Unmet need is already on the 0-100 scale, so Drug A scores 70 and Drug B scores 50.
5Add up weight times score for Drug A. Total value = 0.5*80 + 0.3*60 + 0.2*70 = 40 + 18 + 14 = 72.
6Add up weight times score for Drug B. Total value = 0.5*60 + 0.3*90 + 0.2*50 = 30 + 27 + 10 = 67.
7Compare and stress-test. Drug A leads by 72 - 67 = 5 points on the back of efficacy; if the committee dropped the efficacy weight to 0.2 (with safety 0.48 and unmet need 0.32), Drug B would win - so report that the ranking turns on the efficacy weight.

Result

Drug A total value = 72, Drug B total value = 67 - Drug A ranks first by 5 points under the committee's weights (0.5 efficacy, 0.3 safety, 0.2 unmet need), and sensitivity analysis shows the ranking flips if the efficacy weight falls far enough, so the deliberation should focus on how firmly the committee holds that weight.

Trade-offs

Pros of this
Carries criteria the QALY cannot hold - unmet need, severity, equity, innovation, delivery burden - as explicit weighted criteria with committee-owned trade-offs, instead of as unexplained modifiers around an ICER.
Pros of this
Handles more than one effectiveness dimension at once - a CEA's single natural-unit outcome (cost per event avoided) cannot trade efficacy against safety against unmet need, which is exactly the multi-attribute structure MCDA formalizes.
Pros of this
MCDA is the decision framework - it consumes preference weights and turns them into a ranked, deliberation-ready comparison of actual alternatives; a DCE alone quantifies preferences but decides nothing.

Runnable example

Minimal additive value model - the core MCDA arithmetic. Inputs: a performance matrix of alternatives x criteria in natural units, per-criterion worst/best anchors (direction-aware - for a harm, worst is the high value), and raw swing-weight points.

requires: pandas
import pandas as pd

# Performance matrix in NATURAL units (3 criteria x 2 alternatives).
perf = pd.DataFrame(
    {"os_gain_months": [8.0, 6.0],      # efficacy: overall-survival gain vs standard of care
     "sae_rate_per100": [8.0, 2.0],     # safety: serious adverse events per 100 patients (lower = better)
     "unmet_need_score": [70.0, 50.0]}, # committee-scored unmet need addressed, already on 0-100
    index=["Drug A", "Drug B"])

# Declared anchors: worst and best PLAUSIBLE levels per criterion (set during problem structuring,
# BEFORE weights are elicited - swing weights are only meaningful relative to these ranges).
anchors = {                      # (worst, best)
    "os_gain_months":   (0.0, 10.0),
    "sae_rate_per100":  (20.0, 0.0),   # harm: worst is the HIGH rate
    "unmet_need_score": (0.0, 100.0),
}

# Raw swing-weight points: top-ranked swing = 100, others rated relative to it.
swing_points = {"os_gain_months": 100.0, "sae_rate_per100": 60.0, "unmet_need_score": 40.0}

def partial_value(x: float, worst: float, best: float) -> float:
    """Linear 0-100 partial value between anchors; direction-aware via the anchor order."""
    return 100.0 * (x - worst) / (best - worst)

def additive_mcda(perf: pd.DataFrame, anchors: dict, swing_points: dict) -> pd.DataFrame:
    total_pts = sum(swing_points.values())
    weights = {k: v / total_pts for k, v in swing_points.items()}   # normalize to sum to 1
    scores = perf.apply(lambda col: partial_value(col, *anchors[col.name]), axis=0)
    contrib = scores * pd.Series(weights)            # weighted contribution per criterion
    out = contrib.add_suffix("_wtd")
    out["total_value"] = contrib.sum(axis=1)
    out["rank"] = out["total_value"].rank(ascending=False).astype(int)
    return out.round(2)

result = additive_mcda(perf, anchors, swing_points)
print(result)
# Drug A: 0.5*80 + 0.3*60 + 0.2*70 = 72.0 (rank 1); Drug B: 0.5*60 + 0.3*90 + 0.2*50 = 67.0 (rank 2)

# Minimal weight sensitivity: at what efficacy weight do the alternatives tie?
for w_eff in (0.50, 0.40, 0.30, 0.20):
    rest = 1.0 - w_eff
    w = {"os_gain_months": w_eff, "sae_rate_per100": rest * 0.6, "unmet_need_score": rest * 0.4}
    scores = perf.apply(lambda col: partial_value(col, *anchors[col.name]), axis=0)
    tot = (scores * pd.Series(w)).sum(axis=1)
    print(f"w_efficacy={w_eff:.2f}: A={tot['Drug A']:.1f}  B={tot['Drug B']:.1f}")

Citations

FOUNDATIONAL / METHODS
  1. [1]Thokala P, Devlin N, Marsh K, Baltussen R, Boysen M, Kalo Z, Longrenn T, Mussen F, Peacock S, Watkins J, IJzerman M. Multiple criteria decision analysis for health care decision making - an introduction: report 1 of the ISPOR MCDA Emerging Good Practices Task Force. Value in Health. 2016;19(1):1-13.
  2. [2]Marsh K, IJzerman M, Thokala P, Baltussen R, Boysen M, Kalo Z, Lonngren T, Mussen F, Peacock S, Watkins J, Devlin N. Multiple criteria decision analysis for health care decision making - emerging good practices: report 2 of the ISPOR MCDA Emerging Good Practices Task Force. Value in Health. 2016;19(2):125-137.
  3. [3]Marsh K, Sculpher M, Caro JJ, Tervonen T. The use of MCDA in HTA: great potential, but more effort needed. Value in Health. 2018;21(4):394-397.
APPLIED EXAMPLES
  1. [4]Baltussen R, Marsh K, Thokala P, Diaby V, Castro H, Cleemput I, Garau M, Iskrov G, Olyaeemanesh A, Mirelman A, Mobinizadeh M, Morton A, Tringali M, van Til J, Valentim J, Wagner M, Jansen MP, Bijlmakers L, Oortwijn W, Broekhuizen H. Multicriteria decision analysis to support health technology assessment agencies: benefits, limitations, and the way forward. Value in Health. 2019;22(11):1283-1288.