← Methods repository
CONCEPTINTERMEDIATEPYTHON · R · SAS5 citations

Fit-for-Purpose Data Assessment

A structured, pre-protocol process that judges whether a candidate real-world data source is relevant (captures the population, exposure, outcome, confounders, and follow-up the question requires) and reliable (accurate, complete, traceable, and consistently curated) enough to answer one specific regulatory or HTA question.

Framework Standardfit-for-purposedata-relevancedata-reliabilityspifdregulatory-readinessdata-quality-assessmenttarget-trialpicots
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

Before designing any real-world study, researchers must ask one question: does the database we are considering actually contain what we need to answer our specific question? Fit-for-purpose data assessment is the structured process that answers that question by checking two things: relevance (does the source capture the right patients, the right drug, the right outcome, and enough follow-up time?) and reliability (are those records accurate, complete, and produced by a trustworthy data process?). The output is a single, defensible verdict — go, no-go, or go only if certain gaps are addressed — tied to that one question, not to the database in general. A database can be perfectly fit for counting prescriptions and completely unfit for measuring whether patients die from a heart attack, even when both studies use the exact same patients.

When to use it
Any regulatory-grade or HTA-facing study, or any analysis where a data limitation could change the conclusion.
When the deliverable is a go/no-go decision for a specific submission; use the scorecard only as a pre-screen.
As the qualifying step for each source; pair with replication for robustness on high-stakes questions.
Watch out for
Front-loaded effort that delays first results and requires data-provenance documentation vendors may not readily provide.
Non-transferable; must be re-executed for each new question and cannot be compared cleanly across databases.
Does not test database-specific artifacts the way replication does; concordant-but-wrong results across sources sharing a flaw can still mislead.

Fit-for-purpose (FFP) data assessment

is the gatekeeping step that decides, before any analysis is programmed, whether a particular real-world data source can credibly answer a particular question. It is not a generic "data quality score" and it is not transferable: a database can be fit for purpose for a comparative drug-utilization study and entirely unfit for a comparative mortality study run on the same patients.

The assessment is organized around two axes made canonical by the FDA RWD guidance and operationalized by Gatto et al.'s Structured Process to Identify Fit-for-Purpose Data (SPIFD): relevance — does the source contain the population, exposure, outcome, key confounders, and follow-up duration the estimand demands? — and reliability — are those elements accurate, complete, traceable, and produced by a stable, documented data-curation (ETL) process?

The output is a documented go / no-go / go-with-mitigations verdict for one PICOTS-defined question, plus the specific sensitivity analyses that will probe the most judgment-dependent thresholds.

Core conceptual distinction

. FFP assessment sits upstream of, and is distinct from, three things it is often confused with. (1) vs database feasibility / attrition counting: a feasibility funnel tells you how many patients survive each eligibility step; FFP tells you whether the surviving cohort and its variables mean what the protocol needs them to mean. Feasibility is a necessary input to FFP, not a substitute.

(2) vs algorithm/outcome validation: validation estimates the operating characteristics (PPV, sensitivity) of one variable; FFP integrates those characteristics with relevance and curation evidence into a question-level decision. (3) vs a global data-quality grade: question-agnostic grading (completeness %, conformance checks) is a reliability input but cannot, by itself, declare fitness, because fitness is defined relative to a specific estimand. The decisive output of FFP is therefore not a number but a defensible decision tied to a question, with the residual risks named and mitigated.

Pros, cons, and trade-offs

(specific & comparative, naming the alternatives).

  • vs proceeding straight to analysis on a convenient database: FFP forces relevance/reliability to be argued in protocol language before code is written, which is exactly what FDA and EMA reviewers expect and what prevents an expensive study from being rejected for an avoidable data limitation (e.g., MA-only person-time with no fee-for-service claims, no death linkage for a mortality endpoint). Cost: it is front-loaded work that delays the first results and requires data-provenance documentation the vendor may not readily supply. Prefer FFP for any regulatory-grade or HTA-facing study.
  • vs a one-time, question-agnostic data-quality scorecard: a reusable scorecard is cheap and comparable across databases, but it systematically over- or under-states fitness because it ignores the estimand — a source with 99% completeness on labs is still unfit for a question that turns on outpatient mortality. FFP is question-specific and therefore more defensible, at the price of being non-transferable and needing re-execution for each new question. Prefer the scorecard only as a pre-screen to shortlist databases, then run FFP on the finalists.
  • vs multi-database replication as the primary safeguard: running the analysis in several databases is powerful against database-specific artifacts, but it is reactive and expensive, and concordant-but-wrong results across sources that share a structural flaw (e.g., all lack reliable cause-of-death) give false reassurance. FFP is proactive and cheaper. Use both when stakes are high: FFP to qualify each source, replication to test robustness.

When NOT to use — and when it is actively misleading or dangerous

.

  • When the question is not yet specified. Running FFP against a vague aim produces a meaningless verdict; fitness is undefined without an estimand. Fix PICOTS first.
  • When the assessment is treated as a checkbox. A FFP memo that recites "relevant and reliable" without source-level evidence (provenance, code-list hit rates, missingness by site/time/arm, linkage denominators) is more dangerous than none, because it manufactures false confidence and is exactly the artifact a regulator will probe. The danger is laundering an unfit source through a process veneer.
  • When a fatal relevance gap is rationalized into a "mitigation." If the outcome is out-of-hospital cardiac death and the source has no death-index linkage, no sensitivity analysis rescues it — that is a no-go, not a go-with-mitigations. Mitigations are for measurable, bounded uncertainty (a quantitative-bias analysis for a validated-but-imperfect outcome algorithm), not for structurally absent data.
  • When reliability is assumed because the database is large or familiar. Size is not accuracy; a marquee claims database can still drop fee-for-service claims for Medicare Advantage enrollees, lag adjudication, reverse claims, or bundle services — all of which silently corrupt the very variables the study depends on.

Data-source operational depth

.

  • Administrative claims (FFS vs MA vs commercial): Relevance strengths are exposure (NDC + `fill_date` + `days_supply`) and healthcare utilization/cost; weaknesses are clinical severity, labs, vitals, and cause of death. Critical reliability failure modes: Medicare Advantage person-time lacks fee-for-service claims — encounter data are incomplete and inconsistently submitted, so an MA enrollee can look like a non-user or a non-utilizer purely from missingness; restrict to enrollees with the relevant benefit (A/B/D, or commercial medical+pharmacy) and exclude MA-only person-time unless complete encounter data are demonstrated. Other failure modes: adjudication lag and claim reversals (right-censor with a data-maturity buffer), bundled/capitated services that hide individual procedures, plan-switching that breaks continuous enrollment, and sample/mail-order fills that distort `days_supply`. Differential competing risks matter: in elderly claims, death competes with the outcome and may be captured only via the Medicare enrollment database, not the claim stream — verify the mortality source before trusting any time-to-event endpoint.
  • EHR: Relevance strengths are labs, vitals, problem lists, and clinician notes (severity, indication); the dominant reliability problem is encounter-driven, network-bounded capture — a patient who seeks care outside the system is differentially unobserved ("leakage"), so absence of a record is ambiguous (no event vs cared-for elsewhere). Structured fields are often sparse or entered inconsistently across sites; note availability varies by visit type. Linkage to claims is the standard fix for completeness, and explicit observation windows plus loss-to-follow-up handling are mandatory.
  • Registry: Relevance strength is adjudicated, clinically rich outcomes and disease staging (e.g., cancer registries); weakness is incomplete longitudinal pharmacy exposure and follow-up. Reliability turns on enrollment eligibility, case-ascertainment completeness, adjudication rules, and reporting lag. Almost always requires linkage to claims (for exposure/utilization) and to a death index (for mortality).
  • Linked claims–EHR–vital-records: The ideal substrate (EHR severity + claims completeness + reliable mortality) but linkage introduces selection (only the linkable subset, which may differ systematically) and date-reconciliation problems across order, fill, and service dates that must be resolved before time-zero assignment. Report the linkage denominator and compare linked vs unlinkable patients.

Worked example (claims-style logic)

Question (PICOTS fixed): among adults ≥18 with type-2 diabetes, does initiating a GLP-1 receptor agonist vs a DPP-4 inhibitor change 3-point MACE risk over 2 years? Candidate source: a commercial + Medicare fee-for-service claims database. (1) Relevance — population: confirm ≥2 T2D diagnoses are codeable and the age band is present. Exposure: both classes are identifiable by NDC with `fill_date` and `days_supply`, so new-user status (no prior fill in a 365-day washout) and on-treatment episodes are constructible.

Outcome

3-point MACE = nonfatal MI + nonfatal stroke + cardiovascular death; the MI/stroke components are claims-codeable, but CV death requires a death source — plain claims give an end-of-enrollment date, not a cause, so without National Death Index or Medicare-enrollment death linkage the outcome is only partially ascertainable: a relevance gap, not a reliability nuance. Confounders: HbA1c and BMI (key effect modifiers) are largely absent in claims — note this as a candidate for linkage or quantitative bias analysis.

Follow-up

2 years requires continuous A/B/D (or commercial) enrollment; check the median observable follow-up against the 2-year horizon. (2) Reliability: obtain ETL/provenance documentation and refresh date; right-censor with a 3-month maturity buffer for adjudication lag; exclude MA-only person-time because fee-for-service claims are missing there; profile missingness of `days_supply` and date fields by calendar quarter and by arm; verify the mortality source completeness against expected age-specific rates.

(3) Verdict: go-with-mitigations — fit for nonfatal MACE components and exposure; conditionally fit for fatal MACE only if death-index linkage is secured; otherwise restrict the endpoint to nonfatal MACE or escalate to a linked source. (4) Pre-specified sensitivity analyses targeting the judgment-dependent thresholds: vary the washout (180 vs 365 days), the data-maturity buffer (1 vs 3 vs 6 months), the MACE algorithm definition (and apply a PPV-based quantitative bias analysis), and the MA-exclusion rule, reporting cohort counts and the estimate's stability at each step.

Decision diagram

flowchart TD
  Q[PICOTS + estimand fixed<br/>one specific question] --> REL{Relevance:<br/>can the source capture<br/>population, exposure, outcome,<br/>confounders, follow-up?}
  REL -->|fatal gap, e.g. no death linkage<br/>for a mortality endpoint| NOGO[NO-GO<br/>wrong source for this question]
  REL -->|yes / partial| RELI{Reliability:<br/>accurate, complete, traceable?<br/>documented stable ETL?}
  RELI -->|exclude MA-only person-time,<br/>verify mortality source,<br/>profile missingness + lag| VERDICT{Fit-for-purpose verdict}
  VERDICT -->|all elements adequate| GO[GO]
  VERDICT -->|bounded, measurable uncertainty| MIT[GO WITH MITIGATIONS<br/>+ pre-specified sensitivity analyses]
  VERDICT -->|structural gap rationalized| NOGO
  MIT --> SENS[Sensitivity: washout length,<br/>data-maturity buffer, algorithm PPV/QBA,<br/>MA-exclusion rule]
SPIFD-style fit-for-purpose decision flow. Relevance is judged first against the fixed estimand; a structural relevance gap is a no-go that no mitigation rescues. Reliability is judged next, and only bounded, measurable uncertainty qualifies for a go-with-mitigations verdict tied to pre-specified sensitivity analyses.
flowchart LR
  subgraph Relevance[What the question needs]
    P[Population] --- E[Exposure] --- O[Outcome] --- C[Confounders] --- F[Follow-up]
  end
  Relevance --> Claims[Claims:<br/>+ exposure, utilization, cost<br/>- severity, labs, cause of death]
  Relevance --> EHR[EHR:<br/>+ labs, vitals, notes, severity<br/>- network leakage, sparse fields]
  Relevance --> Registry[Registry:<br/>+ adjudicated outcomes, staging<br/>- longitudinal exposure/follow-up]
  Claims --> Linked[Linked claims-EHR-death index:<br/>severity + completeness + mortality<br/>- linkage selection + date reconciliation]
  EHR --> Linked
  Registry --> Linked
Relevance-by-source matrix. No single source is universally fit; each trades a relevance strength against a structural weakness, and linkage is the standard route to close the gaps at the cost of selection and date-reconciliation problems.

Worked example

Scenario

A research team wants to study whether a GLP-1 receptor agonist (a diabetes drug) reduces the risk of serious heart events compared with a DPP-4 inhibitor (another diabetes drug) over two years. The primary outcome is called 3-point MACE: nonfatal heart attack, nonfatal stroke, or cardiovascular death. Before writing a single line of analysis code, the team runs a fit-for-purpose assessment on a commercial plus Medicare fee-for-service claims database. The table below lists each assessment dimension, what the team checks, and whether it passes or raises a concern.

Dataset

Fit-for-purpose checklist for a GLP-1 vs DPP-4 MACE study in a commercial plus Medicare FFS claims database

DimensionCheckWhat the analyst looks forVerdict
RelevancePopulationAre adults with type-2 diabetes identifiable using diagnosis codes?PASS
RelevanceExposure captureAre both drug classes recorded by prescription fill date and days supply so new-user status can be defined?PASS
RelevanceNonfatal outcomeCan heart attack and stroke be identified using validated hospital diagnosis codes?PASS
RelevanceFatal outcomeIs there a linked death source that records cause of death for cardiovascular deaths?CONCERN — death linkage not confirmed
RelevanceKey confoundersAre HbA1c (blood sugar control) and BMI recorded in the claims?CONCERN — largely absent in claims
RelevanceFollow-up durationCan most patients be observed continuously for the full two-year horizon?PASS — median observable follow-up exceeds 2 years
ReliabilityMedicare Advantage person-timeAre Medicare Advantage enrollees excluded or is complete encounter data confirmed? MA-only records lack fee-for-service claims and can make patients look like non-users.CONCERN — MA-only person-time must be excluded
ReliabilityAdjudication lagAre the most recent months of data complete, or do recent claims still need time to be processed and paid?CONCERN — a 3-month maturity buffer is required
ReliabilityDays supply completenessIs the days-supply field populated and plausible (1 to 180 days) for the study drugs?PASS
ReliabilityData provenanceHas the vendor provided documentation of how the database is built and updated?PASS — documentation available

Steps

1Work through the relevance checks first: confirm that the source can capture each piece of the PICOTS question before checking data quality.
2The nonfatal MACE components (heart attack, stroke) are identifiable by hospital diagnosis codes — these checks pass.
3Cardiovascular death is a fatal outcome and requires a cause-of-death source such as the National Death Index; plain claims only record when enrollment ended, not why the patient died — this is a relevance gap, not a quality nuance.
4HbA1c and BMI are clinical measurements rarely captured in claims, so the team notes them as a confounder gap requiring either a linked EHR or a sensitivity analysis.
5Move to reliability: Medicare Advantage enrollees in the database may appear to have no prescription fills simply because their insurer does not submit fee-for-service claims — including their person-time would silently corrupt the exposure measure, so it must be excluded.
6The maturity buffer check confirms that very recent claims are incomplete due to adjudication lag; censor the data three months before the extraction date.
7Summarize the verdict: the source is fit for the nonfatal MACE components and for exposure measurement; it is not fit for the full 3-point MACE endpoint unless death-index linkage is secured.

Result

Verdict: GO WITH MITIGATIONS for nonfatal MACE (heart attack + stroke); CONDITIONAL NO-GO for cardiovascular death until death-index linkage is confirmed. Key gap: fatal outcome ascertainment. Required mitigations before analysis:

  1. exclude Medicare Advantage-only person-time,
  2. apply a 3-month data-maturity censor,
  3. secure death-index linkage or restrict the primary endpoint to nonfatal MACE. Pre-specified sensitivity analyses: vary the washout length (180 vs 365 days), vary the maturity buffer (1 vs 3 vs 6 months), test the MACE diagnosis algorithm with and without a PPV-based correction, and report cohort counts under each MA-exclusion rule.

Trade-offs

vs. Proceeding straight to analysis on a convenient database
Pros of this
Surfaces fatal relevance/reliability gaps (MA-only person-time, missing death linkage, unvalidated outcomes) in protocol language before code is written, matching FDA/EMA expectations and avoiding avoidable rejection.
vs. A one time, question agnostic data quality scorecard
Pros of this
Question-specific and tied to the estimand, so the fitness verdict is defensible to a reviewer rather than a generic completeness statistic.
vs. Multi database replication as the primary safeguard
Pros of this
Proactive and cheaper; prevents a structurally flawed source from entering the study in the first place.

Runnable example

Fit-for-purpose profiling for a candidate claims source against a fixed question. This does what the SPIFD process does operationally: quantify the relevance/reliability evidence a reviewer will demand. It is profiling/feasibility code, not an estimation step.

requires: pandas · numpy
import pandas as pd
import numpy as np

REQUIRED_FOLLOWUP_DAYS = 730   # 2-year estimand horizon
EXPOSURE_NDCS = {...}          # NDC set for the study + comparator drug classes
OUTCOME_DX = {...}             # validated MI/stroke code set for nonfatal MACE

def ffp_claims_profile(enroll, rx, dx, death):
    out = {}

    # --- RELIABILITY: MA-only person-time lacks fee-for-service claims -> must be excludable ---
    person_days = (enroll["enroll_end"] - enroll["enroll_start"]).dt.days.clip(lower=0)
    ma_only_days = person_days[enroll["plan_type"].eq("MA")].sum()
    out["pct_person_time_MA_only"] = 100 * ma_only_days / max(person_days.sum(), 1)
    out["pct_enrollees_with_AB_D_benefit"] = 100 * enroll.groupby("person_id")["ab_d"].max().mean()

    # --- RELEVANCE: follow-up duration vs the estimand horizon ---
    span = (enroll.groupby("person_id")
                  .apply(lambda g: (g["enroll_end"].max() - g["enroll_start"].min()).days))
    out["median_observable_followup_days"] = float(span.median())
    out["pct_with_full_horizon"] = 100 * (span >= REQUIRED_FOLLOWUP_DAYS).mean()

    # --- RELEVANCE: exposure capture (NDC + days_supply usable?) ---
    exp = rx[rx["ndc"].isin(EXPOSURE_NDCS)]
    out["pct_exposed_with_valid_days_supply"] = 100 * exp["days_supply"].between(1, 180).mean()
    out["n_exposed_persons"] = exp["person_id"].nunique()

    # --- RELEVANCE: nonfatal-outcome code-list hit rate ---
    out["n_persons_with_outcome_code"] = dx.loc[dx["dx_code"].isin(OUTCOME_DX), "person_id"].nunique()

    # --- RELEVANCE/RELIABILITY: fatal endpoint depends on a real death source w/ cause ---
    out["pct_deaths_with_known_source"] = 100 * death["death_source"].notna().mean()
    out["pct_deaths_with_cause"] = 100 * death["cause_of_death"].notna().mean()  # ~0 => CV death NOT ascertainable

    # --- RELIABILITY: date-field missingness by calendar quarter (data-maturity / lag signal) ---
    rx_q = rx.assign(q=rx["fill_date"].dt.to_period("Q"))
    out["fill_date_missing_by_quarter"] = (rx_q["fill_date"].isna()
                                           .groupby(rx_q["q"]).mean().mul(100).round(2).to_dict())
    return pd.Series(out)

# profile = ffp_claims_profile(enroll, rx, dx, death)
# Verdict logic: fatal MACE is fit ONLY if pct_deaths_with_cause is non-trivial; otherwise restrict to nonfatal MACE
# or escalate to a death-index-linked source. MA-only person-time must be excluded before any time-to-event analysis.

Citations

FOUNDATIONAL / METHODS
  1. [1]Gatto NM, Campbell UB, Rubinstein E, Jaksa A, Mattox P, Mo J, Reynolds RF. The Structured Process to Identify Fit-For-Purpose Data: A Data Feasibility Assessment Framework. Clinical Pharmacology & Therapeutics. 2022;111(1):122-134.
  2. [2]Gatto NM, Vititoe SE, Rubinstein E, Reynolds RF, Campbell UB. A Structured Process to Identify Fit-for-Purpose Study Design and Data to Generate Valid and Transparent Real-World Evidence for Regulatory Uses. Clinical Pharmacology & Therapeutics. 2023;113(6):1235-1239.
  3. [3]FDA. Real-World Data: Assessing Electronic Health Records and Medical Claims Data To Support Regulatory Decision-Making for Drug and Biological Products. Guidance for Industry. 2024.
APPLIED EXAMPLES
  1. [4]Schneeweiss S, Patorno E. Conducting Real-world Evidence Studies on the Clinical Outcomes of Diabetes Treatments. Endocrine Reviews. 2021;42(5):658-690.
REPORTING & GUIDANCE
  1. [5]Hernán MA, Robins JM. Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. American Journal of Epidemiology. 2016;183(8):758-764.