← Methods repository
CONCEPTINTERMEDIATEPYTHON · R · SASlast reviewed 2026-08-25 · updated 2026-08-25 · 4 citations

CPRD (Clinical Practice Research Datalink)

The UK's primary-care research database: anonymized electronic health records from NHS general practices covering roughly one in six Britons, with diagnoses, prescriptions, labs, and referrals linkable to hospital episodes (HES), death registration, and disease-specific registries — the reference source for UK pharmacoepidemiology.

Data Sourcecprdukprimary-careehrdata-sourceread-codesnomedhes-linkage
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

CPRD holds longitudinal primary-care EHRs from UK GP practices: every consultation, diagnosis (Read-coded), prescription issued, lab result, and referral, linkable to hospital admissions, deaths, and registries like cancer and pregnancy. Because NHS primary care is universal and most chronic-disease management happens in general practice, CPRD offers near-population-based coverage with decades of continuous records per patient.

When to use it
—Chronic-disease and safety questions; US claims for utilization/cost endpoints.
—Long-latency safety and UK-policy-relevant work; Truveta for large-scale US inference.
—Phenotype precision matters; SNDS for whole-population utilization rates.
Watch out for
—No itemized costs; weaker procedure granularity; OTC exposure blind spot.
—Smaller absolute population; UK-specific care patterns limit transportability.
—Partial linkage coverage vs SNDS's effectively total population capture.

CPRD

(Clinical Practice Research Datalink) collects anonymized primary-care electronic health records from NHS general practices in the UK — around 2,400 practices, roughly one in six Britons, with GOLD and Aurum products covering different practice software ecosystems. Records include diagnoses (Read codes in GOLD; SNOMED CT in Aurum), prescriptions issued by the GP, laboratory results, referrals, and immunizations, linkable to HES hospital admissions, ONS death registrations, and disease registries (cancer, congenital anomalies, pregnancy).

Why it matters for RWE

Universal NHS registration means near-complete primary-care capture regardless of payer or employment; patients rarely leave their GP practice, producing record spans measured in decades. Prescription issuance (rather than dispensing) plus complete problem documentation makes exposure and comorbidity ascertainment unusually reliable. Linkage to hospital and mortality data closes much of the out-of-system gap that plagues US EHR sources.

Operational characteristics

  • Acceptable-patient flag: practices report data quality ('up to standard'); analyses filter to acceptable patients and practices.
  • Prescriptions are issued, not dispensed: assume dispensing unless secondary evidence; duration comes from quantity/dose fields.
  • Coding systems: Read codes (GOLD) vs SNOMED CT (Aurum) — phenotype algorithms are not portable between them without mapping.
  • Linkage subsets: HES/death linkage covers roughly half of practices and requires separate approval; linkage availability shapes cohort definitions.

Common pitfalls

  • Private/secondary care prescribing: specialist-initiated drugs may not appear in primary-care issue records until GP uptake.
  • Consultation-driven recording: 'healthy user' gaps mean absence of codes during quiet periods is not confirmation of absence of disease.
  • Over-the-counter medications invisible: e.g., NSAIDs and aspirin bought OTC escape exposure measurement.
  • UK-specific practice patterns: dosing conventions, formularies, and screening programs limit direct transportability to other systems.

Pros, cons, and trade-offs

  • vs US claims: clinical depth (labs, BMI, smoking, lifestyle fields recorded in primary care) vs procedure/cost granularity.
  • vs US multi-system EHR (Truveta): population-based registration and decades-long spans vs larger scale but encounter-driven capture.
  • Trade-off: primary-care completeness vs specialist/oncology treatment detail that lives in hospital systems (partially recoverable via HES).

When NOT to use

Hospital-only acute outcomes without linkage approval; specialty biologics initiated exclusively in secondary care; US-generalizable cost estimates.

Decision diagram

flowchart LR
  GP[NHS general practices\ndiagnoses - scripts - labs] --> G[GOLD product]
  GP2[EMIS practices] --> A[Aurum product]
  G --> L[CPRD linkage service]
  A --> L
  H[HES admissions] --> L
  O[ONS deaths] --> L
  RG[Disease registries] --> L
  L --> S[Population-based RWE]
CPRD structure: GP records via GOLD/Aurum linked to HES, ONS, and registries.

Worked example

Scenario

Estimate stroke risk reduction for anticoagulation adherence in atrial fibrillation patients 2010-2020.

Dataset

Adherence-stratified stroke incidence in AF cohort.

adherence_groupnstroke_eventsir_per_100pyadj_hr
high_pdc_ge80 - 18400 - 412 - 1.21 - 1.00
low_pdc_lt80 - 15250 - 505 - 1.79 - 1.44

Steps

1Define AF cohort via Read/SNOMED code list with anticoagulant issue 2010-2020, age 40+.
2Compute proportion of days covered from prescription issue quantities and daily doses.
3Link HES strokes (I61,I63,I64) and ONS deaths; Poisson/Cox models adjusted for CHA2DS2-VASc components.
4Restrict to acceptable patients in linked practices; test Read-vs-SNOMED algorithm concordance if mixing GOLD/Aurum.

Result

Adjusted HR 1.44 (95% CI 1.25-1.66) for low vs high adherence; results stable across GOLD/Aurum implementations after code-list harmonization.

Trade-offs

vs. US commercial claims
Pros of this
—Clinical/lifestyle variables (BMI, smoking, labs); population-based registration; decades-long spans.
vs. Truveta style US EHR
Pros of this
—Universal-registration completeness vs encounter-driven capture; longer continuity.
vs. SNIIRAM/SNDS France
Pros of this
—Richer clinical detail from GP records vs comprehensive nationwide claims.

Runnable example

AF cohort construction with PDC adherence computation from issued prescriptions.

requires: pandas
\
import pandas as pd

def af_cohort(clinical, therapy, code_lists):
    af = clinical[clinical.medcode.isin(code_lists["af"])]
    index_dx = af.sort_values(["patid","eventdate"]).groupby("patid").first()
    ac = therapy[therapy.prodcode.isin(code_lists["anticoagulant"])]
    first_ac = ac.merge(index_dx.reset_index()[["patid","eventdate"]], on="patid")
    first_ac = first_ac[first_ac.eventdate_y <= first_ac.eventdate_x]
    return first_ac.sort_values(["patid","eventdate_x"]).groupby("patid").first()

def pdc(therapy_rows, start, end):
    # Sum covered days from qty/daily_dose, cap overlaps at calendar span
    span = (end - start).days
    covered = min((therapy_rows.qty / therapy_rows.daily_dose).sum(), span)
    return covered / span

Citations

FOUNDATIONAL / METHODS
  1. [1]Herrett E, et al. Data Resource Profile: Clinical Practice Research Datalink (CPRD). International Journal of Epidemiology. 2015.
  2. [2]Wolf A, et al. Data resource profile: Clinical Practice Research Datalink (CPRD) Aurum. International Journal of Epidemiology. 2019.
APPLIED EXAMPLES
  1. [3]NICE. Real-world evidence frameworks referencing CPRD for health technology evaluation.
REPORTING & GUIDANCE
  1. [4]Herrett E, et al. Data Resource Profile: Clinical Practice Research Datalink (CPRD). International Journal of Epidemiology. 2015.