← Methods repository
CONCEPTINTERMEDIATEPYTHON · R · SAS5 citations

eSource, EDC, and Source Data Verification

The regulatory and operational discipline for defining electronic source data, capturing or transferring those data into electronic data capture systems, preserving audit trails and data originator metadata, and verifying critical data against source records using risk-based source data verification rather than blind 100 percent checking.

Data Quality Assessmentesourceedcelectronic-data-capturesource-data-verificationsdvrisk-based-monitoringaudit-traildata-integrity
On this page
Methods reference only. Use primary source citations and local policy before applying this in a study protocol, regulatory submission, payer dossier, or clinical decision.
In plain language

eSource means the original study data are electronic, EDC is the study database that stores case report form data, and source data verification checks important EDC fields against the original source. The point is traceability: a reviewer should be able to see where a value came from, who or what created it, when it changed, and whether critical fields were checked. More checking is not always better; risk-based verification targets the data that can change safety, eligibility, endpoints, or the study conclusion.

When to use it
When the original data are already electronic and the source system can preserve traceability.
Most large clinical investigations, pragmatic trials, registries, and RWE-adjacent primary data studies.
Regulatory-grade or HTA-facing studies where source reliability and data lineage must be defended.
Watch out for
Requires validated systems, interface governance, access controls, audit trails, and source-system documentation.
Requires prospective risk assessment, central monitoring, and defensible sampling rules.
Adds operational burden and may require access to systems a data vendor does not normally expose.

eSource, electronic data capture (EDC), and source data verification (SDV)

describe the chain from original clinical investigation data to the analyzable trial or registry dataset. eSource is source data initially recorded in electronic form or transferred directly from an electronic source. EDC is the system that stores case report form data and study-specific fields. SDV is the monitoring activity that checks selected EDC data against the source record, such as an EHR note, lab result, device file, imaging report, ePRO entry, or the EDC itself when the eCRF is the source.

In RWE-adjacent primary research, this chain matters because data increasingly originate outside paper charts: EHR extracts, direct EHR-to-EDC transfer, central labs, imaging vendors, wearables, ePRO apps, registries, claims extracts, and home devices. Each source can be valid, but only if the protocol and data-management plan identify the authorized source data originator, the source location, the data element identifier, the audit trail, and the reconciliation path into the analysis dataset. A clean-looking EDC field is not evidence of quality unless its source and transformations are traceable.

Core conceptual distinction

eSource is about where the original data live and how they are captured. EDC is about where study data are stored and curated. SDV is about how selected values are checked against their source. These are often collapsed into one operational phrase, but they answer different questions. "The value is in EDC" does not tell you whether it came from an EHR interface, manual transcription, a device vendor, or the subject's direct entry. "The monitor verified it" does not tell you whether the right records were selected, whether the source itself had an audit trail, or whether the check targeted variables that matter to the estimand and safety.

The modern quality question is not "did we verify every field?" It is "did we preserve reliability, integrity, traceability, and fitness for purpose for the data elements that can change the study conclusion or participant safety?" FDA's eSource guidance emphasizes authorized source data originators, data element identifiers, audit trails, capture into the eCRF, investigator responsibilities, and computerized systems. FDA's electronic systems/records/signatures guidance frames the broader requirement that electronic records be trustworthy, reliable, and generally equivalent to paper records.

Pros, cons, and trade-offs

  • vs paper source plus manual EDC transcription: eSource reduces transcription, enables near-real-time review, and preserves metadata. Cost: interfaces, system validation, access controls, audit trails, and vendor governance become part of the evidence package. Prefer eSource when data already exist electronically and the source system can preserve traceability.
  • vs EDC-as-source: Direct entry into an eCRF can be efficient when the eCRF is genuinely the first place the observation is recorded. Cost: the investigator must be able to corroborate or explain the value, and the eCRF audit trail becomes the source audit trail. Do not use EDC-as-source to hide missing clinical documentation.
  • vs 100 percent SDV: Risk-based SDV focuses monitoring on critical variables, high-risk sites, unexpected patterns, and safety/endpoint data. Cost: it requires prospective risk assessment and central monitoring analytics. Prefer risk-based SDV for large pragmatic or registry studies; reserve 100 percent SDV for small, high-risk, or highly manual settings where every field is critical.
  • vs secondary RWD extracts: Claims, EHR, and registry extracts often arrive as curated datasets rather than study EDC. The eSource discipline still applies: document provenance, extraction logic, data transformations, audit trail, refresh date, and reconciliation to source when challenged.

When NOT to use - and when it is actively misleading

Do not treat SDV as a substitute for data quality by design. Verifying an incorrect or ambiguous source record only proves the EDC copied the ambiguity. Do not claim eSource traceability if the upstream device, app, spreadsheet, or EHR extract lacks an audit trail or version history. Do not use 100 percent SDV of noncritical fields as theater while ignoring endpoint algorithms, consent status, eligibility, randomization, safety events, and key covariates. Do not call an EHR extract source-verified unless the extraction logic, source tables, transformation rules, and sampled record checks are documented. It is actively misleading to report "source verified" without stating which variables, what source, what sampling fraction, what discrepancy threshold, and what corrective action process were used.

Data-source operational depth

  • EHR eSource: High value for labs, vitals, medications, diagnoses, encounters, and notes, especially in EHR-embedded trials. Failure modes are local build differences, flowsheet reuse, late-entered notes, copied-forward text, interface mapping errors, and source table changes. Preserve source system, table/field, extraction date, transformation code, and investigator review status.
  • EDC direct entry: Appropriate when the eCRF is the first capture point, such as a study-specific assessment performed only for the protocol. Failure modes are missing corroboration, user-role confusion, late corrections, and inadequate audit-trail review. The data originator and timestamp must be retained.
  • Vendor DHT/ePRO/lab feeds: Useful for high-frequency or central measurements, but the source may be outside the EDC. Failure modes include file reprocessing, algorithm updates, timezone changes, missing device metadata, and vendor-side corrections. Require data transfer specifications and retained raw/source files.
  • Claims/registry extracts: In pragmatic trials and RWE studies, source verification is usually not field-by-field chart comparison. It is provenance and extract verification: confirm the data supplier, extract criteria, code lists, refresh date, record counts, linkage, and sampled source-to-extract concordance for critical elements.

Worked example

A pragmatic oncology registry trial collects baseline ECOG performance status, randomization arm, grade >=3 adverse events, and progression date. ECOG is entered directly into EDC by the clinician during the visit and is therefore EDC-as-source. Randomization comes from the EHR trial module. Lab-based safety events come from an EHR interface. Progression date comes from radiology report abstraction. A risk-based SDV plan selects 100 percent verification for informed consent, eligibility, randomization, death, progression, and grade >=3 adverse events; 20 percent verification for key baseline covariates; and central monitoring for all sites. The monitor finds that one site has high EHR-to-EDC lab discrepancy rates because local units changed after an interface update. The fix is not to verify more manually forever; it is to correct the interface mapping, reprocess affected records, document the audit trail, and re-run discrepancy checks.

Decision diagram

flowchart TD
  Source[Authorized source data originator<br/>EHR, lab, device, ePRO, EDC-as-source] --> Capture[Capture or transfer into EDC]
  Capture --> Audit[Audit trail + data element identifier]
  Audit --> DM[Data management checks and queries]
  DM --> SDV[Risk-based SDV of critical fields]
  SDV --> Root[Root-cause correction<br/>interface, source, transcription, training]
  Root --> Analysis[Traceable analysis dataset]
eSource quality depends on source-originator definition, audit trail, EDC capture, targeted verification, root-cause correction, and traceable analysis data.

Worked example

Scenario

An oncology registry trial uses several electronic sources. The monitor checks whether critical EDC values match their authorized sources and whether discrepancies trigger corrective action.

Dataset

Simplified source-to-EDC verification sample.

subject_iddata_elementsource_systemsource_valueedc_valuecriticaldiscrepancy
S001randomization_armEHR trial moduleArm AArm ATrueFalse
S002ECOGEDC direct entry11TrueFalse
S003grade3_neutropeniaEHR lab interfaceTrueFalseTrueTrue
S004smoking_statusEHR abstractionformerunknownFalseTrue

Steps

1Identify the authorized source for each data element before verification starts.
2Verify all critical fields or use a pre-specified high sampling fraction for them.
3Separate critical discrepancies from noncritical discrepancies because they carry different corrective actions.
4Investigate clustered discrepancies by site, source system, interface, or calendar time.
5Correct the root cause, not only the visible EDC value.

Result

S003 is a critical discrepancy that can affect safety analysis and must be queried and root-caused. S004 is still a data-quality issue, but it is lower priority because it is noncritical to the primary endpoint and safety rules.

Trade-offs

vs. Paper source plus manual EDC transcription
Pros of this
Reduces transcription, speeds review, preserves metadata, and supports remote monitoring.
vs. 100 percent source data verification
Pros of this
Targets variables and sites most likely to affect participant safety or study conclusions; scales better for pragmatic and registry trials.
vs. Secondary RWD extract without source verification
Pros of this
Creates audit-ready provenance, traceability, and sampled concordance evidence for critical elements.

Runnable example

Risk-based SDV discrepancy profile from a sampled source-to-EDC comparison. Inputs: checks : subject_id, site_id, data_element, source_system, source_value, edc_value, critical (bool) Returns discrepancy rates by element/site and a critical-field query list.

requires: pandas
import pandas as pd

def sdv_profile(checks):
    df = checks.copy()
    df["source_norm"] = df["source_value"].astype(str).str.strip().str.upper()
    df["edc_norm"] = df["edc_value"].astype(str).str.strip().str.upper()
    df["discrepancy"] = df["source_norm"] != df["edc_norm"]

    by_element = (df.groupby(["data_element", "critical"])
                    .agg(n_checked=("subject_id", "size"),
                         discrepancies=("discrepancy", "sum"))
                    .reset_index())
    by_element["discrepancy_rate"] = by_element["discrepancies"] / by_element["n_checked"]

    by_site = (df.groupby("site_id")
                 .agg(n_checked=("subject_id", "size"),
                      discrepancies=("discrepancy", "sum"),
                      critical_discrepancies=("discrepancy", lambda x: int((x & df.loc[x.index, "critical"]).sum())))
                 .reset_index())
    by_site["discrepancy_rate"] = by_site["discrepancies"] / by_site["n_checked"]

    queries = df[df["discrepancy"] & df["critical"]].copy()
    return {"by_element": by_element, "by_site": by_site, "critical_queries": queries}

Citations

FOUNDATIONAL / METHODS
  1. [1]U.S. Food and Drug Administration. Electronic Source Data in Clinical Investigations. Guidance for Industry. September 2013.
  2. [2]U.S. Food and Drug Administration. Electronic Systems, Electronic Records, and Electronic Signatures in Clinical Investigations: Questions and Answers. Guidance for Industry and Other Interested Parties. 2024.
  3. [3]Sheetz N, Wilson B, Benedict J, et al. Evaluating Source Data Verification as a Quality Control Measure in Clinical Trials. Therapeutic Innovation & Regulatory Science. 2014;48(6):671-680.
APPLIED EXAMPLES
  1. [4]Yamada O, Chiu SW, Takata M, et al. Clinical trial monitoring effectiveness: Remote risk-based monitoring versus on-site monitoring with 100% source data verification. Clinical Trials. 2021;18(2):158-167.
REPORTING & GUIDANCE
  1. [5]U.S. Food and Drug Administration. Digital Health Technologies for Remote Data Acquisition in Clinical Investigations. Guidance for Industry, Investigators, and Other Stakeholders. December 2023.