Term of the Day

Natural history study

A natural history study is a preplanned observational study intended to track the course of a disease over time, identifying demographic, genetic, environmental and other variables that correlate with its development and outcomes in the absence of intervention, or under standard of care. Designs may be retrospective (chart review of existing records) or prospective (longitudinal follow-up of a cohort or registry).

Natural history data is particularly important in rare and paediatric diseases, where randomised placebo-controlled trials may be infeasible or unethical. The FDA (guidance on rare disease natural history studies, 2019) and the EMA accept well-designed natural history studies to define endpoints and biomarkers, identify patient subgroups, estimate sample sizes and, in some cases, serve as external or historical control arms for single-arm trials supporting orphan products.

Because they are non-interventional, natural history studies fall outside the CTR and are governed by national law (for example France's MR-003 or MR-004 reference methodologies) and by the GDPR. They typically involve secondary use of medical records, long-term follow-up, genetic data and small populations in which anonymisation is rarely achievable, so pseudonymisation, a DPIA and a robust research legal basis under Art. 9(2)(j) are essential. Registries maintained by patient organisations or academic consortia raise additional questions of joint controllership and data access governance.

K

Key-coded data (coded data)

Key-coded data, also called coded data or single- or double-coded data, is the term used in clinical research for personal data in which direct identifiers such as name, address and hospital number have been replaced by a code (the subject number), with the correspondence between code and identity, the key, held separately by a defined party, typically the investigator at the site. Double coding adds a second layer, with a trusted third party holding the link between the first and second codes, and is used in biobanks and genomic research to separate the sponsor further from identities.

Key-coded data is the standard form in which trial data reaches the sponsor, its CRO and vendors, and in which it is submitted to regulators in SDTM and ADaM format. ICH GCP requires that the confidentiality of records identifying subjects be protected and that the sponsor's data be coded, and the EMA and FDA use the term "key-coded" in their own guidance on data sharing and publication.

Under the GDPR, key-coded data is pseudonymised data and therefore remains personal data in the hands of the sponsor, even though the sponsor cannot itself re-identify participants, because re-identification through the site is reasonably possible (Recital 26, EDPB Guidelines 01/2025). The consequences are that all GDPR obligations apply to the sponsor as controller of the coded dataset, that sharing it outside the EEA is an international transfer, and that it cannot be described as "anonymous" in informed consent forms, clinical trial agreements or data-sharing contracts. Under US HIPAA, by contrast, coded data can qualify as de-identified if the code is not derived from identifiers and the recipient has no access to the key (45 CFR 164.514(c)), which explains why US sponsors often treat coded EU data as non-personal, one of the most common transatlantic misunderstandings iliomad corrects in contract reviews.