Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Term of the Day

Natural history study

A natural history study is a preplanned observational study intended to track the course of a disease over time, identifying demographic, genetic, environmental and other variables that correlate with its development and outcomes in the absence of intervention, or under standard of care. Designs may be retrospective (chart review of existing records) or prospective (longitudinal follow-up of a cohort or registry).

Natural history data is particularly important in rare and paediatric diseases, where randomised placebo-controlled trials may be infeasible or unethical. The FDA (guidance on rare disease natural history studies, 2019) and the EMA accept well-designed natural history studies to define endpoints and biomarkers, identify patient subgroups, estimate sample sizes and, in some cases, serve as external or historical control arms for single-arm trials supporting orphan products.

Because they are non-interventional, natural history studies fall outside the CTR and are governed by national law (for example France's MR-003 or MR-004 reference methodologies) and by the GDPR. They typically involve secondary use of medical records, long-term follow-up, genetic data and small populations in which anonymisation is rarely achievable, so pseudonymisation, a DPIA and a robust research legal basis under Art. 9(2)(j) are essential. Registries maintained by patient organisations or academic consortia raise additional questions of joint controllership and data access governance.

D

De-identification (HIPAA)

De-identification is the process defined in the US HIPAA Privacy Rule (45 CFR 164.514(a) to (c)) by which protected health information is rendered not individually identifiable, so that it ceases to be PHI and may be used and disclosed without the Privacy Rule's restrictions. Two methods are recognised. Under the Safe Harbor method, all 18 listed identifiers of the individual and of relatives, employers and household members are removed, and the covered entity has no actual knowledge that the remaining information could identify the individual. Under the Expert Determination method, a qualified statistician or scientist applies generally accepted principles and determines, and documents, that the risk of re-identification by an anticipated recipient is very small.

The Privacy Rule also permits the covered entity to assign a re-identification code to de-identified data, provided the code is not derived from information about the individual and the mechanism for re-identification is not disclosed (45 CFR 164.514(c)); this is why key-coded clinical trial data can be treated as de-identified in the US when the recipient has no access to the key. A related but distinct concept is the limited data set (45 CFR 164.514(e)), which retains dates and some geographic data and may be shared for research, public health and healthcare operations under a data use agreement. HHS guidance of 2012 explains both methods in detail.

De-identification must not be equated with anonymisation under the GDPR. The GDPR test is contextual and considers all means reasonably likely to be used by anyone, not only the anticipated recipient, so a Safe Harbor dataset that retains year of birth, three-digit ZIP codes and rich clinical detail will often remain personal data in the EU, particularly for rare diseases; and key-coded data that HIPAA deems de-identified is pseudonymised personal data under the GDPR because the key exists. Expert Determination is methodologically closer to the EDPB approach, and its documented risk analysis can support a GDPR anonymisation assessment, but the conclusion must be re-evaluated under EU criteria. In transatlantic data-sharing agreements, iliomad recommends stating both the HIPAA status and the GDPR status of each dataset explicitly rather than using "de-identified" as a shorthand for "not regulated".