Term of the Day

Natural history study

A natural history study is a preplanned observational study intended to track the course of a disease over time, identifying demographic, genetic, environmental and other variables that correlate with its development and outcomes in the absence of intervention, or under standard of care. Designs may be retrospective (chart review of existing records) or prospective (longitudinal follow-up of a cohort or registry).

Natural history data is particularly important in rare and paediatric diseases, where randomised placebo-controlled trials may be infeasible or unethical. The FDA (guidance on rare disease natural history studies, 2019) and the EMA accept well-designed natural history studies to define endpoints and biomarkers, identify patient subgroups, estimate sample sizes and, in some cases, serve as external or historical control arms for single-arm trials supporting orphan products.

Because they are non-interventional, natural history studies fall outside the CTR and are governed by national law (for example France's MR-003 or MR-004 reference methodologies) and by the GDPR. They typically involve secondary use of medical records, long-term follow-up, genetic data and small populations in which anonymisation is rarely achievable, so pseudonymisation, a DPIA and a robust research legal basis under Art. 9(2)(j) are essential. Registries maintained by patient organisations or academic consortia raise additional questions of joint controllership and data access governance.

Q

Query (clinical data query)

A query, in clinical data management, is a formal request for clarification or correction issued when a data point recorded in the case report form appears missing, inconsistent, out of range or implausible. Queries are generated automatically by edit checks programmed in the EDC system (for example a visit date before the enrolment date, or a laboratory value outside physiological limits), manually by data managers or medical reviewers during data review, or by monitors during source data verification. The site responds by confirming, correcting or explaining the entry; every step is recorded in the audit trail with user, timestamp and reason for change, as required by 21 CFR Part 11 and the EMA guideline on computerised systems.

Query management is a core process of the data management plan: it defines query text conventions, response timelines (often 5 to 10 working days), escalation, closure rules and metrics such as query rate per subject and ageing. High query volumes signal poor eCRF design, unclear protocol definitions or site training needs, and unresolved queries block database lock. Under ICH E6(R3) risk-based approaches, queries are concentrated on critical data and processes rather than on every data point.

Queries are a frequently overlooked channel for personal data leakage. Free-text query responses written by site staff sometimes include the participant's name, initials, hospital number or clinical details that go beyond the data point in question, which breaks the pseudonymisation boundary between site and sponsor. Good practice is to train site staff and data managers to write queries and responses that refer only to the subject number and the field concerned, to configure the EDC so that query text cannot be exported without review, and to include query text in periodic data reviews for identifying information. Query metadata also constitutes personal data about site staff and data managers, covered by the site-staff information notice.