Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Term of the Day

Natural history study

A natural history study is a preplanned observational study intended to track the course of a disease over time, identifying demographic, genetic, environmental and other variables that correlate with its development and outcomes in the absence of intervention, or under standard of care. Designs may be retrospective (chart review of existing records) or prospective (longitudinal follow-up of a cohort or registry).

Natural history data is particularly important in rare and paediatric diseases, where randomised placebo-controlled trials may be infeasible or unethical. The FDA (guidance on rare disease natural history studies, 2019) and the EMA accept well-designed natural history studies to define endpoints and biomarkers, identify patient subgroups, estimate sample sizes and, in some cases, serve as external or historical control arms for single-arm trials supporting orphan products.

Because they are non-interventional, natural history studies fall outside the CTR and are governed by national law (for example France's MR-003 or MR-004 reference methodologies) and by the GDPR. They typically involve secondary use of medical records, long-term follow-up, genetic data and small populations in which anonymisation is rarely achievable, so pseudonymisation, a DPIA and a robust research legal basis under Art. 9(2)(j) are essential. Registries maintained by patient organisations or academic consortia raise additional questions of joint controllership and data access governance.

L

Large language model (LLM)

A large language model (LLM) is a neural network, typically based on the transformer architecture, trained on very large corpora of text (and increasingly images, audio and code) to predict and generate language. Through pre-training on broad data and subsequent fine-tuning and alignment, LLMs acquire the ability to summarise, translate, answer questions, write code and hold dialogue; models such as GPT, Claude, Gemini, Llama and Mistral power chat assistants, copilots and an expanding range of enterprise applications, often augmented with retrieval from an organisation's own documents.

In life sciences, LLMs are used for medical writing and regulatory document drafting, literature screening and systematic reviews, pharmacovigilance case intake and coding, medical information responses, protocol and informed consent drafting, clinical trial matching, coding of electronic health records, and patient-facing chatbots. The EMA reflection paper on AI in the medicinal product lifecycle (2024) and FDA discussion papers set expectations on validation, transparency and human accountability, and applications that support diagnosis or treatment may qualify as software as a medical device.

Under the EU AI Act, the underlying models are general-purpose AI models whose providers carry documentation, copyright and, for the largest models, systemic-risk obligations; systems built on them may be high-risk depending on use, and chatbots must disclose that users interact with an AI (Art. 50). Under the GDPR, the EDPB Opinion 28/2024 confirms that models trained on personal data may themselves contain personal data, that legitimate interest can be a valid basis for training subject to a strict three-step test, and that unlawful training can taint downstream use. Operational issues for companies include: preventing staff from entering patient or health data into public tools; vendor contracts covering processor terms, non-use for training and transfers to US-hosted providers; DPIAs for internal deployments; hallucination and bias controls; and human review of outputs that inform decisions about individuals. An enterprise generative AI policy and an approved-tool list are now standard elements of an iliomad AI compliance programme.