Privacy AI
Regulatory

Health data compliance, handled.

Tell us what you need. We'll tell you what it takes and how long, before you commit.

Contact us

Summary

Contact us
A research conducted by Google DeepMind and numerous universities looks at how easily an outsider, without any prior knowledge of the data used to train a machine learning model, can get this information just by asking the model questions. We found that an adversary can pull out a lot of data, even gigabytes, from publicly available language models like Pythia or GPT-Neo, semi-public ones like LLaMA or Falcon, and private models like ChatGPT.

Key Findings:

- Vulnerabilities in Language Models: The research identifies vulnerabilities across various types of Language Models, ranging from open source (Pythia) to closed models (ChatGPT), and semi-open models (LLaMa). The vulnerabilities in semi-open and closed models are particularly concerning due to the non-public nature of their training data.

- Focus on Extractable Memorization: The study delves into the risks of extractable memorization, where data can be efficiently extracted from a machine learning model without prior knowledge of the training dataset.

- Enhanced Data Extraction Capabilities: The attack model developed by the researchers enables the extraction of training data at rates exceeding 150% compared to normal Language Model usage.

- Ineffectiveness of Data Deduplication: The research indicates that deduplication of training data does not significantly reduce the amount of data that can be extracted.

- Uncertainties in Data Handling: The study highlights ongoing uncertainties in how training data is processed and retained by Language Models.

Contact us

FAQs

Our frequently questions

No items found.

Seamus Larroque

CDPO / CPIM / ISO 27005 Certified

Find out how iliomad can help your company.

[Map placeholder]
Only visible in production
38.709099
-39.182035
1.6
6d17042a3425c5b3
Your message has been received!
We'll get back to you as soon as possible.
Something went wrong, please try again.
Home

Discover our latest articles

View All Blog Posts
Illustrated weekly digest header showing a digital shield, a DNA helix and a regulatory document, representing AI governance, clinical trial reform and healthcare data protection themes for September 2026
September 16, 2026
Regulations & Guidelines
Events
GDPR
Regulation
AI

iliomad Weekly Digest: AI Governance Failures, UK Regulatory Reform and Healthcare Cybersecurity Surge

This week: ICO becomes the Information Commission, Medicare AI prior-auth failures, FDA pilots accelerate trials, and ransomware hits 3.5 million patient records.

World map with data flow lines connecting clinical trial sites across continents, illustrating cross-border data transfer compliance in global studies
September 14, 2026
Clinical Trial Sponsor
DPIA
USA
GDPR
Regulations & Guidelines

Data transfer compliance in global clinical trials: a sponsor's guide

Understand data transfer compliance obligations for global clinical trials under GDPR, SCCs and local frameworks. A practical guide for sponsors from iliomad.

Illustration of the CNIL MR-001 framework applied to a clinical trial data compliance workflow in France, showing patient data flow and security controls
September 7, 2026
GDPR
Regulation
Guideline
Regulations & Guidelines

MR-001 CNIL: what clinical trial sponsors must know about French health data compliance

Understand MR-001 CNIL obligations for clinical trial sponsors in France, from Article 32 security requirements to breach notification and cross-border data transfers.