Key-coded data (coded data)
Key-coded data, also called coded data or single- or double-coded data, is the term used in clinical research for personal data in which direct identifiers such as name, address and hospital number have been replaced by a code (the subject number), with the correspondence between code and identity, the key, held separately by a defined party, typically the investigator at the site. Double coding adds a second layer, with a trusted third party holding the link between the first and second codes, and is used in biobanks and genomic research to separate the sponsor further from identities.
Key-coded data is the standard form in which trial data reaches the sponsor, its CRO and vendors, and in which it is submitted to regulators in SDTM and ADaM format. ICH GCP requires that the confidentiality of records identifying subjects be protected and that the sponsor's data be coded, and the EMA and FDA use the term "key-coded" in their own guidance on data sharing and publication.
Under the GDPR, key-coded data is pseudonymised data and therefore remains personal data in the hands of the sponsor, even though the sponsor cannot itself re-identify participants, because re-identification through the site is reasonably possible (Recital 26, EDPB Guidelines 01/2025). The consequences are that all GDPR obligations apply to the sponsor as controller of the coded dataset, that sharing it outside the EEA is an international transfer, and that it cannot be described as "anonymous" in informed consent forms, clinical trial agreements or data-sharing contracts. Under US HIPAA, by contrast, coded data can qualify as de-identified if the code is not derived from identifiers and the recipient has no access to the key (45 CFR 164.514(c)), which explains why US sponsors often treat coded EU data as non-personal, one of the most common transatlantic misunderstandings iliomad corrects in contract reviews.
