Mynd Healthcare

Home / Safety, Ethics & Governance

De-identification and Re-identification Risk

Removing identifying details can reduce privacy risk, but remaining information may still allow a person to be recognised.

#Removing names is only one step

De-identification means changing information to reduce the possibility of linking it to a person. Removing names, addresses and record numbers is a common starting point. However, other details may still identify someone when combined, especially in small communities or datasets containing unusual events or conditions.

Dates, locations, occupations and detailed medical histories can act as indirect identifiers. Free-text notes, images and file metadata may also contain identifying information. A dataset that appears anonymous to one reader may be recognisable to someone with additional knowledge, such as a colleague, neighbour or family member.

#Risk depends on context

Re-identification can occur when records are linked with other available information. The risk depends on the data, who can access it, what other sources exist and how much effort identification would require. Public release creates different risks from access within a tightly controlled research environment.

Replacing a name with a code is often called pseudonymisation. If a separate key allows the code to be linked back to a person, the information remains linkable. Legal definitions differ, but pseudonymised information commonly remains subject to personal-data obligations rather than becoming unrestricted anonymous data.

#Reducing and reviewing risk

Risk reduction may involve grouping ages, reducing location detail, suppressing rare combinations or limiting access. Each approach has trade-offs: removing detail may make data less useful or hide differences between groups. The aim is a justified balance, not an unsupported promise that identification is impossible.

Assessment should consider the intended use, likely recipients and consequences if someone is identified. Agreements, secure environments and restrictions on onward sharing can complement technical changes. Risk may need reassessment as new information becomes available, because a release considered low risk today may become easier to link later.

#Common misunderstandings

De-identification does not necessarily make information anonymous. Removing a name or contact details can reduce risk, but combinations of remaining details, such as age, location, occupation and an unusual medical history, may still point to someone.

Replacing names with codes is also not the same as making identification impossible. If a separate key connects those codes to individuals, someone with access to that key may be able to reconnect records. Even without a key, matching details against other available information can sometimes reveal an identity.

Another misunderstanding is that information is either completely safe or completely unsafe. Risk varies with what a record contains, who can access it and what other information they hold. A dataset suitable for restricted research access may not be suitable for public release.

Finally, de-identification is not a permanent guarantee. New information or ways of linking records can change the risk after data has been shared.

#Questions worth asking a clinician

  • Which identifying details will be removed from my health information, and which details will remain?
  • Could someone recognise me from my diagnosis, age, location, or treatment dates after my name is removed?
  • Could my de-identified health information be matched with public records or other datasets to identify me?
  • Who can access my de-identified health information, and what safeguards prevent them from trying to identify me?
  • What steps would you take if someone identified me from my de-identified health information?