Skip to content
PlagiatScanner.de
Research data · DAT-14

Anonymising sensitive research data: limits and evidence

Do more than remove names. Examine direct and indirect attributes whose combination could single someone out, then assess the dataset against plausible external knowledge and its intended audience. Choose controls such as generalisation, suppression, aggregation or restricted access, and record both the transformations and residual risk. Pseudonymised data is not automatically anonymous: a person may remain linkable through separately held information.

Risk and controls matrixGeneral information, not legal advice

Removing identifiers is only the first pass

Names, email addresses and participant numbers are direct identifiers, but they are rarely the whole problem. Age, occupation, location, diagnosis, an exact interview date or a rare event can become identifying when combined. In transcripts, a job title, personal history and local reference may reveal more than the name field. Faces, voices, coordinates and technical metadata can carry identity in image, audio and geospatial collections.

Under the GDPR, pseudonymisation means data cannot be attributed to a person without additional information that is kept separately and protected. The resulting data remains personal data. Anonymisation aims at a context in which the person is no longer identifiable by means reasonably likely to be used. That judgement is contextual and can change as new linking data or techniques become available; it should not be treated as an everlasting badge.

Assess the environment as well as the file

Build a field-level inventory. Mark direct identifiers, quasi-identifiers, special-category data, free text and technical metadata. Next ask what outside information a plausible recipient could obtain. A rare occupation might be harmless in a national table but obvious inside a small organisation. Staff pages, professional profiles, news reports and public registers can turn an otherwise vague combination into a recognisable individual.

Audience matters. An open download, an accredited data centre and a small supervised examination panel create different opportunities for linkage. Time matters too: a dataset that appears low-risk today may become linkable later. State the dataset version, intended setting and assessment date, and schedule a fresh review before expanding access or depositing a revised file.

Balance disclosure control against research value

Suppression removes individual values or rare records. Generalisation replaces an exact age with a band or a small place with a wider region. Aggregation releases group summaries rather than person-level rows. Carefully redacting or paraphrasing transcript passages can reduce biographical exposure. Every measure changes what the data can support.

Test analytical damage alongside privacy risk. Removing every rare group may make the study unusable for precisely those participants. Record the fields affected, rule used and effect on key analyses. Synthetic data may demonstrate a workflow, but it does not automatically replace the source and can inherit disclosure risks from how it was produced. Where open release would require destructive alteration, controlled access may preserve more scientific value with stronger safeguards.

Authority asset: disclosure risk and controls matrix

Connect an identification signal to a proportionate control
SignalQuestionPossible controlEvidence to retain
Direct identifierIs the field needed for analysis?Remove it or hold a key separatelyField inventory and key controls
Rare combinationCould a small group be singled out?Band, pool or suppress valuesBefore-and-after frequencies
Narrative textDoes context reveal a person or site?Redact, paraphrase or restrict accessEditorial decision log
Location or timestampCan precision support external linkage?Reduce spatial or temporal detailTransformation rule
Severe loss of utilityWould open release destroy the analysis?Use controlled rather than open accessAccess conditions and review route

Add an owner, review date and accepted residual risk to every row. This is a decision record, not an automatic declaration that a file is anonymous.

Create an auditable transformation record

Keep the untouched source in an appropriately protected location and work on a versioned derivative. The record should name the source field, transformation, reason, affected cases and resulting field. Use reproducible code for repeatable structured changes. For manual qualitative redaction, define categories of intervention and have borderline passages reviewed without distributing sensitive material more widely than necessary.

Inspect the exported artefact itself. Search for residual names and identifiers; check document properties, comments, tracked changes, embedded thumbnails and attachments. Hidden spreadsheet columns may survive an apparently clean view. Separate files may contain compatible IDs that recreate a link. An anonymisation decision must cover the release package as a whole rather than a single visible worksheet.

Match the release route to consent and governance

Compare proposed use with participant information, consent, ethics approval, contracts and institutional policies. If open reuse was not covered, or residual risk remains excessive, define access tiers. Public metadata can make a resource discoverable while a data access committee assesses reasoned applications for detailed files. Terms may prohibit re-identification and onward sharing, but rules complement rather than replace technical minimisation.

High-risk processing may require formal assessment and engagement with a data protection officer or ethics body. Start before collection, when fields and consent can still be redesigned. Document who can access the re-identification key, how access is logged and when it will be destroyed or retained. Our guide to planning consent for data reuse addresses this earlier stage.

Make bounded claims rather than absolute promises

Avoid saying that identification is impossible. Describe the released version, method, assumed external information, audience and remaining uncertainty. State which variables were removed, banded or protected through access controls. A conclusion reached for a secure research environment does not automatically apply to public download. Re-run the assessment whenever the dataset, audience or linking environment changes.

Official and institutional sources

  1. General Data Protection Regulation – definitions, processing principles and safeguards for research.
  2. Information Commissioner's Office: anonymisation guidance – official UK guidance on anonymisation, pseudonymisation and identifiability.
  3. UK Data Service: anonymisation – institutional practice for quantitative and qualitative data.