Anonymization & Pseudonymization Methodology
When to anonymize versus pseudonymize, which technique to use depending on the case, and how to actually verify the result protects people — not just that it "looks" anonymous.
What's included
- PDF with the anonymize vs. pseudonymize decision tree by use case, and the main techniques (generalization, suppression, statistical noise, tokenization)
- Excel with a verification checklist: how to check the anonymization withstands the 3 re-identification risks (singling out, linkability, inference)
- Legal Notice
Why this document exists
"Anonymized" is one of the most misused words in data protection — many companies remove the name and email from a dataset and call it anonymous, when in reality cross-referencing it with another source is enough to re-identify the person (postal code + birth date + gender already identifies most of the population). This methodology gives a concrete process for choosing the right technique and, above all, verifying it — instead of assuming "removing the name" is enough.
Frequently asked questions
What's the difference between anonymizing and pseudonymizing?
Anonymized data can no longer be linked back to a person by any reasonable means — it stops being personal data under GDPR. Pseudonymized data replaces direct identifiers with a code, but remains personal data.
Why does AI training data make this especially relevant?
A model trained on poorly anonymized data can memorize and leak identifiable information in its responses. Truly verifying anonymization before training reduces that risk.
What format is it delivered in?
PDF with the methodology and decision tree, plus Excel with a verification checklist by technique applied.