Anonymization & Pseudonymization Methodology
When to anonymize versus pseudonymize, which technique to use depending on the case, and how to actually verify the result protects people — not just that it "looks" anonymous.
What's included
- PDF with the anonymize vs. pseudonymize decision tree by use case, and the main techniques (generalization, suppression, statistical noise, tokenization)
- Excel with a verification checklist: how to check the anonymization withstands the 3 re-identification risks (singling out, linkability, inference)
- Legal Notice

Why this document exists
"Anonymized" is one of the most misused words in data protection — many companies remove the name and email from a dataset and call it anonymous, when in reality cross-referencing it with another source is enough to re-identify the person (postal code + birth date + gender already identifies most of the population). This methodology gives a concrete process for choosing the right technique and, above all, verifying it — instead of assuming "removing the name" is enough.
Frequently asked questions
What's the difference between anonymizing and pseudonymizing?
Anonymized data can no longer be linked back to a person by any reasonable means — it stops being personal data under GDPR. Pseudonymized data replaces direct identifiers with a code, but remains personal data.
Why does AI training data make this especially relevant?
A model trained on poorly anonymized data can memorize and leak identifiable information in its responses. Truly verifying anonymization before training reduces that risk.
Does this replace legal advice?
No. It's a technical working methodology, but it doesn't constitute legal advice or guarantee GDPR or other regulatory compliance — see the included Legal Notice.
What format is it delivered in?
PDF with the methodology and decision tree, plus Excel with a verification checklist by technique applied.
