De-identification is the work of removing or masking the details that point to a person, such as names, emails and account numbers. Anonymization is the result you reach when nobody can reasonably link a record back to that person, even by combining it with other data.
Every anonymized dataset has been de-identified, but plenty of de-identified datasets are still personal data in the eyes of the law. The difference decides which rules apply when you license data to an AI company. This page is general information, not legal advice, so check your own case with a privacy lawyer.
Three words that get mixed up
| Term | What was done | Can the person be found again? | Usual legal status |
|---|---|---|---|
| Pseudonymization | Names and IDs replaced with codes; a key links codes back to people | Yes, by anyone holding the key | Still personal data under GDPR |
| De-identification | Direct identifiers removed or masked; risky values generalized | Depends on what's left and who might look | Depends on the law and the safeguards |
| Anonymization | Enough removed that identification is no longer reasonably likely | No, by any means reasonably likely to be used | Outside GDPR; outside CCPA if the conditions are met |
Re-identification runs the other way: tracing records back to the people in them. A buyer who tries it breaks our license.
Why removing names isn't enough
A record can identify someone without a name in it. The fields that do this are called quasi-identifiers: ordinary details that become unique in combination.
The research on this is old and consistent:
- Latanya Sweeney found that 87% of the US population was likely unique on just three fields: five-digit ZIP code, gender and date of birth.
- A 2019 study in Nature Communications estimated that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes.
- Researchers linked records in the Netflix Prize dataset, about 100 million ratings from roughly 480,000 subscribers with names removed, to public IMDb ratings and identified individual subscribers.
Business data has its own quasi-identifiers. A company domain plus a job title can point to one person. An unusual invoice amount on a known date can point to one client. Free-text notes often contain names nobody thought to remove.
What the GDPR says
The GDPR covers personal data, meaning information about an identified or identifiable person. Recital 26 says the regulation does not apply to anonymous information. The test for anonymity looks at all the means reasonably likely to be used to identify someone, by you or by anyone else, taking into account cost, time and available technology.
Pseudonymized data stays inside the GDPR. Article 4(5) defines pseudonymisation as processing that stops data being linked to a person without additional information kept separately. The European Data Protection Board's 2025 guidelines confirm that pseudonymized data is still personal data. They also say pseudonymization reduces risk and can help support a legitimate interest or a compatible new purpose.
The EDPB has applied the same thinking to AI. In Opinion 28/2024, it said a model trained on personal data is not automatically anonymous, and anonymity has to be assessed case by case.
What US law says
US laws mostly use the word "deidentified" and attach conditions to it.
Under the California Consumer Privacy Act, deidentified information must be impossible to reasonably link to a consumer, and the business holding it must do three things: take reasonable measures so it can't be associated with a consumer or household, publicly commit not to re-identify it, and contractually bind every recipient to the same. Miss one and the data is still personal information. Our CCPA guide covers what that means for a sale.
Health data under HIPAA has two routes, described by HHS. Safe Harbor removes 18 listed types of identifiers. Expert Determination has a qualified expert conclude that the risk of identification is very small and document how they reached that view.
How de-identification is done
Good de-identification uses several techniques together:
- Remove direct identifiers. Names, emails, phone numbers, account numbers, street addresses.
- Replace with stable pseudonyms where structure matters. A client becomes
ORG_7F3Aeverywhere it appears, so a model can still see that one client paid late three times. - Generalize. Exact amounts become ranges, exact dates become months or quarters, job titles become broader roles. See generalization.
- Suppress. Drop the rare records or fields that stay unique after generalizing. See suppression.
- Drop or review free text. Notes fields hide names and details that column rules miss.
- Test. Check whether any record can still be singled out from a combination of fields. k-anonymity is one common test.
The pseudonym key never goes to the buyer. If it did, the data would be plainly re-identifiable.
How we handle it
Our process aims for data that can't be linked back to a person or a client, and we test for that instead of assuming it.
- Scrubbing runs on your machine through our SDK. Raw data never leaves your company that way; only the scrubbed output and the scrub report are uploaded.
- We remove names, contact details, account numbers and secrets such as API keys and passwords. We round rare amounts, coarsen dates and broaden job titles, and drop free text unless it has been reviewed.
- We test whether records can be re-identified before every sale and keep the report.
- Every buyer license bans re-identification, resale and combining your data with other sources, which also meets the contract condition in California law.
Health, finance, government and export-controlled data get a separate review or stay out.
Which one do AI labs need?
Labs want data that keeps its structure: which client, which step, what happened next. Stable pseudonyms keep that structure. The legal work is making sure the released dataset, with pseudonyms and generalized values, meets the anonymization or deidentification standard for the law that applies, and that the buyer is bound not to undo it.
Check your data before you sell it
Our calculator estimates what your records are worth, and the readiness report we run on a sample under NDA shows what has to go before anything is offered. Start at sell data to AI companies.
Frequently asked questions
Is de-identified data the same as anonymous data?
Not always. De-identified data has had identifiers removed or masked, but it only counts as anonymous if no one could reasonably identify a person from it, including by linking it with other data. Pseudonymized data, where a key still exists, is personal data under the GDPR.
Does the GDPR apply to de-identified data?
It applies unless the data is anonymous under the Recital 26 test. If a person can still be identified by means reasonably likely to be used, by you or anyone else, the GDPR applies in full.
Is removing names and emails enough?
No. Combinations of ordinary fields, such as ZIP code, gender and date of birth, identify most people. Amounts, dates, job titles and free text need to be generalized, dropped or reviewed, and the result should be tested.
Can a buyer re-identify the data I license?
Our licenses forbid it, and California law requires that contract term for data to count as deidentified. The stronger protection is technical: data that has been properly generalized and tested gives a buyer little to work with.