Docs/Privacy and scrubbing

Re-identification testing

What k-anonymity measures, how the scrubber computes it on every run, and what to do when records stand out.

Removing names is not enough. A record can still point to one person or one company through a combination of ordinary details, such as a postcode, a month and an unusual amount. The scrubber measures that risk on every run with a test called k-anonymity and reports the result.

k-anonymity in plain terms

Pick the columns an outsider might already know about someone, such as where they are, when something happened and roughly how much it cost. Those are the quasi-identifiers. Group the records that have exactly the same values in all of those columns. Each group is an equivalence class, and its size is k.

A record in a group of 1 is unique: anyone who knows those few details can pick it out. A record in a group of 20 hides among 19 others. The smallest group in the file is the file's minimum k.

An example

Here are six invoices after scrubbing, with month, postcode area and amount as the quasi-identifiers.

Invoice Month Postcode area Amount
ID_1a… 2024-03 LS6 1200
ID_2b… 2024-03 LS6 1200
ID_3c… 2024-03 LS6 1200
ID_4d… 2024-04 M1 860
ID_5e… 2024-04 M1 860
ID_6f… 2024-05 EH3 48000

There are three groups. The first has three records (k = 3), the second has two (k = 2), and the last invoice is alone (k = 1). With the default threshold of 5, all six records are below it, and the invoice for 48,000 in EH3 in May stands out the most. Someone who knows that a company in that area paid a large invoice that month could find it.

What to do when records stand out

The report lists the minimum k, the number of groups and how many records fall below the threshold. To raise k:

  • Coarsen dates from month to year with "dates": { "mode": "year" }.
  • Bucket amounts into wide ranges with "amounts": { "mode": "bucket", "edges": [0, 1000, 10000, 100000] }.
  • Take a column out of the quasi-identifier list if it does not need to be there, or drop it from the output.
  • Remove the records that remain unique with --suppress-below 5, which deletes every record whose group is smaller than 5 and reports k before and after.

Coarsening keeps more records but less detail. Suppression keeps detail but drops the outliers, which are sometimes the most interesting records to a buyer. We agree the trade-off with you before a batch goes out.

Which columns count

By default, the quasi-identifiers are the date, amount, postcode and category columns that the scrubber generalized. You can set the list yourself with quasiIdentifiers in the config. Company columns are pseudonymized and left out of the check.

What k does not tell you

k-anonymity is computed per file, over the chosen columns only. It says nothing about columns outside that set, about linking several delivered files on their shared tokens, or about outside data a buyer might hold. A small sample has small groups, so a sample's k will usually be lower than the full dataset's. A company can still stand out from its amounts, dates and volumes even when its name is a token. The report states these limits, and so does Limits.

Every license also forbids the buyer from trying to re-identify anyone. See Licensing terms.