Docs/Privacy and scrubbing

Readiness report

Send a sample under NDA and get back a report of what is sellable, what has to go and how hard the records are to re-identify.

The readiness report tells you what a buyer would receive and what we would remove, before you commit to anything.

How to get one

  1. Sign the mutual NDA. Sample uploads open once it is signed.
  2. Export a sample from one system as .csv, .json or .jsonl, up to 5 MB. We usually ask for about 1,000 rows.
  3. Upload it in the portal under Readiness (/app/readiness).

The server reads the file into memory, scrubs it with the same scrubber the SDK uses, and discards the raw file without writing it to storage or the database. It keeps the report and the scrubbed output. If the leak check finds a known identifier in the scrubbed output, the output is not kept either, only the report. Your deal timeline records the upload, the row count and the verdict.

If you would rather no raw record reach us, tell us before you upload. The scrubber is a command-line tool you can run on your own machine, and you can send us its output and report instead.

The verdict

The portal gives each sample one of two verdicts.

  • Sellable: no known identifier survived, no secrets were found, and every record shares its quasi-identifiers with at least the threshold number of others.
  • Needs work: one or more of those checks failed.

Either way, the portal lists anything still to fix, such as secrets to rotate, records that stand out, or free-text columns that need a human read or should be dropped.

Each column is also marked sellable or not. Free-text columns, removed columns and columns where a secret turned up are not sellable as they stand.

What the report contains

Summary. Rows in and out, secrets found, how many distinct identifier values were collected, and how many output cells still contain one, which must be zero.

Residual risk. Plain-language notes on what remains: free-text columns that need a human read, records that stand out, values that could not be parsed and were blanked, and secrets that need rotating.

Columns. For every column: the role it was given and why, what kind of data it holds, the transform applied, how many cells changed, and two masked examples.

Redactions by type. How many emails, phone numbers, names, secrets and other items were replaced in free text.

Secrets found. The row, column and type of each secret. The values are never shown.

k-anonymity. The quasi-identifier columns used, the minimum k, the number of equivalence classes and how many records fall below the threshold, before and after any suppression. See Re-identification testing.

Review needed. If the optional column reviewer ran, every column where it disagreed with the rules, with its suggested action and reason.

Settings. The k threshold, the date and amount modes and the free-text setting used, plus where the salt came from. The salt itself is never written out.

What happens next

We go through the report with you. If a column was classified wrongly, we fix it with a config override and run it again. If records stand out, we agree how to coarsen or suppress them. The report then becomes part of the evidence a buyer sees during diligence, alongside the rights review.

The report shows reduced risk, not legal anonymization. We use it as evidence in the privacy review, never in place of one.