Docs/Privacy and scrubbing

What leaves your machine

The SDK scrubs records where they live and uploads only the scrubbed output and its report. The one exception is the sample you upload for the readiness report.

Raw data never leaves your machine through the SDK. The SDK pulls records from your systems onto the machine where it runs, scrubs them there, and uploads two things: the scrubbed records and the scrub report. If the leak check finds a known identifier in the scrubbed output, the run stops and nothing is uploaded.

What the SDK uploads

Uploaded Contents
Scrubbed records Gzipped JSON Lines, one scrubbed record or episode per line.
Scrub report Column roles and the reason for each, the transform applied, counts of cells changed and redactions by type, secrets found (by row, column and type, never the value), and the re-identification result.
Batch metadata Source name, connector type, product, period, record count, and the schema of the scrubbed records: each column's name, role, kind and transform.
Checksum The SHA-256 of the uploaded file, so we can confirm it arrived intact.

The report never contains raw values. Where it shows examples, they are masked to their first character, such as j***@***.com or P***.

What stays on your machine

  • Raw records. During a run they sit in a private temporary folder readable only by your user, and the SDK deletes that folder when the run ends, including when it is interrupted.
  • The secret salt used to make pseudonyms. Anyone holding the salt can confirm a guess by hashing a candidate value, so it stays with you and never reaches a buyer or us.
  • The credentials for your source systems.
  • Run state, such as the time of the last successful run for each source.

The readiness report exception

Before you license anything, you can upload a sample export in the portal to get a readiness report. The NDA covers that upload. Our server scrubs the sample with the same scrubber, stores only the scrubbed output and the report, and deletes the raw file.

If you would rather no raw record reach us at all, tell us before you upload. You can run the scrubber on your own machine and send us only its output and report.

Optional model stages

The scrubber has two optional stages, both off by default. The SDK does not use them; they apply when you run the scrubber on its own. See How scrubbing works.

  • The column reviewer sends column names, the roles the rules chose and shape-masked samples (Aaaaa Aaaaa for a name) to the reviewer service you configure. It never sends a raw value.
  • The entity detector sends free-text cells, with secrets stripped, to the Presidio URL you configure. That text can contain personal details, so point it at a Presidio instance running on your own machine.

What the portal sees about your sources

The portal knows each source's name, its connector type and when it last ran. It does not see your source credentials, your configuration file or the raw records.

Who sees the scrubbed output

We review every batch before a buyer receives it. Buyers receive batches only for datasets they hold a license to, and only after you approved them. See Licenses and approving buyers.