Docs/For AI labs

What comes with every dataset

The schema, scrub report, re-identification result and rights evidence that ship with each dataset and batch.

Every batch arrives with the evidence you need for your own diligence and privacy review.

Schema

The columns of the scrubbed records as delivered, grouped by object type, with each column's role (such as identifier, quasi-identifier or free text), kind and the transform applied. A column the scrubber removed, such as a credentials column or a free-text column set to drop, is absent from the schema.

Scrub report

The JSON report produced by the scrubber on the seller's machine for that batch. It contains:

  • each column's role, the reason for it and the transform applied
  • the number of cells changed in each column
  • redaction counts by type in free text
  • the row, column and type of every secret removed, never the value
  • the k-anonymity result over the quasi-identifier columns, before and after any suppression
  • a leak check confirming no known identifier value remains in the output
  • the settings used: date and amount modes, the k threshold and the free-text setting

The report never contains raw values. See How scrubbing works for what each setting does.

Batch metadata

The source's connector type, the product, the period, the record count, and a SHA-256 checksum of the delivered file.

Rights evidence

Before a dataset is offered, we review the seller's customer agreements, data processing agreements and privacy notices to confirm the seller can license the data. We share the outcome of that review with buyers during diligence. See Rights review.

What the evidence does not cover

The scrub report shows reduced re-identification risk, not anonymization under GDPR, UK GDPR or CCPA. The limits of the method are listed in Limits, and the license terms that cover the remaining risk are in Licensing terms.