A good export for an AI data buyer contains the full history of one system, the audit trail that shows how records changed, and the IDs that link one table to another, with personal details and secrets scrubbed out before anything leaves your machine. Most of the value is lost by exporting too little context, so export broadly, then remove what has to go.
The steps below follow the order we use with sellers. You can do most of them yourself, and our SDK automates the scrubbing and packaging for the systems it supports.
Step 1: decide what the buyer is paying for
Labs buy three kinds of product, and each needs a slightly different export. The episodes, records and environments guide covers them in depth.
- Records: de-identified rows from one domain, such as invoices, experiments or tickets.
- Episodes: a situation, the action a person took and the outcome, in order. These are built from audit trails and event histories.
- Environments: a working replica of a system, built from its structure plus synthetic data that matches the real distributions.
If you only export current records, you can sell records. If you also export history, you can sell episodes too, and episodes are where task-level prices apply. So when a system offers a field history, audit log or event export, include it.
Step 2: pick systems and a date range
Start with one system that holds the core of your work: the accounting package for a bookkeeping firm, the ELN for a chemistry lab, the ticket system for a support team. A second system adds value when it shares IDs with the first, because linked records show cause and effect across tools.
Take all the history you have. Years of records matter more than most sellers expect, because labs want to see how work changed over time and how outcomes turned out. Our calculator weights years of history directly.
Leave out any system you cannot license. Read step 6 below before you export from systems that hold customer data.
Step 3: export with structure intact
The most common mistake is an export that flattens the data into a spreadsheet and loses its shape.
| Keep | Why it matters |
|---|---|
| Primary and foreign keys | They link an invoice to its payments, a ticket to its pull request, an experiment to its analysis |
| Timestamps on every change | They put actions in order, which is what an episode needs |
| Status and outcome fields | Paid, rejected, merged, rolled back, failed: labs value outcomes highly |
| Field history and audit logs | They show who changed what and when |
| Schema and field descriptions | A buyer can't use a column it can't interpret |
| Attachments metadata | File type and size, even if the file content is excluded |
Use the system's native export or API where you can. JSON or CSV with a separate schema file both work. Avoid PDFs and screenshots, which strip structure.
Keep IDs consistent across tables. If you replace an internal customer ID with a pseudonym, replace it the same way everywhere, so the links survive scrubbing.
Step 4: scrub on your own machine
Scrubbing is the step that makes a sale possible, and it should happen before the data leaves your control. Our SDK runs the scrub locally. Raw data never leaves your machine through the SDK; only the scrubbed output and the scrub report are uploaded.
What the scrub does:
- Removes direct identifiers: names, email addresses, phone numbers, account numbers and street addresses. Where a link must survive, the value is replaced with a consistent pseudonym such as
ORG_7F3A. - Removes secrets: API keys, tokens, passwords and internal hostnames. These turn up in ticket comments and notes more often than anyone expects.
- Generalizes quasi-identifiers, the values that identify someone in combination. Exact amounts become ranges, exact dates become a month or quarter, and rare job titles become broader ones.
- Drops free-text fields unless they have been reviewed, because notes and comments carry names and details that pattern matching misses.
Then the scrub tests whether any record can still be singled out from a combination of fields, and removes the ones that can. Removing names alone is not enough, and the de-identification vs anonymization guide explains why with published research.
By default we exclude raw source code, security logs, credentials and free text containing personal details. Any of these can be added back only after a dedicated review.
Step 5: read the scrub report
The scrub report lists every column's role, how many values were removed or generalized by type, and the re-identification result. Read it before anything goes to a buyer. Check that:
- every column you consider sensitive was caught;
- the pseudonyms still link the tables you care about;
- the generalization didn't wipe out the signal (if every amount became "over $1,000", the data lost its value);
- no free-text field slipped through unreviewed.
Run a preview first. datayield preview runs the whole pipeline without uploading anything and prints the report, so you can adjust settings and run again. The quickstart walks through it.
Step 6: check rights before you send anything
Ownership of a database is a different question from the right to license what is in it. Before any export leaves, confirm:
- your customer agreements and data processing agreements allow it, and that no clause limits the data to "providing the service";
- you are not holding the data as a processor for your customers, in which case you usually need their written permission (see can I sell data I hold for my customers?);
- sector rules don't apply, or have had their own review. Health, finance, government and export-controlled data need special handling or are excluded.
We read these contracts with you during the rights review. If we can't get a clear answer for a system, we leave that system out.
Step 7: send a sample, not the whole thing
The first data to leave is a sample for the readiness report, about 1,000 rows, under a mutual NDA. If you upload it in the portal, the server scrubs it, keeps only the scrubbed output and the report, and deletes the raw file. The report tells you what is sellable, what has to go and the re-identification result. NDAs and data rooms covers what protects that sample.
Step 8: set up the refresh
Once a license is signed, the history goes first. After that the SDK can run on a schedule, monthly by default, pulling only records created or changed since the last run, scrubbing them locally and uploading a new batch. Each license with a refresh term pays per accepted batch, which is how one export becomes recurring data revenue. Setting the schedule is one command: datayield schedule --monthly.
Start with a number
If you are deciding whether an export is worth the effort, get an estimate first. The valuation calculator takes two minutes and needs no access to your systems. For the full process, see selling data to AI companies.
Frequently asked questions
What file format should the export be in?
Structured formats that keep keys and timestamps: JSON lines or CSV, plus a schema file describing each field. Avoid PDFs and screenshots. Our SDK uploads scrubbed batches as gzipped JSON lines.
Should I delete fields I think are useless?
Not before talking to a buyer or broker. Status fields, timestamps and IDs that look boring are often what lets a lab build episodes. Remove fields for privacy and rights reasons, and keep the rest.
Do I have to clean up messy data first?
No. Real data is messy, and labs want to see how work actually happened. Fix only what blocks interpretation, such as missing field descriptions. Deduplication and packaging happen later in the pipeline.
Can I export everything and let you scrub it?
Through the SDK, scrubbing runs on your machine and raw data never leaves it. The only raw data we ever receive is the readiness report sample under NDA, which the server scrubs and then deletes.