Each dataset comes from one source system at one company and is sold as one product: records, episodes or an environment. Records and episodes arrive as monthly batches.
Batches
A batch is one scrubbed delivery of a dataset for one period, written as YYYY-MM. The first batch carries the history. Each batch after that holds the new records for its period, so a license with a refresh term receives a stream of monthly batches. Every batch comes with its record count, its schema and its scrub report.
Files are gzipped JSON Lines: one JSON object per line. The examples below are formatted across several lines for reading.
Records
A record is one row from the source system after scrubbing, such as an invoice, a deal, a ticket or a lab sample. Field names follow the source system, and each connector page lists them. Nested values are flattened into dotted names, such as CustomerRef.name. The _object field names the object type, such as invoice or invoice_line, since one batch can hold several.
{
"_object": "invoice",
"invoice_id": "ID_9c1e2a7f04b3",
"customer": "ORG_8b21e0c4a9f1",
"contact_email": "EMAIL_51d0c9e2a7b8",
"invoice_date": "2024-03",
"postcode": "LS6",
"total": 1200,
"status": "PAID",
"description": "Monthly retainer for ORG_8b21e0c4a9f1, [NAME] approved"
}The scrubber changes values in predictable ways:
| Value | In the delivered record |
|---|---|
| Names, emails, phones, addresses, account numbers, tax IDs, URLs, IPs | Tokens such as PERSON_…, EMAIL_…, PHONE_… |
| Company names | ORG_… tokens |
| Record IDs | ID_… tokens |
| Dates | Month (2024-03) by default; year, or shifted full dates in some datasets |
| Amounts | Two significant figures by default; rounded or bucketed (1000-5000) in some datasets |
| Postcodes | UK outward code (LS6) or ZIP3 |
| Rare categories | OTHER |
| Free text | Placeholders such as [EMAIL], [NAME], [SECRET], or the field is absent |
Tokens are consistent within a dataset. The same customer has the same ORG_ token in every record and every batch, so joins and per-entity histories work. The scrub report for each batch states the exact settings used.
Episodes
An episode is the ordered history of one thing in a system, such as a ticket, an issue, a pull request or an opportunity: the state it was in, what a person did, and what happened next. Episodes are built from the source's audit trail or change history.
{
"episode_id": "EP_4a2f9e1b7c30",
"source": "jira",
"entity": "issue",
"actor": "PERSON_3f9a1c2b7d4e",
"started_at": "2026-09",
"steps": [
{
"at": "2026-09",
"actor": "PERSON_3f9a1c2b7d4e",
"state": {
"issue_type": "Bug",
"project": "ENG"
},
"action": "created",
"outcome": "created"
},
{
"at": "2026-09",
"actor": "PERSON_9d02b5e1c6a7",
"state": {
"issue_type": "Bug",
"project": "ENG",
"status": "To Do"
},
"action": "status: To Do -> In Progress",
"outcome": "status = In Progress"
},
{
"at": "2026-09",
"actor": "PERSON_9d02b5e1c6a7",
"state": {
"issue_type": "Bug",
"project": "ENG",
"status": "In Progress"
},
"action": "commented: Fixed in the retry handler, [NAME] to verify",
"outcome": "comment added"
}
],
"outcome": "status: In Progress"
}| Field | Meaning |
|---|---|
episode_id |
Pseudonymous ID of the episode, derived from the scrubbed ID of the thing it describes |
source |
The connector it came from, such as jira or zendesk |
entity |
What the episode is about: ticket, issue, pull_request, opportunity and so on |
actor |
Token for the person who acted most often in the episode |
started_at |
Time of the first step, at the dataset's date granularity |
steps |
The steps in order |
steps[].at |
When the step happened |
steps[].actor |
Token for the person who took the step |
steps[].state |
The tracked fields before the step, such as status, priority or assignee |
steps[].action |
What happened: created, a field change written as field: old -> new, or a comment, review, merge or label written as commented: <text> |
steps[].outcome |
The immediate result, such as status = In Review, approved, merged or comment added |
outcome |
The final status, state, resolution or stage, with merged first for merged pull requests, or open if none is set |
Every value in an episode goes through the same scrubber as records, with the same pseudonyms. People in actors, assignees and owners become PERSON_ tokens, dates are generalized, and comment text is redacted like any other free text.
An episode covers the whole history of its item up to the end of the batch's window. When the item changes again later, it appears in a later batch with the same episode_id and its longer history. Keep the latest version.
Each connector page lists whether it produces episodes and from which history.
Environments
An environment is a working replica of a system: its structure plus synthetic data drawn from the distributions in the real records, for agents to practice in. The SDK does not produce environments, and there is no standard delivery format for them yet. Ask us about a specific system.