Docs/API reference

Batches

Create a batch, upload the scrubbed file, mark it complete and list batches with their review status.

A batch is one scrubbed delivery of a dataset for one period. Delivering one takes three calls: create the batch, upload the file to the URL you get back, then mark the batch complete with the file's checksum.

Statuses

Status Meaning
awaiting_upload Created. The file has not been confirmed yet.
received The upload was confirmed with its checksum.
in_review We are reviewing the batch and its report.
accepted The batch passed review. Licenses with an active refresh term pay for it.
rejected The batch failed review and is not delivered.

POST /api/v1/batches

Registers a batch and returns a URL to upload the file to.

Body

Field Type Meaning
source_id string From POST /api/v1/sources
product string records, episodes or environment
period string The month the batch covers, as YYYY-MM
record_count number Number of records or episodes in the file
schema object The columns of the scrubbed output, with each column's role, kind and transform
report object The scrub report JSON for this batch

The report is the report.json the scrubber writes, so the portal can show the same column roles, redaction counts and re-identification result the seller saw locally.

The SDK sends one batch per product per run. For records, schema is { "format": "jsonl", "objects": { <object>: [columns] } }, with one column list per object type. For episodes, it is { "format": "episodes/v1", "steps": [columns] }, describing the scrubbed step table the episodes were built from.

bash
curl https://datayield.ai/api/v1/batches \
  -H "Authorization: Bearer $DATAYIELD_API_KEY" \
  -H "Content-Type: application/json" \
  -d @batch.json
json
{
  "source_id": "...",
  "product": "records",
  "period": "2026-09",
  "record_count": 1840,
  "schema": {
    "format": "jsonl",
    "objects": {
      "invoice": [
        { "name": "invoice_id", "role": "safe", "kind": "id", "transform": "re-keyed with HMAC token" },
        { "name": "date", "role": "quasi", "kind": "date", "transform": "generalized to month" },
        { "name": "total", "role": "quasi", "kind": "amount", "transform": "rounded to 2 significant figures" }
      ]
    }
  },
  "report": { "tool": "datayield scrub 0.1.0", "files": [], "totals": {} }
}

Response

json
{
  "batch": { "id": "...", "status": "awaiting_upload" },
  "upload": {
    "url": "https://...",
    "method": "PUT",
    "headers": { "Content-Type": "application/gzip" }
  }
}

PUT the upload URL

Upload the gzipped JSON Lines file of scrubbed records to upload.url with the method and headers from the response. Send your API key only if the upload URL is on the API's own host; a signed storage URL needs only the headers the API returned. A successful upload returns 200.

bash
curl -X PUT "$UPLOAD_URL" \
  -H "Content-Type: application/gzip" \
  --data-binary @batch.jsonl.gz

POST /api/v1/batches/:id/complete

Confirms the upload. Send the SHA-256 of the file exactly as uploaded, as lowercase hex.

bash
curl "https://datayield.ai/api/v1/batches/$BATCH_ID/complete" \
  -H "Authorization: Bearer $DATAYIELD_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{ \"sha256\": \"$(shasum -a 256 batch.jsonl.gz | cut -d' ' -f1)\" }"

Response

json
{
  "batch": { "id": "...", "status": "received" }
}

GET /api/v1/batches

Lists batches for one source.

Query parameter Meaning
source_id The source to list batches for
bash
curl "https://datayield.ai/api/v1/batches?source_id=$SOURCE_ID" \
  -H "Authorization: Bearer $DATAYIELD_API_KEY"

Response

json
{
  "batches": [
    {
      "id": "...",
      "period": "2026-09",
      "product": "records",
      "record_count": 1840,
      "status": "accepted",
      "created_at": "2026-10-01T02:00:00Z"
    }
  ]
}