Docs/SDK

CLI reference

Every datayield command, its flags, what it prints and its exit codes.

bash
datayield <command> [--config path/to/datayield.config.json] [options]

Every command accepts --config to use a config file other than ./datayield.config.json, and --help to print its options. datayield --version prints the SDK version.

Exit codes

Code Meaning
0 Success
1 An error: bad options, bad credentials, an API error, or a failed upload
2 The leak check found a known identifier in the scrubbed output. Nothing was uploaded.

datayield init

bash
datayield init [--api-key dy_live_...] [--api-url URL] [--skip-verify]

Writes datayield.config.json in the current folder, keeping any sources already in it. Stores the API key in ~/.datayield/credentials with mode 600, and creates the pseudonym salt if there is none. Without --api-key, it uses the key already stored, or asks for one in a terminal.

Unless you pass --skip-verify, it calls the API to check the key and prints the organization name and the key's mode. See Install.

datayield connect

bash
datayield connect <connector> [--name NAME] [-o key=value ...] [--products records,episodes] [--skip-check]

Adds a source to the config and checks its credentials against the system.

Flag Meaning
<connector> One of csv, postgres, xero, quickbooks, hubspot, salesforce, zendesk, jira, linear, github
--name NAME The source name. Defaults to the connector name. Use it to connect two systems of the same kind, such as xero-uk and xero-us.
-o key=value A connector option. Repeat for each option. true, false and numbers become JSON values, and values starting with [ or { are parsed as JSON.
--products records, episodes or both. Defaults to everything the connector supports.
--skip-check Save the source without checking the credentials

In a terminal, connect asks for any required option you did not pass. Secret options (accessToken, apiKey, apiToken, clientSecret, password, refreshToken, token, url) go to ~/.datayield/credentials, never to the config. A value of env:NAME is kept in the config and read from that environment variable at run time.

bash
datayield connect csv -o path=./exports -o timestampColumn=updated_at
datayield connect jira -o site=acme.atlassian.net -o email=ops@acme.com -o apiToken=env:JIRA_TOKEN
datayield connect github -o repos=acme/api,acme/web -o token=github_pat_...

Running connect again for an existing source updates its options. If the system rotates a refresh token during the check, the new token is the one saved.

datayield run

bash
datayield run [--source NAME] [--since DATE] [--period YYYY-MM] [--dry-run]

Pulls records changed since the last successful run, scrubs them on this machine, prints a summary, and uploads one batch per product. Without --source, every source in the config runs, one after another. A failure in one source does not stop the others.

Flag Meaning
--source NAME Run one source. Repeat to run several.
--since DATE Start the window here instead of at the last successful run, such as 2026-09-01 or 2026-09-01T00:00:00Z
--period YYYY-MM Label the batch with this period instead of the month in the middle of the window
--dry-run Pull and scrub, then stop before anything is sent. A dry run needs no API key.

For each source, run:

  1. checks the key with the API, and says so when it is a test key
  2. registers the source with the API, which returns the existing source if it is already registered
  3. works out the window, as described in Incremental runs
  4. pulls records, and audit events if the source produces episodes, into a private temporary folder
  5. scrubs everything in one pass, so a name seen in a record is also caught in episode text and gets the same token
  6. prints rows in and out, redactions, secrets removed, minimum k and leaks
  7. stops with exit code 2 if the leak check failed
  8. uploads a records batch and an episodes batch, and saves the new state
  9. deletes the temporary folder

A run that finds nothing new prints nothing new since the last run, uploads nothing and moves the state forward.

datayield preview

bash
datayield preview [--source NAME] [--since DATE] [--out DIR]

The same as run --dry-run, and then prints the full scrub report: column roles and the reason for each, transforms, redaction counts, secrets found and the re-identification result.

--out DIR also writes what would be uploaded, so you can read it: records.jsonl, episodes.jsonl, report.json and report.md. With several sources, each gets its own subfolder.

A preview makes no call to the Datayield API and needs no API key. It uses your stored salt, or a random salt for that preview only if none exists yet.

datayield schedule

bash
datayield schedule [--monthly | --weekly] [--install] [--github [--repo DIR] [--force]]

Sets up datayield run to run on its own. Monthly runs are at 06:00 on the 1st, and weekly runs at 06:00 every Monday.

Flag Meaning
--monthly Run monthly. The default.
--weekly Run weekly
--install Install the entry instead of printing it: a crontab line on Linux, a launchd agent on macOS
--github Write .github/workflows/datayield.yml and list the repository secrets it needs
--repo DIR The repository to write the workflow into. Defaults to the current folder.
--force Overwrite an existing workflow file

See Scheduling for what each option writes.

datayield status

bash
datayield status

Prints the API URL and key mode, then for each source its last successful run and the point the next run starts from. With an API key set, it also lists the source's ten most recent batches with period, product, line count, status and batch ID.

text
API: https://datayield.ai (live key)

xero-uk (xero)
  last successful run: 2026-10-01T06:04:12.000Z, next run pulls changes after 2026-10-01T06:00:00.000Z
  2026-09  records       1840 lines  accepted        <batch id>