You need an organization in the portal, an API key from Settings, and a machine with Bun 1.2 or later that can reach the system you want to connect. The SDK runs on that machine, and raw records never leave it.
1. Create an API key
In the portal, open Settings → API keys and create a key. Keys starting with dy_test_ are for trying things out and never create billable batches. Keys starting with dy_live_ deliver real batches. Start with a test key. See API keys.
2. Install and initialize
Install Bun and the SDK as described in Install, then run:
datayield initdatayield init asks for your API key, checks it, and writes datayield.config.json in the current folder. It stores the key in ~/.datayield/credentials, readable only by your user, along with a new salt that keeps pseudonyms stable from one batch to the next. Neither is written to the config file. Back up the salt. See Install and Configuration.
3. Connect a source
datayield connect hubspotconnect adds a source to the config, asks for any required option you did not pass with -o key=value, and checks the credentials straight away, so a bad token fails here instead of on the first scheduled run. Tokens go to the credentials file, not the config. Each connector page lists the token it needs and the scopes to grant; start with Connectors.
4. Preview
datayield preview --out reviewpreview pulls records, scrubs them on your machine and prints the scrub report without uploading anything. --out review also writes the scrubbed records and the report to a review folder, so you can read exactly what would be sent. Read the column roles in the report. If a column is classified wrongly, fix it with an override in the config. See How scrubbing works.
5. Run
datayield runrun does the same work as preview and then uploads the scrubbed batch with its report, one batch for records and one for episodes if the source produces them. The first run of a source sends its history. Later runs send only records created or changed since the last successful run. If the leak check finds an identifier in the output, the run stops with exit code 2 and uploads nothing.
6. Schedule
datayield schedule --installschedule prints a cron line on Linux or a launchd agent on macOS. Add --install to install it, or use --github to write a GitHub Actions workflow instead. Once it is in place, new batches go out every month without anyone touching them. See Scheduling.
Check what happened
datayield statusstatus shows the last run of each source and every batch with its status. The same batches appear in the portal under Datasets.