Data types

Selling software work history: plans, tickets and pull requests

AI labs pay for the record of how software teams decide and ship: plans, tickets, pull requests and outcomes. What sells, what stays out, and how to record it.

·5 min read

Software companies can license the record of how their team plans and ships work: the decisions behind each feature, the tickets that broke it down, the pull requests that implemented it and what happened after release. Labs training coding agents value that chain more than the code itself, because it shows how a real team turned a goal into a working change.

What labs want from software teams

AI coding agents can already write code. What they get wrong is the work around it: understanding what the team actually agreed to build, choosing between approaches, breaking the job into steps and knowing when a change is done. The records that show this are the ones labs pay for.

Plans and decisions

What the team chose to build, the options it turned down and why. A brief that says "retry webhooks instead of polling, because the vendor rate-limits us" carries more training value than the code that implements it, because the reasoning rarely survives anywhere else.

Work history

How each plan became tasks, how tasks became commits and pull requests, the review comments and the test results. Linked records matter here, because a ticket tied to the change that closed it shows a step and its result. Archives of Slack messages, Jira tickets and emails from shut-down startups have sold for $10,000 to $100,000 per company, according to Forbes.

Outcomes

What shipped, what was rolled back and what broke afterward. An outcome tells a model whether the approach worked. Labs pay for checked results: Epoch AI reports that RL tasks typically cost $200 to $2,000 each, with especially complex software engineering tasks reaching around $20,000 in rare cases. A real history of goal, steps and verified outcome is the raw material for that kind of task.

We package this chain as episodes: the situation, what each person did in order, and the result. See episodes, records and environments for how they differ from plain records.

What we leave out by default

Some material stays out unless you decide otherwise after a dedicated review.

  • Source code. We sell the decisions and outcomes your team recorded, and never your source code unless you choose to include it.
  • Secrets. API keys, tokens, passwords and internal hostnames are removed during scrubbing, wherever they appear, including inside ticket comments and PR descriptions.
  • Security logs and credentials, which are excluded entirely.
  • Free text with personal details, unless it has been reviewed.

People are replaced with stable pseudonyms such as PM_04, so a sequence still shows that the same person wrote the brief and approved the merge without naming them. Customer names mentioned in tickets are replaced the same way.

Rights to check first

Your own work versus client work

Code and records you build for your own product are usually yours to license. Work done for clients under contract may belong to the client, or the contract may restrict how you use anything you learned. Agencies and consultancies should read can I sell data I hold for my customers? first.

Employees and contractors

Tickets and reviews contain personal data about the people who wrote them. In the EU and UK that is personal data even in a work context, so selling it needs a lawful basis or proper de-identification. Our GDPR guide covers this.

AI coding-agent transcripts

Many teams now work through coding agents, and the transcripts look like ideal training data. They carry a specific risk: AI providers' terms generally forbid using their outputs to build competing models, so selling transcripts from one provider's tool to another lab may breach those terms. We confirm with a lawyer before collecting any transcripts. The model-written parts of a session are also worth little, since they are what a frontier model already produced. The human parts are worth more: the goal, the corrections, which edits were accepted or rejected and whether the change merged.

This is general information, not legal advice. Have a lawyer review your contracts and tool terms before any data moves.

Where the record usually breaks

Most teams have the pieces scattered. The plan lives in a doc that went stale the week after it was written. The tradeoff was settled on a call nobody recorded. The ticket links to the PR, but nothing links the PR back to the reason for the change. A buyer can still use that history, but the missing decisions lower its value.

How Hamster records the chain as you work

Hamster is our partner for software teams. Its context graph captures the chain while the team plans and ships: the brief that records what the team agreed to build, the decisions made along the way, the tasks that broke the work down, the pull request that implemented it and the outcome after release. Because the links are recorded as the work happens, a team using Hamster builds the decision-to-outcome history labs pay for as a side effect of its normal work.

That history can then be licensed like any other source. The first license sells the past, and new briefs, tasks and merged changes each month become refresh batches. Our SDK pulls only the new records, scrubs them on your machine and uploads the batch, so the revenue recurs without extra work. See recurring data revenue.

What drives the price for software data

  • Linked records from plan to outcome, beyond loose tickets
  • Years of history and the number of people contributing
  • Verified results such as passing tests, merges and rollbacks
  • Unusual domains, such as payments, infrastructure or regulated software
  • Clear ownership of the work

There is no public price list for company engineering histories. Our software data prices page collects the signals that do exist, and the software industry page shows the sample records we work with.

Estimate your team's data

Our calculator includes plans, tickets and decisions as a source, along with code repositories and coding-agent sessions. Value your data in about a minute, or read how a deal works on our seller page.

Frequently asked questions

Do I have to sell my source code?

No. We sell the decisions, work history and outcomes your team recorded, and leave source code out unless you choose to include it after a dedicated review. Secrets such as API keys and passwords are always removed.

What software data do AI labs pay for?

Labs pay most for linked records that show how a goal became a shipped change: plans and decisions, tickets, pull requests with reviews and test results, and what happened after release. Isolated code with no context is worth less.

Can I sell transcripts from AI coding tools?

Be careful. AI providers' terms generally forbid using their outputs to build competing models, so selling those transcripts to another lab may breach the terms. We confirm with a lawyer before collecting any, and the human parts of a session are worth more than the model-written parts.

How does Hamster help?

Hamster's context graph records the chain from brief to decision, tasks, pull request and outcome while your team works. That gives you the linked history labs value without reconstructing it later.

Find out what your records are worth.

Value my data