Yes, AI labs license electronic lab notebook (ELN) and LIMS records, and the most useful part is usually the experiments that failed. A notebook that records reagents, conditions, observations and yields for every run, including the ones that produced nothing, gives a science model information that published papers leave out.
Why failed experiments are worth money
Journals mostly publish what worked. A model trained only on published chemistry sees the successful corner of reaction space and has to guess where the boundaries are. Your notebooks hold the boundaries.
The best-known evidence comes from Alexander Norquist and colleagues, who took archived "dark reactions" from laboratory notebooks, the failed and unpublished hydrothermal syntheses, and trained a model on them. According to the 2016 Nature paper, the model outperformed traditional human strategies and predicted conditions for new products with a success rate of 89 percent. A later study in Angewandte Chemie, Machine Learning for Chemical Reactivity: The Importance of Failed Experiments, found that removing negative results changed the models' conclusions fundamentally, while added experimental noise mattered less.
Those findings explain why we tell chemistry and materials companies that the runs they consider embarrassing are often the most valuable records they have. You can read more on the chemistry and materials page and the life sciences page.
What buyers look for in notebook data
A notebook entry is useful when a stranger can reconstruct what was tried and what happened. The fields that matter most:
- Reactants, reagents, catalysts and solvents, with amounts or concentrations
- Temperature, time, atmosphere and work-up
- The expected product and what was observed, including decomposition or no reaction
- Analytical results (HPLC, NMR, QC) linked to the experiment that produced them
- Whether the run was performed as planned
- Revisions to the entry over time
That last item matters more than most sellers expect. ELN revision history shows how a chemist changed course after a result, which lets us build episodes: the situation, the action a person took and the outcome, in order. Episodes and plain records are different products with different prices, and our guide to episodes, records and environments explains both.
Scale-up history is the other strong category. What changed between bench and production, and why, is knowledge that rarely leaves a company.
What has to come out before anything is shared
Lab data raises three separate questions: identity, confidentiality and sector rules.
People
Chemist names, initials and sign-offs are personal data. We replace them with stable pseudonyms such as CHEMIST_12, so a buyer can still see that one person ran a series without learning who. Dates are generalized to a month or quarter. Free-text observations are dropped unless someone has reviewed them, because a note like "per call with the client's plant manager" can identify people and customers. The guide to de-identification vs anonymization covers why pseudonyms alone do not make data anonymous.
Clients and projects
Contract research labs often run work for clients. Project names, client compound codes and target names come out or get replaced, and the bigger question is whether your client agreements allow the data to be licensed at all. If you hold the data on behalf of a client, you usually need the client's written permission. Our guide on selling data you hold for customers walks through that test.
Health data
Life sciences labs that handle patient samples face health privacy rules. Our default is to leave patient data out unless it is already fully de-identified under your existing approvals. In the US, HIPAA allows de-identification by removing 18 types of identifiers (Safe Harbor) or by an expert finding that the re-identification risk is very small (Expert Determination), as described in HHS guidance. We do not treat our own scrub as a substitute for either.
How pharma companies already share data for AI
Large drug companies have treated their assay data as too sensitive to pool. The EU-funded MELLODDY project let several pharma companies train a shared model by federated learning, so each company's data stayed on its own systems. That project shows both sides of the market: the data is valuable enough to build infrastructure around, and owners want control over where it goes.
Our process follows the same instinct. When you use the SDK, scrubbing runs on your machine and only the scrubbed output and the scrub report are uploaded. The one exception is the sample for the readiness report, which you upload under NDA; our server scrubs it, keeps only the scrubbed output and report, and deletes the raw file.
What lab notebook data sells for
There is no public price list for ELN or LIMS data. Deals in this area are private, and the few public numbers in life sciences mix data with services. Tempus, for example, announced a three-year, $200 million agreement with AstraZeneca and Pathos that covered data licensing and model development fees for an oncology foundation model. That deal involves clinical and molecular data at very large scale, so it tells you buyers pay for science data, not what a mid-size lab's notebooks are worth.
What moves the price for a single seller, in our experience of how buyers evaluate samples:
| Factor | Why it matters |
|---|---|
| Failed runs included | Labs can't get them from papers |
| Linked analytical results | Outcomes make records checkable |
| Years of history | More conditions covered |
| Revision history | Lets us build episodes, which price higher than records |
| Clear rights | Unclear client rights cut value sharply |
Our data prices page for chemistry and the life sciences data prices page collect the public signals that do exist. For a range on your own records, use the calculator; lab notebooks and LIMS carry the highest base value of any system in it, and the number is confirmed on a call.
Keep earning as the lab keeps running
A lab produces new notebook entries every week. The first license sells your history. After that, the SDK can run monthly, pull only entries created or changed since the last run, scrub them on your machine and upload a new batch. Each license with an active refresh term pays its refresh price for every accepted batch, with no work from your team once the schedule is set. The recurring data revenue page shows how that adds up.
Find out what your notebooks are worth
Start with the valuation calculator, or read how the whole process works on the seller hub. On the first call we ask about client agreements before anything else, then systems, volume and years of history.
Frequently asked questions
Do buyers want failed experiments or only successful ones?
Both, and failed runs are often worth more because they are missing from published literature. Research such as the Nature paper on dark reactions showed that models trained with failed experiments predicted new syntheses better than human strategies did.
Will my client project data be included?
Only if your agreement with that client allows it. Data you generate for a client is often the client's, or covered by confidentiality terms that block a sale even after scrubbing. If we can't get a clear answer, we leave that data out.
Do you sell our proprietary compounds or recipes?
Nothing is shared without your approval. You choose which systems, date ranges and fields go into a dataset, and you can exclude any project, compound series or process you consider a trade secret.
What happens to patient-derived data?
We leave patient data out unless it is already fully de-identified under your existing approvals, such as HIPAA Safe Harbor or Expert Determination. Health data needs its own review before it goes near a buyer.