What data do AI companies buy? A guide for company owners
AI companies buy data that isn't already online and that shows people thinking and working: long written conversations, decisions with their reasons, and what happened afterwards. For a company, that means years of email, chat, tickets, documents and code that link together, that you have the right to license, and that can be properly de-identified.

Why labs now want private company data
Public text is running short. Epoch AI estimates the stock of public human-written text at about 300 trillion tokens and expects it to be fully used between 2026 and 2032. Its estimate leaves out private data, such as company messages, because that data is fragmented across platforms and legally harder to reach.
That is the gap labs are now trying to fill. OpenAI's data partnership programme asked for datasets that are not already easily accessible online, and particularly for data that "expresses human intention", such as long-form writing or conversations rather than disconnected snippets. In April 2026, TechCrunch reported that Meta would record some employees' keystrokes and clicks to train its models, because they need real examples of how people actually use computers.
Decisions are worth more than finished work
Finished output is common: the internet is full of code, documents and published answers. What is scarce is the reasoning around the work. OpenAI's research on maths problems found that feedback on each step of a solution trained a better model than feedback on the final answer alone.
| Level | What it is | Value to a lab |
|---|---|---|
| Outcomes | "Ticket opened, ticket closed" | Low: says nothing about how |
| Finished work | Shipped code, final documents, sent emails | Low to medium: close to what is already public |
| The process | Drafts, reviews, handovers, the back-and-forth | High: shows how experienced people improve work |
| The decisions | Options weighed, the one chosen, the ones rejected and why, and the result | Highest: judgement is what models find hardest to learn |
Builders of training environments, who turn real workflows into practice worlds for AI agents, say the same in their own terms. One buyer's checklist from Invisible Technologies puts it bluntly: "Generic simulation trains generic agents." It asks for real tool interfaces, exceptions and the points where an experienced person would escalate rather than proceed.
What raises value, and what rules data out
- Years of unbroken history on the same tools. A move to new software three years ago leaves three usable years, however old the company is.
- Systems that link: one piece of work you can follow from an email to a chat thread to a ticket to the change that shipped.
- Work done in writing. If decisions happen on calls, the record only shows that something was opened and closed.
- Outcomes recorded: whether the fix held, the deal closed or the customer stayed.
- A field where public data is thin. What buyers want shifts over time, so timing matters.
Data you hold for clients under their contracts, privileged legal advice, health records you can't de-identify properly and anything about children are usually out, whatever their quality.
Does your data have what labs look for? Tick what is true
Tick everything that is true. Your result appears here, and changes as you go.
How to read this
- Thin for now: Buyers will see little reasoning or history. Keep your records in writing and on stable tools, and look again later.
- Some of what buyers want: There is something here. A review can show which systems and years are worth preparing, and what must stay out.
- Worth a proper review: Your data has depth, written decisions and linked systems. Take the full check on our data page.
Who uses company data, and for what
Most company data goes into training environments: simulated workplaces where AI agents practise real tasks. Forbes reported that the Slack, email and Jira archives of shut-down startups sold through one wind-down firm were used for this, at $10,000 to $100,000 per company.
Buyers fall into three groups: AI labs themselves, companies that build training environments for labs, and data marketplaces or brokers that package data from many sources. Our guide on how to sell data to OpenAI explains how each route works.
Questions
What kind of data do AI companies buy?
Data that isn't freely online and shows people reasoning: long written conversations, decisions with their reasons, and outcomes. In companies, that is email, chat, tickets, documents and code over many years.
Is old company data still valuable?
Yes, if it is continuous. Years of history on the same tools are worth more than a large volume from a short or broken period.
Do AI labs want our code?
Code alone is common. The discussion around it, such as reviews, rejected approaches and the tickets that led to it, is scarcer and usually worth more.
Is phone-based work valuable to AI labs?
Less so. If decisions are made on calls, the written record shows little of the reasoning, unless calls were recorded and can lawfully be used.
What data can't we license?
Usually data you hold for clients, privileged legal advice, health records you can't de-identify to the required standard, and anything about children.
Sources
- Will we run out of data? Limits of LLM scaling based on human-generated data (June 2024) Epoch AI
- OpenAI Data Partnerships (November 2023) OpenAI
- Meta will record employees' keystrokes and use it to train its AI models (21 April 2026) TechCrunch
- Let's Verify Step by Step (May 2023) Lightman et al., OpenAI, on arXiv
- RL environments for enterprise workflows: a buyer's checklist (14 July 2026) Invisible Technologies
- Failed companies are selling old Slack chats and email archives to train AI (17 April 2026) Gizmodo, reporting Forbes