Sell your company's data to AI labs, the safe way
AI labs pay for private company data that shows how work really gets done: written discussions, decisions and their outcomes, over years. Data is worth more with a long history on the same tools and systems that link together. Check your contracts and platform terms first, and have it de-identified independently before anything leaves your systems.
Why AI labs want company data now
The public internet is running out as a source of new training text. Epoch AI estimates the stock of public human-written text at about 300 trillion tokens and expects it to be fully used between 2026 and 2032. Its estimate leaves out private data, such as company messages, because that data is fragmented and legally harder to reach.
That private data is exactly what labs now seek. OpenAI's data partnership programme asks for data that is not already easily accessible online, and especially data that "expresses human intention", such as long conversations rather than disconnected snippets. In April 2026, TechCrunch reported that Meta would record some employees' keystrokes and clicks to train its models, because they need real examples of how people actually use computers.
A market has formed. Forbes reported that one wind-down firm had handled about 100 sales of shut-down startups' Slack, email and Jira archives in a year, at $10,000 to $100,000 each. In 2026, Google bid $10 million for Spirit Airlines' internal data in its bankruptcy, according to a law firm briefing that also set out the objections it drew.
What labs pay for: decisions, not output
Finished work is common, because the internet is full of it. What is scarce is the trail of reasoning around the work, which only exists inside companies. OpenAI's research found that feedback on each step of a solution trains a model better than feedback on the final answer alone, and builders of training environments for labs stress real workflows, with their exceptions and escalations.
| Level | What it is | Value to a lab |
|---|---|---|
| Outcomes | "Ticket opened, ticket closed" | Low: says nothing about how the work was done |
| Finished work | Shipped code, final documents, sent emails | Low to medium: much of it looks like what is already public |
| The process | Drafts, reviews, handovers, the back-and-forth | High: shows how experienced people improve work |
| The decisions | Options weighed, the one chosen, the ones turned down and why, and what happened next | Highest: this is judgement, which models find hardest to learn |
That is why a Slack thread where your team argued over three fixes, rejected two and explained why, then linked to the ticket and the change that shipped, can be worth more than the code itself.
What makes your data worth more, and what rules it out
Raises value
- Depth: years of unbroken history on the same tools. A switch of tools three years ago means three years of usable history, however old the company is.
- Linked systems: one piece of work you can follow from email to chat to ticket to change to result.
- Decisions in writing: remote, written teams keep their reasoning. If the work happens on calls, the record only shows that a ticket was opened and closed.
- Outcomes: whether the fix worked or the customer stayed.
- Rarity: fields where public data is thin. What buyers want shifts over time.
Rules data out
- Data you hold for clients under their contracts, unless each client agrees.
- Privileged legal advice, which generally needs each client's informed consent.
- Health records you cannot de-identify to HIPAA's standard, and anything about children.
- Promises in your privacy notice, customer terms or handbook that rule out new uses.
Is your data worth licensing?
Seven questions, about two minutes. No files, and nothing leaves your systems.
Where is the company based?
How many people work there?
How long on the same main tools?
How does most of the work happen?
What language is most of it in?
Where does the work live? Tick all that apply
Do any of these apply? Tick all that do
Answer the questions above. Your read appears here straight away.
A rule-of-thumb read, not a valuation or legal advice. Nothing you tap is saved or sent.
Who buys company data
- AI labs, which train general models and take data through partnership programmes and licensing deals.
- Training environment builders, which turn real workflows into practice worlds where AI agents learn to do work. The shut-down startup archives in the Forbes report went to this kind of use.
- Data marketplaces and brokers, which package data from many sources for buyers.
You don't need to approach a lab directly. Most deals run through intermediaries who know what each buyer is looking for at the moment.
Why you shouldn't de-identify it yourself
Removing names is the easy part. The hard part is everything else in free text: signatures, quoted replies, nicknames, project codenames, customer names and numbers. Even widely used detection tools say plainly that they can't guarantee to find all sensitive information, and a 2026 study showed that AI can link pseudonymous writing back to real people at scale.
- Keep the relationships. Each person should become the same stand-in, such as "employee 632", across every system, so who-said-what survives. Ad hoc redaction breaks it.
- Keep the reasoning. In-house clean-ups often delete the threads that look messy, which are the ones buyers value most.
- Keep it independent. A specialist that processes the raw files, deletes them afterwards, and documents the method is easier to stand behind than a script your team ran.
The FTC fined Avast $16.5 million in 2024 after it sold browsing data it had described as anonymous. "Anonymous" is a claim you can be held to.
The legal ground in the US, in brief
- De-identified, properly. Under California's privacy law, data counts as de-identified only if you take reasonable measures against re-identification, publicly commit not to re-identify it, and contractually bind every recipient. Virginia, Texas, Oregon and others use the same three steps.
- Employees. Most state laws exclude employees, but California's has covered them since January 2023.
- Your promises. The FTC says a company "cannot then unilaterally renege" on the privacy commitments it collected data under.
- Your tools. Slack's API terms ban using its API data to train large language models, and Google's Workspace API policy bans transfers for general AI models. Check how your data can be exported before you plan anything.
- Sector laws. HIPAA for health records, GLBA for financial institutions, and Illinois' biometric law for voice and face data.
This is general information, not legal advice. Every deal needs its own legal review.
How it works with HerWorkCircle
- Take the free check above, then a free call. You share record counts and the tools you use, not files.
- We review what you could license: which systems, how many years, what has to stay out, and what your contracts and tools allow.
- An independent specialist de-identifies a sample, and our trained data-checking team reviews it by hand for anything software missed. You approve it before anything goes further.
- We take the prepared dataset to buyers looking for that kind of data, and help you weigh the terms: scope, length, exclusivity and deletion at the end.
- You sign the licence you choose, the data is delivered, and you keep your business and your records.
Questions
Can I sell my company's data to AI companies?
Often, yes, if the data is yours to license, it can be properly de-identified, and your contracts, privacy promises and software terms allow it. Data you hold for clients usually belongs to them.
How much is company data worth to AI labs?
There is no price list. Forbes reported that shut-down startups sold their Slack, email and Jira archives for $10,000 to $100,000 each through one wind-down firm, and Google bid $10 million for Spirit Airlines' internal data in bankruptcy. Depth, linked systems and clean rights move the price most.
Who buys company data for AI training?
AI labs, companies that build training environments for labs from real workflows, and data marketplaces. OpenAI, for example, has invited datasets that are not easily available online.
Is selling company data to AI labs legal in the US?
It can be. Under California's law and similar state laws, data counts as de-identified only if you take reasonable measures, publicly commit not to re-identify it, and bind every buyer by contract. Contracts, privacy promises, platform terms and sector laws such as HIPAA also apply.
Do I need my employees' consent?
Most state privacy laws exclude employees, but California's has covered them since January 2023. The FTC also warns against quietly changing privacy promises to allow new uses. Check what your handbook and notices say, and get legal advice.
Can I sell our Slack data?
Slack's API terms ban using data from its API to train large language models and ban bulk export without a separate agreement. A workspace owner's own export is governed by the customer agreement instead, so check it with counsel before planning a sale.
Why shouldn't we de-identify the data ourselves?
In-house clean-ups often delete the discussions buyers value, miss names hidden in free text, or replace people inconsistently, which breaks who-said-what. An independent process is also easier to stand behind if anyone asks.
Do we lose our data if we license it?
No. A licence lets a buyer use a copy on agreed terms: scope, length, whether it is exclusive, and what happens at the end. You keep your records and your business.
Sources
- Will we run out of data? Limits of LLM scaling based on human-generated data (June 2024) Epoch AI
- OpenAI Data Partnerships (November 2023) OpenAI
- Meta will record employees' keystrokes and use it to train its AI models (21 April 2026) TechCrunch
- Failed companies are selling old Slack chats and email archives to train AI (17 April 2026) Gizmodo, reporting Forbes
- Selling company data for AI training: what to know before you say yes (1 October 2026) Frankfurt Kurnit Klein & Selz, law firm
- Let's Verify Step by Step (May 2023) Lightman et al., OpenAI, on arXiv
- RL environments for enterprise workflows: a buyer's checklist (14 July 2026) Invisible Technologies
- California Consumer Privacy Act, section 1798.140(m), deidentified California Privacy Protection Agency
- AI (and other) companies: quietly changing your terms of service could be unfair or deceptive (13 February 2024) US Federal Trade Commission
- FTC order on Avast's sale of browsing data (22 February 2024) US Federal Trade Commission
- Slack API Terms of Service (29 May 2025) Slack
- Google Workspace API user data policy protections (3 February 2024) Google Workspace
- Presidio: data protection and de-identification SDK Microsoft (open source)
- Large-scale online deanonymization with LLMs (February 2026) Lermen et al., on arXiv
Know the ground before you talk to a buyer
What labs pay for, de-identification, the law in each country and deal terms, with sources on every guide.
- Written 11 Oct 2026What data do AI companies buy? A guide for company ownersWhat AI labs look for in private company data: decisions over output, depth over volume, linked systems, and what rules data out. With a quick self-check.Read the guide
- Written 11 Oct 2026How to sell data to OpenAI and other AI labsThere is no upload form or price list. What OpenAI has asked for, the routes companies actually use, what buyers check, and the steps to take first.Read the guide