With every prompt, AI labs are ingesting more of your business data every day. Your only protection? A loosely-worded promise not to train on it.
A recent Lawfare analysis(nieuw venster) argues that when businesses hand their data to a few powerful AI labs, they are teaching those labs how to make them redundant. The authors call this replacement through knowledge acquisition (RKA).
While there is no way of knowing what AI labs do with your business data today, three converging forces are in play: largely unchecked power to set the terms, a legal system still grappling with what “training” even means, and both the ability and the incentive to compete with their own customers.
Your business may be feeding the beast that could ultimately consume it whole.
Do AI labs train on enterprise data?
The largest AI labs have made a concerted effort to assure business users that their data is private by default:
- OpenAI claims that it does not train on inputs or outputs from ChatGPT for business users.
- Anthropic claims that business data is excluded from training for Claude’s model by default.
- Microsoft claims that prompts, responses, and data aren’t used to train Copilot.
- Google claims that Gemini doesn’t train its model on business data without prior permission or instruction.
- X claims that interactions on Grok business plans don’t improve its model, but will apparently use its own SpaceX employee data for training purposes.
In practice, their promises to protect their users’ privacy are intentionally vague, and thinner than they look:
- Key terms are loosely defined. Lawfare points out that OpenAI’s enterprise agreement defines customer content as the input and output, then defines those using the words “input” and “output.” Anthropic’s terms rule out training models on customer content without defining what training means.
- Some data may fall outside the promise. Patterns, metadata and summaries an AI generates but never shows you may not count as inputs or outputs at all.
- “Training” is narrower than it sounds. Using your data to build test benchmarks, tune system prompts or train separate safety tools may not count.
- The terms can change. Providers typically reserve the right to amend them, and every new model release is a chance to reset them.
- You can’t check. Only the lab knows what goes into its training. A breach would be nearly impossible for you to spot.Proton founder and CEO Andy Yen has warned that once you share information with an AI, you can’t really take it back. Before long, he argues, AI “could know you better than even you yourself.”(nieuw venster) For a business, the stakes are higher.
What your business stands to lose
What an AI learns about your business is what makes it competitive. Big Tech knows that.
In a case that ran from 2019 to 2022, the European Commission investigated Amazon for using nonpublic data from third-party sellers to inform its own retail business — in effect, using sellers’ own data to compete against them. Amazon committed to stop, and the ban was later written into the Digital Markets Act.
In some industries, it’s not just a competitive edge at stake. Compliance is at risk too. Businesses in medicine and finance have a legal duty to protect sensitive information under rules such as HIPAA and the GDPR. Those rules are clear about securing personal data, but the AI industry is so new that a gray area is growing.
And with AI now built into meeting note-takers, email and everyday workflows, it touches far more of your business than it did when it lived in a chat window.
What AI labs stand to gain
AI labs have a financial and competitive incentive to ingest as much data as possible.
The AI market is set to keep growing dramatically, with Gartner speculating that AI will touch all IT work by 2030. It also forecast that AI spend would reach $2.5 trillion in 2026, a 44% increase on the previous year. Given the challenge of continuing to develop more intelligent AI models without significant training materials, AI firms are more likely to begin finding ways to benefit from their enterprise customers.
Exploding demand is forcing AI labs to become increasingly creative about the ways they can develop their models. And business data is clearly valuable to these labs: Google won a $10 million bid to purchase data from bankrupt airline Spirit Airlines in August. According to the Association of Flight Attendants, this data packet includes employee sensitive data. That data could be ingested and used to train Google’s AI models without the consent of those employees.
OpenAI claims that it does not train on inputs or outputs from its products for business users. But knowing about the enormous value of this data to businesses, we now need to consider the AI industry’s legal standing when it comes to processing data.
We can’t rely on regulation that might not come
Can regulation protect your business data from the largest AI labs? To grasp this, we must examine the financial and the political realities of the industry.
The US dominates AI. According Stanford’s Artificial Index Report, 59 notable models operate out of the US, the largest number of any country in the world. This dominance is important to note; in March, the White House released a national policy framework for regulating AI. It calls for state-level laws to be replaced with a federal AI policy framework to “protect American rights, innovation, and prevent a fragmented patchwork of state regulations that would hinder [US] national competitiveness.” The framework notes that state-level laws regulating AI must not be allowed to “act contrary to the United States’ national strategy to achieve global AI dominance.”
Is it more likely that the US government will prioritize the data privacy rights of individual businesses, potentially located outside the US? Or the expansion rights of its national AI empire? European businesses simply can’t guarantee that American AI tools will put their data sovereignty first. With the largest AI labs operating in the US, there’s an inflection point on the horizon.
A better choice for enterprise AI
Instead of hoping for regulation or trying to keep up with privacy policies, businesses have an easier and more effective option: choosing an AI tool built for privacy. There’s room for innovation without exploitation. The AI tool you choose should stay just that: an AI tool to assist you; not a competitor you’re training to take your place. It should seek to fulfill its function, growing over time, without poaching trade secrets.
Proton built Lumo to meet this need. With no incentive to collect or exploit user data, AI can be a transformative tool for businesses. Lumo keeps no logs of business chats, protecting your data from third parties and government surveillance. With an ad-free business model and no funding from Big Tech or venture capital, Lumo’s only incentive is to provide the best business support possible. Proton as a business is governed by the Proton Foundation, meaning that we act in the interest of people, not industrial domination.
Based in Switzerland, Proton is also regulated by stringent data privacy laws and the GDPR. Its code and model are both open source, so any business can verify our claims. Lumo is designed to do exactly what AI should do: provide assistance without exploitation. Choosing a European alternative to Big Tech’s AI tools is a business’s best way to protect itself from unregulated and untrustworthy AI labs.






