
The great personal data training pipeline
2026-09-24 · Alex Wall, AIGP, CIPM, CIPP/US, CIPP/E, FIP, PLS, CISSP
It's 4:40pm on a Thursday. The approved assistant has hit its limit, isn't the model she trusts for contracts, or can't reach the drive where the contract lives. The deadline hasn't moved, and she's expected to be "AI Native". So she pastes the contract into her personal ChatGPT and, thirty seconds later, has a plain-English reading and a decision.
Nobody in that story thinks anything leaked, and it's a perfectly reasonable use of AI. I think it's the largest data pipeline in the industry, and it runs through a few small words in the terms of service.
Two rulebooks
OpenAI doesn't train on business customers' data by default, but for individuals using ChatGPT it "may use your content to train our models." Anthropic trains on consumer chats "when this setting is on"; business and API use is exempt. Outside Europe, Google's free developer API trains on what it receives; the paid one doesn't. xAI goes further still: in some regions outside the EU and UK, if you use Grok without logging in, you can't opt out of training at all.
The consumer side rests on a clean premise: whoever types owns what they type, and agreed to share it. OpenAI's terms have users "represent and warrant" that they have every right to provide it.
The problem is that the contract is her employer's confidential information, and the names in it are a client's personal data. Neither approved anything, and now it's stuck to the great retention flypaper in the cloud. Layoff fears, faster turnaround expectations and the pressure to look "AI Native" have produced a systematic, constant technical data breach: a violation of customer confidence and of written company policy.
The laundry loop
Under consumer terms, what she pasted, how she worked through it and what she concluded can all become training data. When her company later buys the enterprise version, it gets that learning back with the source washed out, re-used as idea particle board and sold by the seat. The split between personal and commercial terms is a laundry loop for confidential and personal data, and a trade secret vector nobody is watching. It costs, too: IBM found one in five organizations had a breach caused by shadow AI, and heavy shadow AI added $670,000 to the bill.
Why a policy won't stop the leak
Netskope finds 47% of people using generative AI at work use personal AI apps, while nine in ten organizations already block at least one AI app. It happens anyway, because the approved tool is missing one thing: the model the person prefers, or access to the file, repository or system the work depends on. In Zapier's survey of companies that pay for AI, 53% of personal-account users preferred their own tools, and 33% said the company's didn't fit their needs. Blocking just moves the paste to a phone.
What closes the laundry loop
I don't think the labs are villains. The grey area suits almost everyone: employees finish the work, employers get the output, labs get the data. The only party missing from the deal is the one that owns it.
What closes it is an approved tool people would choose anyway. That's why I built Caiioo.
It runs every leading model, and open models on your own hardware. Switching models is like swapping the engine in a car: your conversations, files and settings stay put.
It brings the AI to your work. Files, mail and calendars are reached from your own device, not uploaded to one more cloud, and conversations sync end-to-end encrypted, so only your devices can read them.
And it shows where every request goes. The Trust Boundary lists which AI companies receive your data and what their terms allow, and one switch refuses any provider that may train on it. We're building the same controls so a company can turn its paper policy into rules that hold for everyone.
I started Caiioo because I value capability, autonomy and security. I don't want to pick a winner or depend on one AI company; there is too much power consolidation already. Companies need to use these tools, protect their data, and stay independent.