Five levels of protection, ordered by effort rather than by topic. Each one states exactly what it closes and what it leaves open, so you can stop at the level that matches what you actually paste.
Why this is arranged by effort Most guides on this topic list every risk and every countermeasure at the same weight, which produces either paranoia or nothing. The realistic question is not how to be safe but how much protection is proportionate to what you are actually doing, and where the return on effort stops. So this is a ladder. Each level costs more than the one below and closes more than the one below. Two thirds of the available protection sits in the first thirty minutes, and most readers should stop at level two. |
Pick by the most sensitive thing you paste, not by the average thing. One row applies to you, and the level it points to is where you should stop.
| The most sensitive thing you send | Stop at |
|---|---|
| Public information, general questions, drafting from scratch | Level 1 |
| Your own work, personal notes, anything you would not mind a stranger reading | Level 2 |
| Work material, colleague or customer names, internal documents | Level 3 |
| Client data, regulated records, anything under a confidentiality obligation | Level 3 plus level 4 habits |
| Material where a breach or a court disclosure would be unacceptable | Level 4, without exception |
The common error People choose their level by how they mostly use a tool, then paste something sensitive once and assume the earlier setup covers it. It does not, because levels one and two act on future behaviour while the exposure happened at the moment you sent it. If you occasionally handle sensitive material, you belong at the level that material requires, not the level your typical use requires. |
The advice below is worth more if you know how common the underlying behaviour is. Several independent bodies of research now measure this, and two of them use browser and endpoint telemetry rather than surveys, which matters because people under report what they paste.

Four measured figures from published industry research, and the data type uploaded most often
Browser telemetry published by LayerX found that 77 percent of enterprise AI users paste data into prompts, that 82 percent of those pastes originate from personal accounts outside company oversight, and that users average roughly fourteen pastes a day into non corporate accounts, of which at least three contain sensitive material. Around 22 percent of all pastes carried personal or payment card data, and roughly 40 percent of files uploaded contained the same. On that measurement, generative AI accounted for 32 percent of all corporate to personal data movement, making it the single largest such channel.
The trend matters more than any single figure, and it is the reason this guide exists in its present form.

The share of AI inputs containing sensitive data, tracked across successive annual reports
Longitudinal telemetry from Cyberhaven puts the share of corporate data entering AI tools that qualifies as sensitive at roughly 10.7 percent in 2023, 27.4 percent in 2024, 34.8 percent in 2025 and 39.7 percent in its 2026 report. The same research frames the cadence memorably, estimating that the average employee inputs sensitive data into an AI tool about once every three days.
Two further figures are worth carrying. The 2026 Verizon breach report analysed 858,440 data loss prevention events involving uploads to generative AI tools and found source code was the most frequently uploaded data type by a large margin, ahead of images and structured data. And IBM’s 2025 breach cost research reported that one in five organisations had experienced a breach involving unsanctioned AI, adding roughly 670,000 dollars to the average breach cost.
What these numbers actually argue for Not a ban. Organisations that attempt outright bans consistently find usage continues through personal devices and accounts, with the only measurable difference being that it becomes invisible to the security team. The 82 percent figure on personal account use is the important one, because it says the exposure is concentrated in exactly the tier where no data processing agreement exists. That is the argument for level three rather than for prohibition. |
Six risks, five levels, and no marketing language. Each level includes everything to its left.

Six risks against five levels of effort, with each level including everything to its left
Two things are worth reading off that grid before the detail. The top row closes at level one, in five minutes, which is why account security rather than anything about AI specifically is the highest return action available. And the bottom row never closes until level four, because retention under legal obligation is decided by a court rather than by a vendor or a setting.

The return curve flattens sharply after the first half hour
Each card gives what the level costs, who it suits, what to do, and an explicit statement of what it does not cover. That last part is the one usually left out.
LEVEL 00 No effort Defaults, meaning whatever you have now Everyone starts here, including people who assume they do not | |
Worth stating plainly, because most people believe they are further up this ladder than they are. On a consumer tier, the usual position is that your content may be used to improve future models unless you changed a setting, your conversations are stored and searchable on your account, some fraction of traffic is subject to human review for safety, and your account is protected by a password alone. Defaults also move. One provider updated its consumer terms in August 2025 and required free and paid consumer users to make a training decision by a deadline that October, with extended retention of up to five years attached to those who agreed. A widely used coding assistant was reported to have changed its individual tier to train on interaction data by default in April 2026. A setting you checked last year is not necessarily the setting you have. | |
WHAT TO DO 10. Nothing, which is the point of including this level | |
WHAT THIS CLOSES Nothing. | WHAT IT LEAVES OPEN Everything below, including the one risk you can fix in two minutes. |
LEVEL 01 -About 5 minutes, once The five minute floor Anyone using AI tools for anything at all, with no exceptions | |
Two actions, and the first is the highest return item in this entire guide relative to effort. Multi factor authentication with a unique password. Account takeover is far more common than a vendor breach and exposes exactly the same thing, which is a searchable archive of everything you have ever asked. Unlike almost every other risk here, it is entirely within your control and fully fixable. If you do one thing from this article, do this one. The training toggle, found under data controls in account settings on most consumer tiers. It stops your content being used to improve future models. It does not stop storage, review or disclosure, which is why it is level one rather than the answer. | |
WHAT TO DO 11. Turn on multi factor authentication and set a unique password 12. Open data controls and turn the training setting off 13. Note the date, so you know when you last checked it | |
WHAT THIS CLOSES Account takeover entirely, and future training use. | WHAT IT LEAVES OPEN Storage, human review, breach exposure and legal disclosure. |
LEVEL 02 -About 30 minutes, once The thirty minute pass Anyone who uses AI tools regularly, or for work of any kind | |
This is where most people should stop, and where the return curve flattens. It is a single sitting that covers everything a settings page can reach. The two items people skip are the ones that matter most. Shared links are the most common self inflicted exposure in this category, because several assistants let you publish a conversation to a URL and those links have repeatedly proved discoverable. Connected apps matter because a tool with standing access to your mail or files can be manipulated through instructions hidden in content it reads, which is the specific risk in agentic tools rather than in chat. Memory and personalisation settings belong here too. They control what a tool retains about you between conversations, which is a separate question from whether it retains the conversations themselves. | |
WHAT TO DO 14. Audit shared links and revoke any you no longer need 15. Review connected apps and integrations, and disconnect what you do not use 16. Set chat history and memory to the level you actually want 17. Run a data export to see what is currently held on your account 18. Learn where temporary chat mode is, so you can use it per conversation | |
WHAT THIS CLOSES Adds accidental publication, standing integration access and personalisation retention. | WHAT IT LEAVES OPEN Vendor side storage, human review, breach and disclosure, all of which are contract questions. |
LEVEL 03 - A few hours, plus budget Move to the right tier Anyone handling client data, regulated material or confidential work | |
Everything above happens inside your account. This level changes which account you have, and it is the only step that alters storage, human access and legal posture. The line that matters is between consumer plans and business or enterprise plans, not between free and paid. A personal paid subscription usually sits on the consumer side, with consumer defaults and no data processing agreement. Business, enterprise and direct API access are contractually excluded from training across the major providers, and enterprise contracts are where zero data retention becomes available at all. The telemetry in section 02 makes this concrete. If 82 percent of sensitive pastes come from personal accounts, then the single highest impact organisational change is not a policy document but providing a governed tier that people will actually use, since research consistently finds that unauthorised use falls sharply when an approved alternative exists. Two properties are worth negotiating for specifically. A signed data processing agreement, which is what turns a privacy promise into a contractual obligation. And regional processing, because jurisdiction had a concrete consequence in 2025, when one provider under a United States preservation order stated it was not required to retain conversations originating from the European Economic Area, Switzerland or the United Kingdom. | |
WHAT TO DO 19. Move confidential work onto a business or enterprise tier 20. Ask for a signed data processing agreement rather than accepting standard terms 21. Ask whether zero data retention is available and what it costs 22. Choose regional processing if your work is jurisdiction sensitive 23. Read the published subprocessor list before signing | |
WHAT THIS CLOSES Adds training exclusion by contract, constrained human review, and with zero data retention, storage itself. | WHAT IT LEAVES OPEN Breach exposure is reduced rather than removed, and legal disclosure remains partly outside your control. |
LEVEL 04 - Ongoing habit, or a one off setup Do not send it at all Anyone with material where a breach or a disclosure would be unacceptable | |
The only level that closes everything, because it removes the data from the equation rather than protecting it. There are two versions, and most people only need the first. Redaction. Almost every task works on placeholder data. Replace names with role labels, real figures with representative ones, and identifiers with markers, then substitute the real values into the output yourself. The model does not need the actual account number to draft the letter, and the version it never received cannot be retained, reviewed, trained on, breached or produced in discovery. This costs nothing and is the habit that separates careful users from lucky ones. Local models. Running a model on your own machine means nothing is transmitted at all. Software for this is now approachable for non specialists and capable open weight models run on a reasonably specified laptop. The honest trade is capability, since a model small enough to run locally is meaningfully behind a frontier model on hard reasoning and long documents. You do not have to pick one tool. Use a hosted assistant for the large majority of work that involves nothing sensitive, and keep a local model for the small fraction that does. That gives you frontier capability where it helps and genuine isolation where it matters. | |
WHAT TO DO 24. Adopt placeholder data as a default habit for anything identifying 25. Keep credentials, government identifiers and account numbers out entirely 26. Treat other people’s personal information as not yours to send 27. Install a local model if you regularly handle material that cannot leave your control | |
WHAT THIS CLOSES Everything, for the data you withhold. | WHAT IT LEAVES OPEN Nothing, for that data. It offers no protection for anything you do send. |
This distinction decides why level one is cheap and level three is not, and it is constantly conflated.
Training opt out means the vendor will not use your data to improve future models. Zero data retention means the vendor does not store your inputs and outputs beyond the time needed to answer. The second is far stronger, and it is generally available only on enterprise contracts.

The distinction that separates a settings change from a contract change
Why a paid personal plan does not move you up A toggle changes exactly one row of that comparison. Everything below it, meaning storage, human access, breach exposure and legal exposure, is decided by which contract you hold. The meaningful line runs between consumer plans and business plans, not between free and paid consumer plans, which is why a personal subscription and a business tier are different products rather than different price points. |

One route sits largely outside the ladder, and it is the one nobody plans for
An honest guide has to say where the ladder ends, and one route sits largely outside it.
In May 2025 a magistrate judge in the Southern District of New York ordered a major AI provider to retain and segregate output log data that would otherwise have been deleted, in a copyright case brought by a newspaper. The order reached conversations users had already deleted and chats marked temporary. The provider objected publicly, described it as a departure from privacy norms, and lost. The obligation to indefinitely retain new data ended that September and standard deletion resumed within thirty days, but in November 2025 the court ordered production of twenty million de identified conversation logs to the plaintiffs under a protective order. Users who tried to intervene were denied, as they were not parties to the case.
• A deletion feature is a policy rather than a guarantee. Deletion typically removes a conversation from your account and starts a retention clock, and a court can override that clock entirely.
• You have no standing in someone else’s litigation. The outcome is decided without you, which is why this is the one risk no setting reaches.
• Jurisdiction had a concrete consequence. The provider stated it was not required to retain conversations originating from the European Economic Area, Switzerland or the United Kingdom, which is a real argument for regional processing at level three.
There is also no professional privilege attached to a chatbot conversation. Material you would protect if it were correspondence with a lawyer is not protected because you typed it into an assistant instead.
The two other routes worth naming Human review continues regardless of your training setting, because safety and abuse monitoring is a separate process with its own retention. And prompt injection affects tools that browse or read files, where instructions hidden in content can be picked up and acted on. Level two reduces the second by limiting what you connect. Neither is closed by any setting. |

Levels one and two are one off tasks that quietly expire
Levels one and two are one off tasks that quietly expire, because defaults in this category move and sometimes move toward more collection. Four triggers are worth acting on rather than filing away.
• Any terms of service email. One provider required consumer users to make a fresh training decision on a deadline in late 2025, with extended retention attached to agreement. Treat the notice as a prompt to open your settings.
• Any major product update. New features arrive with new defaults, and memory, sharing and connector settings are the ones that tend to reset or expand.
• Connecting anything new. Every integration is standing access that outlives the task you added it for.
• Changing jobs or clients. The level appropriate to your work changes when your work changes, and level two habits do not carry over to a new confidentiality obligation.
A realistic cadence A five minute check twice a year covers this for almost everyone. Put it in a calendar rather than relying on noticing, because the changes that matter arrive in emails most people archive unread. |
Thirty minutes closes or reduces five of the six risks, for a single sitting and no money
The usual response to a list of AI privacy risks is either to stop using the tools or to decide the situation is hopeless and change nothing. The ladder exists to make a third option obvious, which is that protection is proportionate and most of it is cheap.
Almost everyone should be at level two, and almost nobody is. Thirty minutes, once, closes account takeover, training use, accidental publication through share links, standing connector access and personalisation retention. That is five of the six risks in the grid either closed or meaningfully reduced, for a single sitting and no money.
Level three is a real decision rather than a task, and it is only necessary if you handle client data, regulated material or work under a confidentiality obligation. If you do, no amount of settings work substitutes for it, because storage, human access and legal posture are contract properties rather than account properties. The telemetry makes the organisational version of this argument for you, since the exposure is concentrated in personal accounts rather than in governed ones.
Level four is the only thing that closes everything, and the version that matters is not the local model. It is the redaction habit. Placeholder data works for almost every task, costs nothing, and material the model never received cannot be retained, reviewed, trained on, breached or produced in discovery.
The sentence to keep Assume everything you type into a hosted AI tool is retained somewhere you cannot see, readable by someone you will never meet, and potentially disclosable in a proceeding you are not party to. That assumption is not paranoid, it is what the last two years documented, and it produces exactly the right behaviour without requiring you to stop using the tools. |
Share your thoughts about this article.
Be the first to post a comment!