Local & private AI · 10 min read

Does ChatGPT train on your data? What the settings actually do.

By James Durkin, JDCS Updated 31 July 2026

People usually want a yes or a no here, and the honest answer is a question back: which plan is your team on? That sounds like a dodge. It is the single most useful thing to understand about AI and your data, and two well-documented episodes from 2025 show exactly why. Neither is a story about a dishonest vendor. Both show where a promise about your data actually lives, and who can move it.

The short version: the training setting on a consumer plan does what it says on the day you read it. What it can't do is bind the vendor's future terms, and it can't bind a court. Which plan you are on decides what happens to your data, because business and enterprise tiers sit on a contract while consumer plans sit on terms the vendor can revise. Both of the cases below prove the same point from opposite directions.

The answer starts with the plan, not the toggle

Consumer plans across the major AI tools generally expose a data control: a switch governing whether your conversations can be used to improve the models. It's a real control and turning it off is worth doing. Business and enterprise tiers work differently, because the commitment lives in a contract rather than a settings page.

That structural difference shows up every time something changes. When a consumer-facing term is revised, the business tiers are consistently carved out. When a legal process reaches a provider, the enterprise arrangements have consistently been the ones that held. Whether that's fair to consumers is a separate argument. For a business deciding where client work can go, it's the operative fact.

So the first thing to establish isn't what the setting says. It's who is signed into what. A firm with a paid business tenant and three staff quietly using personal accounts on their phones has two different data regimes running at once, and only one of them is the one the partners think they bought.

Case one: the terms changed under existing users

On 28 August 2025 Anthropic changed its consumer terms for Claude. Training on user chats went from no to yes unless the user opted out. Retention went from 30 days to five years for anyone who opted in. The change applied to Claude Free, Pro and Max, and existing users were given until 8 October 2025 to make a choice.

The carve-out is the part worth reading closely. Excluded from the change were "Claude for Work, Claude for Government, Claude for Education, or API use, including via third parties such as Amazon Bedrock and Google Cloud's Vertex AI".

Nothing improper happened here. Users were notified, given a deadline and given a choice, which is more than several of the episodes further down this page. The lesson is narrower and more useful: the terms you signed up under are a snapshot, not a fixture. A business that made a decision about Claude in early 2025 based on the data terms of the day would have needed to make that decision again in August. Most businesses have no process that would have caught it.

Case two: a court overrode the delete button

On 13 May 2025, in the copyright litigation brought by The New York Times, Magistrate Judge Ona T. Wang ordered OpenAI to preserve and segregate output log data that would otherwise have been deleted going forward. That included data covered by user deletion requests.

Scope again did the deciding. The order covered ChatGPT Free, Plus, Pro and Team, along with API use that wasn't under a Zero Data Retention agreement. ChatGPT Enterprise was excluded, and Zero Data Retention API customers weren't affected. The order was lifted around 26 September 2025, so it ran for roughly four and a half months.

Sit with what that means for an Australian business. Your privacy posture, your retention policy and your delete button were all overridden for a third of a year by a court in a foreign country, in a dispute you had no relationship with, over a matter you'd never heard of. OpenAI didn't breach a promise. It complied with a court, which is what any company would have to do. The exposure isn't the vendor's character. It's the jurisdiction the vendor sits in, which is the same reason data residency and data sovereignty aren't the same thing, a distinction covered in our guide on whether Copilot is safe for confidential information.

The pattern across the industry, and what it says about defaults

Four episodes are worth knowing, because together they describe how these terms tend to move.

  • Zoom, March 2023. Terms granted a perpetual, sublicensable licence over customer content. Full reversal followed on 15 August 2023, with the chief executive describing it as a process failure.
  • Slack, May 2024. Machine learning training on customer data surfaced publicly, with an opt-out that had to be requested by email.
  • Adobe, June 2024. Terms updated on 18 June 2024 to pledge no training on user content, after a public reaction to the previous wording.
  • LinkedIn, September 2024. Members opted in by default for AI training. The approach was expanded in November 2025 to the EU, EEA, Switzerland, Canada and Hong Kong, again opted in by default.

Three of those four were rolled back under public pressure, which is genuinely reassuring about the direction of travel. Less reassuring is the default. In every case the term was live before anyone outside the company noticed, and it took public attention rather than a customer's contract review to surface it. A 20-person firm reading terms of service is not the mechanism that caught any of these.

Draw the sensible conclusion rather than the cynical one. These companies respond to scrutiny, and the enterprise tiers they sell are meaningfully different products. What they can't offer is a guarantee that today's default survives next year's commercial pressure.

What to actually do about it

Avoiding AI tools entirely costs most businesses more than it saves, so that isn't the recommendation here. Know where your line sits, and put it somewhere you could defend to a client.

  1. Find out what your team is actually signed into. Personal accounts on work devices are the usual gap, and browser extensions are the one people forget entirely.
  2. Read the terms attached to your plan, not a summary. That includes this one. Note the date you read them.
  3. Write down what must never be pasted in. Client files, health records, identity documents, credentials, privileged material, anything under a contract clause restricting AI processing. Give people a concrete list in their own language, because a vague instruction to be careful gets interpreted generously under deadline.
  4. Treat delete as a request. It's honoured in normal circumstances and it can be overridden, so don't build a compliance story on it.
  5. Diarise a re-read. Annually, or at renewal. The Anthropic change arrived with a deadline attached, and the businesses that missed it weren't negligent, they just had nobody watching.
  6. For the work that can't tolerate any of this, keep it in the building. The OAIC's guidance says deploying AI systems locally is "likely to be more privacy-preserving as it limits the risks of third party access to the data". Our local and private AI page covers what that involves and, just as importantly, when it isn't worth it.

For the broader picture on data handling, our guide on whether your business data is safe with AI is a good companion read, and an AI consulting conversation is the fastest way to work out which of your work sits on which side of the line.

Bottom line: the setting is real and worth using, and it's the smaller half of the answer. Your plan tier decides whether a vendor's commitments are a contract or a preference, and a court in another country can override either one for months at a time. Know what your team is signed into, write down what never goes in, re-read the terms on a schedule, and keep the genuinely sensitive work somewhere no third party sits in the path.

Not sure what your team is pasting in?

The first conversation is free. You'll get a plain-English read on which tools your business is actually using, what their terms mean for your data, and where to draw the line.

Start a conversation

Data questions, answered.

Does ChatGPT train on your data?
It depends on which plan you are on and what that plan's terms say today. Consumer plans generally expose a setting that governs whether your conversations can be used to improve the models. Business and enterprise tiers are governed by a contract instead, and across this industry those tiers are consistently carved out of consumer training changes. Read the terms attached to your own plan rather than a blog summary.
If I delete a chat, is it really gone?
Deletion is a request to the vendor, and it can be overridden. On 13 May 2025 a US court ordered OpenAI to preserve and segregate output log data that would otherwise have been deleted, including data covered by user deletion requests. The order applied to ChatGPT Free, Plus, Pro and Team and to API use without a Zero Data Retention agreement. It was lifted around 26 September 2025.
Is ChatGPT Enterprise different from ChatGPT Plus?
Materially, yes, and the 2025 preservation order is the clearest evidence. ChatGPT Enterprise was excluded from its scope and Zero Data Retention API customers were not affected, while the consumer and Team plans were covered. Which plan you are on decides what happens to your data, which is the single most useful thing to understand about AI data handling.
Can an AI company change its data terms after I sign up?
Yes, and there is a well-documented example. On 28 August 2025 Anthropic changed its consumer terms so that chats would be used for training unless the user opted out, with retention extending from 30 days to five years for those who opted in. It applied to Claude Free, Pro and Max, and existing users had to make a choice by 8 October 2025. Claude for Work, Government, Education and API use were excluded.
What should a small business never paste into a public AI tool?
The OAIC recommends against entering personal information, and especially sensitive information, into publicly available generative AI tools. In practice that means client files, health records, identity documents, credentials, legally privileged material and anything covered by a contract clause restricting AI processing. Write that list down and give it to your team, because they will not guess it.