If you remember one thing from this guide, make it this: on every major AI platform, the tier — not the brand — decides what happens to your company data. The consumer version of ChatGPT trains on your conversations unless you opt out. The consumer version of Gemini does the same, and human reviewers may read a subset of chats. The business and API tiers of the very same products do not train on your data by default. Same model, same chat box, opposite data policy.
That gap is the classic SME failure mode: a team lead pastes a client contract into a personal account because the company never bought the business tier. Nothing dramatic happens — which is why it keeps happening.
This guide covers the facts behind that rule for OpenAI, Anthropic and Google, what SOC 2 and ISO badges actually prove (and don't), the PDPA-family rules across ASEAN, and a 10-minute vetting workflow for any AI tool. Still choosing between the big three? Start with the assistant comparison — this guide is about what happens to your data once you've picked one.
The consumer-vs-business divide
| Consumer tier | Business / API tier | Retention | |
|---|---|---|---|
| OpenAI | Trains by default opt-out in Data Controls | No training by default | API: up to 30 days, ZDR available; deleted consumer chats removed within 30 days |
| Anthropic | Trains if you opted in (choice forced Oct 2025) | No training by default | Opt-in consumers: 5 years; declined: 30 days; API: deleted within 30 days |
| Trains by default + human review ("Keep Activity") | No training, no human review (Workspace, paid API) | 72 hours (activity off) / 18 months (on) / up to 3 years if human-reviewed |
OpenAI. Consumer ChatGPT (Free, Plus, Pro) uses your conversations to improve models unless you switch off "Improve the model for everyone" under Settings → Data Controls. ChatGPT Business, Enterprise, Edu and the API are the opposite: no training on inputs or outputs by default (openai.com/enterprise-privacy). API data may be retained up to 30 days for abuse detection, Zero Data Retention is available on eligible endpoints, and the Enterprise, Edu and API tiers offer in-region data storage, including Singapore.
One caveat: retention promises bend to courts. A preservation order in the New York Times litigation temporarily forced OpenAI to retain data that would normally be deleted. The hold on new data has been lifted — deleted and temporary chats again auto-delete within 30 days — but a historical April–September 2025 dataset stays under legal lock while the case continues (openai.com/index/response-to-nyt-data-demands). "Deleted in 30 days" is a policy, not a law of physics.
Anthropic. Claude's consumer plans changed materially in late 2025. Free, Pro and Max users were asked to choose, by 8 October 2025, whether their chats and coding sessions could be used to improve Claude. Allowing it extends retention to five years for new and resumed sessions; declining keeps the previous 30-day retention (anthropic.com/news/updates-to-our-consumer-terms). If your team uses personal Claude accounts, check Privacy Settings → "Help improve Claude" today. The change does not touch commercial products: for Claude for Work, the API, and Claude via Amazon Bedrock or Google Vertex AI, the default remains no training on your inputs or outputs, with API data deleted within 30 days (privacy.claude.com). Zero-data-retention agreements exist, though Anthropic's newest "Covered Models" require 30-day retention for safety review, narrowing ZDR's reach.
Google. Consumer Gemini Apps train on your activity by default when "Keep Activity" is on, and a subset of chats is reviewed by human reviewers — Google's own privacy hub tells you not to enter confidential information (support.google.com/gemini). Retention is 72 hours with activity off, 18 months with it on, and up to three years for human-reviewed chats. Gemini for Google Workspace (Business/Enterprise) flips all of that: submissions aren't used to train models and aren't reviewed by humans. The trap is the Gemini Developer API's unpaid tier: its terms say Google uses submitted content to "provide, improve, and develop" its products and machine-learning technologies, that human reviewers may read your input and output, and that you should not submit sensitive, confidential, or personal information (ai.google.dev/gemini-api/terms). The paid tier does not train on your prompts. Any "free AI tool" built on the unpaid tier inherits the unpaid terms.
What the certificates actually mean
SOC 2 Type II is not a certificate — it's an attestation report. An independent CPA firm examines a vendor's controls against the AICPA Trust Services Criteria (security, availability, processing integrity, confidentiality, privacy); Type II means the controls were tested over a period, not just on one day. Because it's a report, the details matter: scope, audit window, and any exceptions noted. A vendor that won't share the report (usually under NDA via its trust center) is asking you to take the badge on faith.
ISO/IEC 27001:2022 certifies an Information Security Management System — a risk-managed system for protecting the confidentiality, integrity and availability of data the organisation handles (iso.org/standard/27001). It certifies the management system, not any product feature, so always read the certificate's scope statement.
ISO/IEC 42001:2023 is the newer badge on AI vendors' trust pages: the first certifiable AI management system standard, covering responsible development and use, accountability and transparency. It signals AI-governance maturity and says nothing about whether the vendor trains on your data.
For ISO claims, ask for the certificate and validate it on IAF CertSearch, the IAF's database of accredited certifications (free for a handful of look-ups a day) — search by company or certificate number and confirm the scope names the service you're buying. For SOC 2 there is no public register; the only real check is obtaining and reading the report. And the overriding point: no certification overrides the contract. A SOC 2-attested vendor can lawfully train on your data if its terms allow it. The terms and the DPA are what bind them.
The ASEAN compliance anchors
Singapore. The PDPC's Advisory Guidelines on personal data in AI recommendation and decision systems (1 March 2024) confirm that PDPA consent and notification obligations still apply when AI is involved, with the Business Improvement and Research exceptions available in defined cases (pdpc.gov.sg). In June 2026 the PDPC opened a public consultation on proposed guidelines specifically for generative AI; final guidance is expected after it closes. The practical point: sending customer or staff data to an AI vendor is an outsourcing-and-transfer scenario — you stay responsible under the Protection Obligation, the Transfer Limitation Obligation requires comparable protection overseas, and the vendor's DPA is the artefact that evidences it.
Malaysia. The PDPA (Amendment) Act 2024 came into force in three stages through 2025: ancillary provisions on 1 January; direct obligations on data processors, "data controller" terminology, biometric data as sensitive data and higher penalties on 1 April; and, from 1 June 2025, mandatory Data Protection Officer appointment (above thresholds) plus breach notification to the Commissioner "as soon as practicable" (pdp.gov.my). If your AI vendor leaks data, breach notification is now a live legal obligation, not a best practice.
Indonesia. The PDP Law (UU No. 27/2022) has been fully in force since 17 October 2024 — a GDPR-style regime with legal bases, DPO requirements in some cases, breach notification and cross-border transfer conditions. The implementing regulation had still not been enacted at the time of writing (July 2026), and there is no dedicated Data Protection Agency yet; supervision sits with Komdigi, with the agency targeted for 2026. The law applies now — the enforcement machinery is still being built.
The 10-minute vetting workflow
- Find the trust center (2 min). Search "vendor trust center" — trust.openai.com, trust.anthropic.com, cloud.google.com/security/compliance. No trust center or security page at all on a paid product weighs heavily against.
- Find the data-training clause (2 min). In the privacy policy or service terms, Ctrl-F "train", "improve", "machine learning". You want an explicit "we do not use your data to train" for the tier you are buying — and, if training is default-on, the location of the opt-out.
- Check retention (1 min). Ctrl-F "retention", "delete", "30 days". Good answers are specific numbers per tier. "As long as necessary" with no numbers is a flag.
- Check the sub-processor list (1 min). Who actually touches your data — cloud hosts, human-review contractors — and where. OpenAI, Anthropic and Google Cloud all publish theirs.
- Check DPA availability (2 min). Anthropic's DPA is auto-incorporated into its commercial terms; OpenAI's is signable online. A paid business plan with no DPA at all can't lawfully process personal data for you under PDPA/GDPR-style regimes. Hard stop.
- Verify the certification claims (2 min). ISO certificates → IAF CertSearch, and check the scope statement. SOC 2 → request the report via the trust center and check the CPA firm, Type I vs Type II, audit period, scope and exceptions.
Decision rule: consumer tier + training on + no DPA → personal experiments only, never company or customer data. Business/API tier + no-training default + 30-day-or-less retention + DPA + verifiable SOC 2 or ISO 27001 → acceptable for most SME data, with residency where clients or regulators require it.
Nine red flags
- "We may use your content to provide, improve, and develop our products" with no opt-out and no tier carve-out. That is the Gemini API unpaid-tier language almost verbatim. Fine as a labelled free tier; a red flag when it's the only tier.
- Training-by-default on the tier you actually use. ChatGPT consumer ("Improve the model for everyone") and Gemini Apps ("Keep Activity") both default on.
- Human review buried in the policy. "Human reviewers may read, annotate, and process your API input and output." If humans can read it, it must never contain client data.
- No retention numbers. "Retained as long as necessary for business purposes" — versus vendors who publish 30-day, 72-hour and 18-month figures per tier.
- No DPA on a paid business plan — or a DPA gated to "Enterprise only" while the Team tier is marketed to businesses.
- No sub-processor list, so you can't know which country or contractor sees the data, or no change-notification mechanism.
- Certification name-dropping without artefacts. "SOC 2 and ISO certified" badges with no report, no certificate number, nothing on IAF CertSearch — or a certificate whose scope doesn't cover the product sold.
- The vendor's own do-not-share warning. When the terms say "do not submit sensitive, confidential, or personal information," take the vendor at its word: that tier is telling you it isn't for company data.
- Free-tier bait-and-switch in wrappers. Many SME-facing "AI tools" are thin wrappers that may call an upstream vendor's unpaid API. Ask which tier they use, and for their own DPA and sub-processor list — the upstream LLM vendor should appear on it.
What this guide doesn't cover
This is a vetting workflow, not legal advice — for contract review or a regulator-ready Data Protection Impact Assessment, engage counsel in your jurisdiction. Sector-specific regimes (banking outsourcing, healthcare data, government classifications) add obligations beyond this guide, as does AI output risk — accuracy, IP and confidentiality of generated content. For ongoing incidents and vendor security news, follow our cybersecurity coverage.
Verification note: every claim was checked against the cited official vendor and regulator pages in July 2026. Two caveats: openai.com and iso.org block automated access, so those pages were confirmed through search excerpts of the official pages rather than direct fetches — spot-check the live pages before relying on exact wording contractually — and Indonesia's PDP status draws on reputable legal-practice summaries, as the official texts are in Bahasa Indonesia. Vendor policies changed materially twice in twelve months; re-run the 10-minute workflow before every renewal.