In our guide on Sun Tzu and the AI war I argued that 地 — ground, terrain — is the decisive factor in this technology, and that jurisdiction is the ground. Then I gave it four sentences and moved on.
This is that debt.
Every business owner in this region eventually asks the same question, usually after somebody has already pasted a customer list into a chat box: where does that text go, who can read it, for how long, and whose law decides?
The distinction almost nobody has been told
"Do you use my data?" is two questions wearing one coat, and they have different answers.
Training is about whether your text gets absorbed into a future version of the model. Most people fixate on this, and on business tiers the answer is usually no by default.
Retention is about how long a copy sits on the provider's systems for abuse monitoring and support. This is almost always a separate, fixed period, and it is usually not a setting you control.
Both statements can be true at once, and they routinely appear on the same page: we do not train on your data, and we retain inputs and outputs for a period. A company that reads only the first sentence has answered the wrong question. The customer list is not in the model. It is on a disk, in a country, reachable by whoever can compel it there.
The third question, sitting under both, is whose courts can reach the copy. This isn't decided by a setting, but by geography and contract.
What "we don't train on your data" is actually promising
Before you rely on that promise, separate three things.
Then check it yourself, because the summaries are wrong
Here I have to be careful, and to disclose something. This guide is drafted with a model built by Anthropic, one of the companies whose terms it discusses. So rather than characterise anybody's policy from memory, I went to the primary pages — and what happened next is the most useful thing in this guide.
Several widely-circulated comparison articles state that Anthropic began training on consumer conversations by default from late September 2025. Anthropic's own documentation, last updated 16 March 2026, describes it as a choice the user makes: conversations are used for model improvement if "you choose to allow us to use your chats and coding sessions to improve Claude", or if they are flagged for safety review, or if you explicitly opt in through something like a tester programme. It also says "your Incognito chats are not used to improve Claude, even if you have enabled Model Improvement", and that feedback data is kept "in our secured back-end for up to 5 years".
I am not going to adjudicate whether the comparison sites were wrong when written or have simply gone stale — policies do change, and that is precisely the point. The point is that the summary and the source disagreed. The source is your contract.
The second finding was quieter. When I tried to pull another major provider's policy pages the same way, they returned 403 Forbidden to automated access. There is nothing sinister in that — plenty of sites block scrapers — but it means the terms you are being asked to rely on are, in practice, checkable only by a human opening a browser. Any comparison table you find has been assembled by someone doing that by hand, at some point in the past, and it decays from the moment it is published.
If you are naming a product in a policy, link the vendor's own page and record the date you read it. Our directory entries for Claude, ChatGPT and Gemini CLI point at the primary sources rather than restating them, for the same reason.
Where it physically runs is a legal question, not a technical one
Engineers hear "where does it run?" and think about latency. Lawyers hear it and think about who can serve process. For this question, the lawyer is right.
Data stored in a jurisdiction is generally reachable by that jurisdiction's legal instruments, regardless of who owns the company. This fact renders "our provider is a reputable firm" useless as an answer to "which government can compel this". It is also why enterprise agreements increasingly specify a processing region, rather than leaving it to the provider's convenience.
It is also why the honest answer to "is this allowed?" in Southeast Asia is: it depends which of ten countries you are asking about.
Ten countries, seven laws, no shared definition
Seven of the ten member states now have a comprehensive personal data protection law. Brunei is the newest — its Personal Data Protection Order was enacted in January 2025 with most substantive provisions commencing on 1 January 2026, phased over a further year. Cambodia, Laos and Myanmar have no comprehensive law.
There is no ASEAN-wide definition of data sovereignty. The 2016 ASEAN Framework on Personal Data Protection sets out principles — consent, security, accuracy — and is not binding. It does not harmonise the underlying laws, and the place they diverge most sharply is exactly the place that matters here: cross-border transfer.
Two developments matter.
- Vietnam now regulates AI directly. Its Personal Data Protection Law took effect on 1 January 2026. Decree 356/2025, issued on the last day of 2025, goes further, explicitly addressing AI systems, big data analytics, and cloud with requirements for stronger controls and purpose limitation. A separate AI Law, adopted in December 2025, took effect in March 2026. Vietnam is not waiting for the region.
- Malaysia's constraint is not where you would look for it. The PDPA, amended in 2024, does not mandate that personal data be stored physically in Malaysia; transfers require safeguards instead. The sharper constraint on regulated firms comes from Bank Negara's technology risk policy, which requires risk assessment before outsourcing material IT services, contractual audit rights, and arrangements that do not impede the regulator's own oversight.
That second point corrects something we had recorded internally as an "in-country storage mandate". It is not one. The obligation is about control and auditability rather than geography — which is a more demanding requirement in some ways and a less demanding one in others, and either way is the wrong thing to solve by picking a data centre.
What to actually do
- Decide the tier before the tool. Consumer and business versions of the same brand are different contracts. If staff will handle customer data, the free app is not the thing you evaluated.
- Read the retention clause, not the headline. Find the number of days and whether you can change it. That figure is what an incident, a subpoena or an audit will turn on.
- Write down the date you read it. Every statement in this guide carries one. Yours should too, because the next person to check will need to know whether anything has moved since.
- Ask which region processes it, and get it in writing. Not who the vendor is — where the work happens. This is the question that decides whose courts are involved, and it is answerable in a contract.
- Assume no regional answer exists. If you operate across ASEAN, "compliant in Singapore" tells you nothing about Vietnam or Indonesia. Seven regimes, and they diverge most on transfer.
"Do you use my data?" is two questions. Training is whether your text enters a future model; on business tiers, this is usually a setting that defaults to off. Retention is how long a copy sits on the provider's systems; this is usually a fixed period you cannot change. A provider can therefore truthfully say it does not train on your data while still holding a copy for months. The question underneath both is whose courts can reach that copy, and geography decides it.
Check the vendor's own page rather than a comparison table. When I did that for this guide, widely-repeated summaries of one provider's consumer training policy disagreed with that provider's own documentation, and another provider's policy pages could not be read by automated tools at all. There is no single answer to this across ASEAN. Seven of the ten member states have a comprehensive data protection law, but three do not. There is no regional definition of data sovereignty, and the national laws diverge most on cross-border transfer. Vietnam now regulates AI directly. Malaysia's real constraint on regulated firms is auditability rather than the in-country storage rule people expect.
Disclosure: this guide was drafted with a model made by Anthropic, whose terms it discusses and quotes. Every provider statement here was read from the provider's own documentation on 29 July 2026 and is dated accordingly; these terms change, sometimes without announcement, and a policy is only as current as the day you checked it. Nothing here is legal advice — for a regulated firm, or for anything with a filing obligation attached, take proper counsel in each jurisdiction you operate in.
- Anthropic — Consumer Data Use for Model Training, page last updated 16 March 2026: the opt-in framing, the Incognito exclusion and the five-year feedback retention are quoted from it. Accessed 29 July 2026.
- A second major provider's privacy and training documentation returned HTTP 403 to automated retrieval on 29 July 2026, so it is described structurally here rather than quoted. This is stated because it affects how checkable the claim is, not as criticism.
- Vietnam: Law No. 91/2025/QH15 on Personal Data Protection, in force 1 January 2026; Decree No. 356/2025/ND-CP issued 31 December 2025 covering AI, big data and cloud; AI Law adopted December 2025, effective March 2026. Accessed 29 July 2026.
- Brunei: Personal Data Protection Order 2025, enacted 8 January 2025, most substantive provisions commencing 1 January 2026 with phased compliance. Accessed 29 July 2026.
- Malaysia: Personal Data Protection Act 2010 as amended by the Personal Data Protection (Amendment) Act 2024; Bank Negara Malaysia's Risk Management in Technology policy for regulated financial institutions. Accessed 29 July 2026.
- ASEAN coverage and the absence of a regional data-sovereignty definition: the ASEAN Framework on Personal Data Protection (2016) and regional practice surveys. Accessed 29 July 2026.