Does Cohere Work in China? Prompts, RAG Corpora, the Generative-AI Filing & Cross-Border Data
Cohere is reachable from mainland China, but runs no in-country region — its managed inference sits on Cohere's own offshore infrastructure in Canada, so every prompt, RAG document, and fine-tuning set from a mainland user is a PIPL cross-border transfer, and a public AI feature turns on a CAC filing. A compliance-first look at Cohere and the lawful China AI path.
Does Cohere work in China?
Cohere is reachable from the mainland, but it runs no mainland-China region — so the question isn't whether the API answers, it's whether your data may lawfully leave the country and whether your public feature needs a filing.
Cohere's managed inference runs on its own offshore infrastructure (the company is based in Toronto, Canada), so every prompt, retrieved RAG document, fine-tuning set, and generated output sent from a mainland user is a PIPL cross-border transfer of personal information — and often trade secrets. On its SaaS platform those prompts and generations are used to train Cohere's models unless you opt out. Separately, any public-facing generative-AI feature you offer to mainland users turns on the CAC Interim Measures filing, which the upstream vendor does not hold for you; Cohere's private VPC/on-prem path narrows where inference runs but is not automatic compliance.
This is a risk map, not a ruling — settle specifics with counsel. Our China team can map your exposure →
What Cohere's own documentation says about China
| Fact | Primary source |
|---|---|
| On Cohere's managed SaaS platform, your prompts and generations are used to train Cohere's models unless you opt out. Cohere's Enterprise Data Commitments state, "You can opt out from your prompts and generations being used to train Cohere models in your dashboard settings at any time," and that SaaS prompts and generations are "logged and automatically deleted after 30 days"; no-logging applies only to accounts approved for zero data retention. For a mainland user, that default is both a PIPL cross-border transfer and a training use of their data. | Cohere — Enterprise Data Commitments (retrieved 2026-10-10) |
| Cohere runs no mainland-China region: managed inference is hosted on Cohere's own infrastructure, and the private VPC/on-prem path is a residency lever, not automatic compliance. Cohere's deployment documentation states, "The models are hosted on Cohere infrastructure and available on our public SaaS platform," and that in private and third-party-cloud deployments "Cohere does not receive any customer inputs (prompts) or outputs (generations)" — while noting it "may also receive certain usage data" such as "aggregate counts of input prompt tokens." None of these is a mainland-China region, and the model and its control still sit with an offshore (Toronto, Canada) vendor. | Cohere — Deployment Options Overview & Enterprise Data Commitments (retrieved 2026-10-10) |
| Sending a mainland user's prompts, RAG documents, or fine-tuning data to Cohere's offshore inference is a cross-border transfer of personal information under PIPL. PIPL Articles 38–40 require notice, a separate consent, and one transfer mechanism (a CAC security assessment, the CAC standard contract, or certification); where a prompt or document contains ID numbers, financial, health, biometric, or minors' data, PIPL Article 28 adds sensitive-personal-information duties. The obligation falls on you, the feature operator — not on the upstream model vendor. | Personal Information Protection Law (PIPL), Articles 28 and 38–40 |
| A public-facing generative-AI feature offered to mainland users turns on China's CAC Interim Measures — a filing the upstream API vendor does not hold for you. The Interim Measures for the Management of Generative AI Services (生成式人工智能服务管理暂行办法, CAC, in force August 15, 2023) require a filing/registration, content-safety controls, and labeling of AI-generated content for services offered to the public in the mainland. Choosing zero retention or a private deployment narrows what crosses the border, but it does not remove the filing duty. | Interim Measures for the Management of Generative AI Services (CAC, effective 2023-08-15) |
Sources verified by the 21YunBox compliance team on 2026-10-10.
For a product aimed at mainland China, the first thing to settle about Cohere is not how fast its API answers — it usually will — but whether the data behind each call may lawfully leave the country. Cohere is an enterprise LLM platform: its Command and Command R/R+ generation models, Embed, Rerank, and the North retrieval-and-agent workspace are offered both as a managed API and for private deployment. The company is based in Toronto, Canada; its managed inference runs on its own offshore infrastructure; and it publishes no mainland-China region. What you feed it is rarely trivial — prompts, the internal documents you embed and retrieve as a RAG corpus, fine-tuning sets, and the generated outputs — and that content routinely carries personal information and often trade secrets. So the decision is a compliance one, and it turns on two doors: whether that data may cross the border, and whether a public AI feature needs a filing.
Cohere in China at a glance
| What decides it | In Cohere's own terms — and China's law |
|---|---|
| Where does inference run? | Cohere is a Toronto, Canada company. Its managed Command, Embed, Rerank, and North services are, in Cohere's words, “hosted on Cohere infrastructure and available on our public SaaS platform”; the cloud options (Amazon Bedrock and SageMaker, Azure AI Foundry, OCI) run on the provider's own infrastructure. None is a mainland-China region, so reaching Cohere from the mainland sends data offshore. |
| What you send it — and why it's personal information | Prompts, the documents you embed and retrieve as a RAG corpus (often your most sensitive internal files), fine-tuning data, and the model's outputs. These carry names, contacts, customer and employee records, source code, and trade secrets — personal information under PIPL, and sensitive personal information under Article 28 where ID numbers, financial, health, biometric, or minors' data appear. |
| Your users' prompts = a cross-border transfer | From a user in China to Cohere's offshore inference, every prompt, retrieved document, and fine-tuning set is a cross-border transfer of personal information under PIPL (Articles 38–40): notice, a separate consent, and one transfer mechanism. On the managed SaaS platform, prompts and generations are also used to train Cohere's models unless you opt out, and are logged and deleted after 30 days. |
| The generative-AI filing gate | Offering a public-facing generative-AI feature to mainland users turns on the CAC Interim Measures for the Management of Generative AI Services (生成式人工智能服务管理暂行办法, in force Aug 15, 2023) — a filing/registration, content-safety controls, and labeling of AI-generated content. The upstream API vendor does not hold that filing for you. |
| Reachability isn't the axis — the lawful path | A responding endpoint changes nothing: the data still left the country, and a public feature still needs its filing. Any China-facing surface also carries an ICP filing (备案) duty. 21YunBox maps the exposure, localizes onto a CAC-filed domestic or in-China sovereign-cloud model, and delivers it in-country — not a route to Cohere's offshore model. |
No mainland-China region, so your prompts and documents leave the country
Cohere’s position is set by where it runs, not by a load-time test. Its own deployment documentation says the managed models are “hosted on Cohere infrastructure and available on our public SaaS platform,” and that the cloud options run on the provider’s own infrastructure; Cohere describes itself as cloud-agnostic, deployable “through any cloud provider.” What none of those is, is a mainland-China region. Cohere is a Toronto company, and its managed inference sits offshore.
That makes the familiar “does it load from Shanghai?” test the wrong question. Whether a request completes on a given day does not change the fact that the prompt, the retrieved document, and the fine-tuning file all had to leave the mainland to reach the model — which is why this page publishes no first-party China latency figure for Cohere; speed is not the axis for a service operated from another jurisdiction. Cohere does offer a genuine residency lever — a private VPC or on-premises deployment, including air-gapped installs — and we weigh it honestly below. But in-country deployment is not the same as compliance, and it is not a way to reach the offshore model from inside China.
What you send it is personal information — prompts, your RAG corpus, and fine-tuning data
A prompt is seldom just a question. With Cohere specifically, the inputs go further than chat: Embed and Rerank are built to take your own documents — knowledge bases, contracts, support tickets, source code, customer records — turn them into vectors, and rank them, and North assembles them into retrieval-augmented answers and agent actions. That RAG corpus is frequently the most sensitive material a company holds, and fine-tuning data is more sensitive still. Sent to Cohere’s offshore inference, all of it is personal information under China’s Personal Information Protection Law wherever it names or identifies people, and it is sensitive personal information under PIPL Article 28 where it contains ID numbers, financial, health, biometric, or minors’ data — each of which raises the bar to a specific, separate consent and a necessity test. It is also, very often, trade secrets.
How that data is used matters too. On Cohere’s managed SaaS platform, training on your inputs is the default: its Enterprise Data Commitments tell you that you “can opt out from your prompts and generations being used to train Cohere models,” which means that until you do, they are in scope; SaaS prompts and generations are “logged and automatically deleted after 30 days,” and no-logging applies only to accounts approved for zero data retention. Cohere’s private and third-party-cloud deployments are better on this point — there, Cohere says it “does not receive any customer inputs (prompts) or outputs (generations)” — but even then it notes it “may also receive certain usage data” such as “aggregate counts of input prompt tokens.” The residency lever narrows what crosses; it does not make the crossing disappear.
The filing gate — and why trimming retention doesn’t close the door
Suppose you still want the assistant, the search box, or the “summarize this” button in a product aimed at mainland users. The first gate is not which model sits behind it — it is whether you may offer a public-facing generative-AI service in China at all. China’s Interim Measures for the Management of Generative AI Services (生成式人工智能服务管理暂行办法, Cyberspace Administration of China, in force since August 15, 2023) are written for exactly this: they reach the use of generative AI “to provide services for generating text, images, audio, video, and other content to the public within the territory of the People’s Republic of China” (Article 2), and they carry duties on training data, personal-information handling, a filing/registration for services able to shape public opinion, and — under the 2025 labeling rules — conspicuous labels on AI-generated content. Where your feature ranks or recommends content to users, the Algorithm Recommendation Provisions filing may also apply. The upstream API vendor does not hold any of these for you.
Here is the part that trips teams up: tightening how Cohere handles your data does not clear this gate. Switching on zero data retention, picking a particular cloud, or standing the model up in your own VPC all change what crosses the border or where inference runs — they do not change that a public generative-AI feature offered in the mainland needs its own filing, nor that the user data still needs a lawful basis under PIPL to leave the country in the first place. And if your RAG corpus includes data a critical information infrastructure operator must keep onshore, Cybersecurity Law Article 39 (formerly Article 37) — the 2025 Cybersecurity Law amendment, in force January 1, 2026, renumbered the data-localization article from 37 to 39, substance unchanged — can bar the transfer outright. None of this is a verdict on your specific feature: it is a map of where the exposure sits, and the specifics — which filings, which consents, whether a transfer is lawful at all — are for your counsel to settle against what you actually ship.
The lawful path — map, localize, deliver
There are lawful ways to put generative AI in front of users in China, and they share one shape: the model is served from inside China by an operator that holds the filings, and the user data stays on a compliant footing. Two patterns recur. One is a CAC-filed domestic model that already carries the generative-AI filing — Baidu Ernie (文心一言) or Alibaba Qwen (通义千问) among them. The other is an in-China sovereign-cloud model operated in the mainland by a licensed local operator, subject to that operator’s own availability and compliance, which you confirm with them directly. Cohere’s own private VPC or on-prem deployment can reduce where inference runs and keep your inputs out of Cohere’s hands — a useful residency step — but it does not by itself discharge the filing duty, it still relies on an offshore vendor’s model and control, and it is not a substitute for the legal analysis above.
Underneath either choice sits the part 21YunBox owns, and it is more than advice. The China-facing surface that carries the AI — the chat panel, the search box, the API edge your mainland users hit — is itself a public service in the mainland, so it carries an ICP filing (备案) duty and needs compliant, in-country delivery like any other China-facing property. Our China team works in three moves. We map your PIPL cross-border, sensitive-data, generative-AI-filing, and residency exposure against your entity, your data, and who your users are. We localize the feature onto a lawful footing — standing up and integrating a CAC-filed domestic model or an in-China sovereign-cloud model in place of the call that cannot lawfully be served offshore, and keeping consented, in-country processing for what must stay on mainland soil. And we deliver the surface in-country on ICP-filed infrastructure — the 21YunBox Optimizer — in front of the stack you already run, with no rebuild and no second codebase. The result is a generative-AI feature that runs legally and compliantly for your users in China. This is a compliant overlay, not a migration — and a partner to your AI vendor, not a competitor. What we do not do — and what no one lawfully can — is turn an offshore model into a mainland one; we stand up and deliver a lawful equivalent. 21YunBox never uses or suggests circumvention of any kind.
Related reading:
- China’s Interim Measures for the Management of Generative AI Services
- Cross-border data transfers under PIPL
- How to get an ICP filing for China
- China’s Cybersecurity Law and data localization
