Does Google Cloud Speech-to-Text Work in China? PIPL Cross-Border, Voiceprint Biometrics & Data Residency
The audio you send Google Cloud Speech-to-Text to transcribe carries the speaker's voiceprint — Article 28 sensitive biometric personal information you cannot anonymize — and Google operates no mainland-China region, so that audio is processed offshore: a PIPL cross-border transfer. A compliance-first look at the voiceprint, data residency, retention, and the lawful in-country speech path.
Does Google Cloud Speech-to-Text work in China?
The recording you send Google Cloud Speech-to-Text to transcribe is the speaker's voiceprint — Article 28 sensitive biometric personal information — and because Google operates no mainland-China region, every request processes that audio offshore.
Speech-to-Text turns audio into text, but the audio itself carries a biometric identifier plus whatever was said — call recordings, voice messages, dictated notes — so sending China-origin audio to Google's offshore servers is a PIPL cross-border transfer of sensitive personal information (Article 28) that redacting the transcript cannot un-make. By default Google does not log or train on your audio, but an opt-in data-logging program does — a setting you disable for China audio. The lawful lever is to keep China audio's processing in-country — a self-hosted/on-prem speech container or a licensed domestic speech service — with the Article 28 separate consent and a PIPIA, not to make the offshore endpoint reachable.
A risk map, not a verdict — settle the specifics with counsel. Our China team can map your exposure →
What Google Cloud Speech-to-Text's own documentation says about China
| Fact | Primary source |
|---|---|
| Google Cloud Speech-to-Text has no mainland-China region — audio is processed offshore. Google's documentation offers only "US and EU regional API endpoints," where "your data at-rest and in-use will stay within the continental boundaries of Europe or the USA," and the V2 API adds regions "such as Belgium or Singapore." Google runs no cloud region on the Chinese mainland (the nearest are asia-east1 in Taiwan and asia-east2 in Hong Kong, both offshore) — and no China partition like other clouds operate. | Google Cloud — Speech-to-Text supported regional endpoints (cloud.google.com), retrieved 2026-10-10 |
| By default Google does not log or train on your audio — but an opt-in program does. Google states "By default, Cloud Speech-to-Text does not log customer audio data or transcripts"; its opt-in data logging program, by contrast, lets "Google ... use this data solely to train and improve Google products and services." For China-origin audio, keep logging disabled so recordings and transcripts are not retained or trained on offshore. | Google Cloud — Speech-to-Text data logging (cloud.google.com), retrieved 2026-10-10 |
| An on-premises container build exists — but it is a private, sales-gated feature. Google offers Cloud Speech-to-Text On-Prem, "a Google Cloud Marketplace application" that "can be deployed as a container to any GKE cluster" and run "on-premises with Anthos." Google's own page states "you must contact Google for full access," so an in-country self-hosted path is possible, but availability and terms are confirmed with Google rather than toggled on. | Google Cloud — Cloud Speech-to-Text On-Prem overview (cloud.google.com), retrieved 2026-10-10 |
| A voiceprint is sensitive biometric personal information under PIPL Article 28. Transcribing a China speaker's audio offshore is a cross-border transfer under PIPL Articles 38–40 (notice, a separate consent, and a transfer mechanism), while the voiceprint in that audio is sensitive personal information under Article 28 — requiring a separate specific consent and a prior protection impact assessment (PIPIA) that you cannot escape by redacting the text. | Personal Information Protection Law of the PRC, Articles 28 and 38–40 (cac.gov.cn), retrieved 2026-10-10 |
Sources verified by the 21YunBox compliance team on 2026-10-10.
For a mainland-China audience, the question to settle about Google Cloud Speech-to-Text is not whether the API can be reached — it is what happens to the audio you send it. Speech-to-Text is a transcription service: audio of someone speaking goes in, text comes back. But a recording is more than its words. A human voice is a biometric identifier — a voiceprint — and under China’s Personal Information Protection Law a voiceprint is sensitive personal information (Article 28) that redacting the transcript cannot anonymize away. The spoken content — support calls, voice messages, dictated medical or legal notes — is often sensitive in its own right. Google operates no mainland-China region, so every request is processed on its offshore servers. Three duties therefore bite at once: a PIPL cross-border transfer (Articles 38–40), the sensitive-biometric consent-and-assessment bar (Article 28), and content residency for a critical-information-infrastructure or high-volume handler.
Google Cloud Speech-to-Text in China at a glance
| What decides it | In Google Cloud Speech-to-Text's own terms — and China's law |
|---|---|
| What you send | Speech-to-Text takes audio of a person speaking and returns text. Every call uploads the raw recording — the speaker's voiceprint (a biometric identifier) plus whatever was said. You can redact a transcript; you cannot anonymize a voice. |
| Where it goes | Google runs no mainland-China region. Its documentation offers only US and EU regional endpoints (and regions "such as Belgium or Singapore"); the nearest datacenters are asia-east1 (Taiwan) and asia-east2 (Hong Kong), both offshore of the mainland. Sending China-origin audio there is a cross-border transfer of personal information (PIPL Articles 38–40, 数据出境). |
| A voiceprint is sensitive data | Under PIPL Article 28, biometric characteristics — including a voiceprint — are sensitive personal information. Processing and exporting them needs a separate specific consent and a prior personal-information protection impact assessment (PIPIA). Spoken content may add further sensitive categories (health, finance, ID). |
| Retention, training & residency | By default Google "does not log customer audio data or transcripts," but an opt-in data-logging program lets it use logged audio to "train and improve Google products and services" — keep it off for China audio. An on-premises container build exists (Cloud Speech-to-Text On-Prem, a private feature). For a CIIO or high-volume handler, Cybersecurity Law Article 39 (formerly Article 37) adds in-country storage. |
| The axis | Reachability is not the question. The lawful lever is to keep China-origin audio's processing in-country — a self-hosted/on-prem speech container or a licensed domestic speech service — with the Article 28 consent, not to make the offshore endpoint faster. The China-facing app that captures the audio also carries an ICP filing duty and needs in-country delivery. |
What you actually send — your speakers’ voiceprints and words
Speech-to-Text is a transcription service: you stream or upload audio to a Google endpoint and receive text. The output is text, which makes it easy to assume the privacy question is about the transcript. It isn’t. The thing that travels on every request is the recording — and a recording of a person speaking is a biometric identifier of that person. A transcript you can redact; a voiceprint you cannot un-say. On top of the voiceprint rides the content itself: a customer-service call, a voicemail, a doctor’s dictation, a deposition, an internal meeting — categories that are routinely sensitive before you even reach the biometric point.
Where does that audio go? Google processes Speech-to-Text in Google Cloud regions, and it has no region on the Chinese mainland. Its own documentation offers “US and EU regional API endpoints,” where “your data at-rest and in-use will stay within the continental boundaries of Europe or the USA,” and the V2 API adds regionalized processing in regions “such as Belgium or Singapore.” Choose any of them to serve a user in China and the audio leaves the country. On retention, Google’s position is favorable but conditional: “By default, Cloud Speech-to-Text does not log customer audio data or transcripts,” yet it also runs an opt-in data-logging program under which “Google uses this data solely to train and improve Google products and services.” For China-origin audio that opt-in is a switch you keep off — but keeping it off does not change the fact that the processing itself happens offshore.
It’s a cross-border transfer of sensitive biometric data — under PIPL
Because the audio is processed outside the mainland, sending a China speaker’s recording to Speech-to-Text is a cross-border transfer of personal information under China’s Personal Information Protection Law. PIPL Articles 38–40 put the duty on you, the handler — not on Google — and require notice, a separate consent distinct from the user’s agreement to use the feature, and one lawful transfer mechanism: a CAC security assessment, the CAC standard contract, or certification.
The voiceprint raises the bar further. Under PIPL Article 28, biometric characteristics are sensitive personal information, and processing them demands a separate specific consent and a prior personal-information protection impact assessment (PIPIA) before the data is processed or transferred. A voice recording is biometric by its nature; stripping names from the transcript does not change what the audio is. Where the spoken content is itself health, financial, or identity data, additional sensitive-category duties stack on top.
Underneath sits residency. A critical information infrastructure operator, or a handler above the volume thresholds, must store personal information collected in China inside the mainland — PIPL Article 40 together with the Cybersecurity Law (Article 39 in the amendment in force since January 1, 2026, which renumbered the data-localization article from Article 37 to Article 39 with the substance unchanged) — a duty an offshore region structurally cannot meet. Where volumes or sensitivity cross the thresholds, the export itself needs a CAC data-export security assessment before it may proceed. None of this is about how fast a transcript returns.
Reaching the API isn’t the question — keeping the audio in-country is
The lever is not to make the offshore endpoint faster or more reachable; it is to stop exporting China-origin audio and to process it on an in-country, consented path. For speech-to-text there are two honest in-country options. First, Google publishes its own on-premises build: Cloud Speech-to-Text On-Prem, “a Google Cloud Marketplace application” that “can be deployed as a container to any GKE cluster” and run “on-premises with Anthos” — so recognition can execute on infrastructure you place inside the mainland instead of calling Google’s offshore regions. It is a private feature; Google’s own page states “you must contact Google for full access,” so treat availability and terms as something you confirm with Google directly, not a self-serve toggle. Second, where that path is not open to you, route China-origin audio through a licensed in-country speech service — a domestic speech provider with the data kept in the mainland.
Either way, the duties travel with the audio: obtain the Article 28 separate consent, run the PIPIA, keep the opt-in data-logging program off so no China recording is retained or used to train, and minimize what you capture. What you do not do is build a tunnel that ships the audio offshore and call it local. This is a risk map, not a verdict — settle the specifics, which path, which consent, which thresholds, with counsel.
The lawful path — map, localize, deliver
There is a lawful way to run speech-to-text for your users in China, and it has a clear shape: China-origin audio is processed from inside the country on a consented footing, the voiceprint is treated as the sensitive data it is, and the China-facing app that captures the audio is itself filed and delivered in-country. The part 21YunBox owns is that footing, and it is more than advice. Our China team does three things. We map what audio flows to the speech service — whose voiceprints and what spoken content it carries, where it is processed and whether it is retained or used to train, and where you lack a lawful basis (the Article 28 separate consent and PIPIA, a transfer mechanism). We localize the processing onto an in-country path — a self-hosted or on-premises speech container, or a licensed domestic speech service — so China audio stays in the mainland instead of being exported, with consent and retention handled correctly. And we deliver the China-facing app that captures and surfaces the audio on ICP-filed, in-country infrastructure — the 21YunBox Optimizer — in front of the stack you already run, with no rebuild and no re-platform.
The result is a speech feature that runs legally and compliantly for your users in China. 21YunBox never uses or suggests circumvention of any kind. This is a lawful, data-resident in-country deployment, and where a service is not offered in the mainland we localize onto a licensed domestic equivalent rather than reach offshore.
Related reading:
- Cross-border data transfers under PIPL
- China’s Cybersecurity Law — Article 39 (formerly Article 37) and data localization
- China’s data-export security assessment measures
- How to get an ICP filing for China
