Does Amazon Polly & Transcribe Work in China? PIPL Cross-Border, Voiceprint Biometrics & Data Residency
The audio you send Amazon Transcribe, and the voices Amazon Polly synthesizes, carry the speaker's voiceprint — PIPL Article 28 sensitive biometric personal information you cannot anonymize — and sending it offshore is a PIPL cross-border transfer. Unusually, both services do run in AWS's separate China partition. A compliance-first look at the cross-border, biometric, residency and deep-synthesis questions.
Does Amazon Polly & Transcribe work in China?
Whether Amazon Polly and Transcribe work in China is a voiceprint-biometric question, not a speed one.
A recording you send Amazon Transcribe, and a Brand Voice you have Amazon Polly synthesize, carry the speaker's voiceprint — sensitive biometric personal information under PIPL Article 28 you cannot anonymize — and on the global service it is processed offshore, a PIPL cross-border transfer AWS stores and uses to improve its models by default unless you opt out. Both services do run in AWS's separate China partition, but that is a different cloud with its own account and ICP duties, not automatic compliance. The lawful lever is to keep China audio's processing in-country — AWS's China partition, a licensed domestic service, or a self-hosted open model (e.g. Whisper) — with Article 28 consent and a PIPIA, not to make the offshore API reachable.
A risk map, not a verdict — settle the specifics with counsel. Our China team can map your exposure →
What Amazon Polly & Transcribe's own documentation says about China
| Fact | Primary source |
|---|---|
| Both services run in AWS's separate China partition. AWS's own “Amazon Web Services in China” Region Table lists Amazon Transcribe (Beijing and Ningxia) and Amazon Polly as available in the mainland partition — Beijing operated by Sinnet, Ningxia by NWCD. That is a different cloud, with its own account under a Chinese business license and its own ICP duty, not a region of your global account. | Amazon Web Services in China — Region Table (Regional Product Services), retrieved 2026-10-10 |
| Audio and text inputs are stored and used to train AWS models by default. By AWS's FAQs, Amazon Transcribe “may store and use voice inputs” and Amazon Polly “may store and use text inputs” to “improve and develop the quality” of the service “and other Amazon machine-learning/artificial-intelligence technologies,” with some content stored “in another AWS region,” unless you set an AWS Organizations opt-out policy. | Amazon Transcribe FAQs and Amazon Polly FAQs — Data Privacy, retrieved 2026-10-10 |
| A voiceprint is sensitive biometric data you can't anonymize. PIPL Article 28 classes biometric information as sensitive personal information, so processing or transferring a voice recording needs a separate specific consent and a prior personal-information protection impact assessment (PIPIA). You can redact a transcript, but stripping the voice from the audio defeats the point — you can no longer transcribe or synthesize it. | PIPL Article 28 (sensitive personal information), retrieved 2026-10-10 |
| Residency bites for a CIIO or high-volume handler. Cybersecurity Law Article 39 (formerly Article 37) requires personal information generated in China to be stored in China (the 2025 amendment, in force January 1, 2026, renumbered 37 to 39, substance unchanged; see also PIPL Article 40) — which an offshore AWS region holding your audio, transcripts and any custom voice model cannot meet. | China Cybersecurity Law, Article 39 (formerly Article 37), retrieved 2026-10-10 |
Sources verified by the 21YunBox compliance team on 2026-10-10.
Whether Amazon Polly and Amazon Transcribe “work” in mainland China is a compliance question before it is a networking one, and for a speech service it has a sharp, under-appreciated shape: a recording of a person’s voice is their voiceprint — a biometric identifier PIPL Article 28 treats as sensitive, and one you cannot anonymize. Amazon Transcribe turns audio into text, shipping the speaker’s voiceprint and spoken words offshore; its speaker-partitioning (diarization) separates people by voice, more plainly biometric still. Amazon Polly turns text into speech, and its Brand Voice builds a custom neural voice from a real person’s recordings — a deep-synthesis / generative-AI door. The prongs: a cross-border transfer (PIPL Articles 38–40); a sensitive biometric voiceprint (Article 28 — separate consent plus a prior impact assessment); content residency (Cybersecurity Law Article 39, formerly Article 37, for a CIIO or high-volume handler); and a deep-synthesis filing for any China-facing synthesized voice. The honest twist that sets this apart from Amazon Translate: both services do run inside AWS’s separate China partition.
Amazon Polly & Transcribe in China at a glance
| What decides it | In Amazon's own terms — and China's law |
|---|---|
| What you send, and that it carries a voiceprint | Amazon Transcribe takes your audio — call recordings, support calls, voice messages, meetings, dictated notes — and returns text; Amazon Polly takes text and returns synthesized speech, and its Brand Voice is trained on a real person's recordings. Audio of a person speaking is that person's voiceprint, a biometric identifier, alongside whatever the words disclose. Under PIPL Article 28 a voiceprint is sensitive personal information you cannot anonymize — stripping the audio defeats the point. |
| Where it is processed | On the global service there is no mainland-China region, so a China-facing caller reaches a non-China AWS region (Tokyo, Singapore, Hong Kong, or farther) and every recording crosses the border — a cross-border transfer of personal information under PIPL Articles 38–40 (数据出境). Unusually, both services are offered in AWS's separate China partition (Beijing/Ningxia), which is a different cloud, not a region of your global account. |
| Sensitive biometric — Article 28 | A voiceprint is biometric data, a sensitive category, so PIPL Article 28 requires a separate specific consent and a prior personal-information protection impact assessment (PIPIA) before it is processed or transferred. Transcribe's speaker diarization — telling individual speakers apart by voice — makes that biometric processing explicit; spoken content (health, finance, ID) can add further sensitive categories. |
| Retention, training & residency | By AWS's own FAQs, Amazon Transcribe “may store and use voice inputs,” and Amazon Polly “may store and use text inputs,” to “improve and develop the quality” of the service “and other Amazon machine-learning/artificial-intelligence technologies” — some content “may be stored in another AWS region” — unless you set an AWS Organizations opt-out policy. For a CIIO or high-volume handler, keeping China audio in-country is also a storage duty (Cybersecurity Law Article 39, formerly Article 37). |
| Reachability is not the axis | The API is reachable from China; that is not the question. The lawful move is to keep China-origin audio's processing in-country — AWS's own China partition, a licensed domestic speech service, or a self-hosted open model (e.g. Whisper) — obtain the Article 28 separate consent and run the PIPIA, and for any China-facing synthesized voice meet the deep-synthesis / generative-AI filing. The China-facing app that uses these features still owes an ICP filing and in-country delivery. |
What you actually send — your speakers’ voiceprints and words
Start with what a speech call actually is. Amazon Transcribe does not transcribe a reference to audio held somewhere safe; it transcribes the recording itself. You hand it the audio in full, and the exposure is the audio — and everything inside it. The jobs that most need transcription are contact-center and support calls, voice messages, interviews, meetings, and dictated medical or legal notes, and those are exactly the payloads most likely to carry health, financial, identity and other sensitive details. More fundamentally, the recording is itself a biometric identifier: a sample of the speaker’s voice. You can redact names from a finished transcript, but you cannot anonymize the audio — a voiceprint stripped of its voice is not a transcription anymore. Transcribe’s optional speaker-partitioning (diarization) sharpens the point: it processes the voice to separate and label the distinct speakers in a recording, which is biometric processing in the plain sense.
Amazon Polly is the other half — text-to-speech. Its stock voices are synthetic personas, so a routine Polly call sends text, not a person’s voice, and carries the sensitivity of whatever that text contains. The exposure changes with Brand Voice: AWS builds a custom neural voice for you, and in its own words, if “you are interested in building a Brand Voice using Amazon Polly,” AWS scopes an engagement because “every voice is unique.” That voice is trained on a real human performer’s recordings, which synthesizes — effectively clones — a specific person’s voice. For a China-facing product that is a deep-synthesis / generative-voice feature, and the performer’s recordings are themselves biometric data with voice-rights and consent attached.
Where does it go? On the global service, to a non-China AWS region — and by AWS’s own FAQs the input is “encrypted and stored at rest in the AWS region where you are using” the service, with “some portion” that “may be stored in another AWS region” for model improvement. Here is the nuance that sets Amazon apart from many speech vendors, and from Amazon Translate: AWS runs a separate China partition (the Beijing Region operated by Beijing Sinnet Technology Co., Ltd. and the Ningxia Region by Ningxia Western Cloud Data Technology Co., Ltd.), and its own China Region Table lists both Amazon Transcribe and Amazon Polly as available there. That is a real in-country option — but it is a different cloud, with its own account under a Chinese business license and its own ICP duties, not a region you toggle on your existing account (see Does AWS work in China? for the two-door partition structure).
It’s a cross-border transfer of sensitive biometric data — under PIPL
Because the audio leaves China to be processed on the global service, sending it to Amazon Transcribe or Amazon Polly is a cross-border transfer of personal information, not a routing detail. The duty sits on you as the handler, not on AWS as the processor. PIPL Articles 38–40 require that, before personal information is sent abroad, you give notice, obtain a separate consent distinct from any general agreement to use your product, and put one transfer mechanism in place — a CAC security assessment, the CAC standard contract, or certification. Above certain thresholds, genuinely necessary exports also run through China’s data-export security assessment (数据出境安全评估) before anything leaves.
Then the biometric prong, which is the heart of a speech service. A voiceprint is biometric data — a sensitive category — so PIPL Article 28 adds a separate specific consent and a prior personal-information protection impact assessment (PIPIA) before the audio is processed or transferred, and you cannot anonymize a recording you need transcribed or a voice you need synthesized. Spoken content frequently layers on more sensitive categories — a recorded support call about a medical claim carries health and financial data as well as the voiceprint. Residency is the other half: if your organization is a critical information infrastructure operator — or a high-volume personal-information handler — the Cybersecurity Law’s Article 39 (formerly Article 37 — the 2025 Cybersecurity Law amendment, in force January 1, 2026, renumbered the data-localization article from 37 to 39, its substance unchanged) requires personal information generated in China to be stored in China, with any genuinely necessary export cleared through a security assessment (see also PIPL Article 40).
For a China-facing synthesized or cloned voice — a Brand Voice offered to the public — there is a second door on top of the transfer. China’s deep-synthesis provisions and the generative-AI measures (生成式人工智能服务备案, the CAC filing) and AI-content-labeling rules (GB 45438) apply to the China-facing feature you build, and synthesizing a real person’s voice without consent is also an Article 28 and voice-rights problem. None of this turns on how fast the API responds; it turns on whose voice is processed, where, and on what lawful basis. For that reason this page publishes no China latency figure for Amazon Polly or Transcribe: speed is not the axis for a decision that turns on biometric data and cross-border transfer.
Reaching the API isn’t the question — keeping the audio in-country is
The reflex is to point the SDK at the nearest region — Tokyo, Singapore, Hong Kong — and treat the distance as the problem solved. But every global AWS region sits outside mainland China, so a nearer one changes the latency, not the law: the audio is still carried across the border, still processed offshore, and — on the default settings — still stored and used to improve AWS’s models. A nearer region is not an in-country one, and Hong Kong is a separate jurisdiction from the mainland for data-export purposes.
So the honest levers are these, and none of them is “make the offshore API reachable.” First — and this is where Amazon differs from most speech vendors — both services are offered inside AWS’s own China partition (Beijing/Ningxia, operated by Sinnet and NWCD), which keeps China-origin audio’s processing in the mainland. That is a lawful in-country option, but it is not automatic compliance: it is a separate account under a Chinese entity, it carries its own ICP duty, you still owe the Article 28 separate consent and the PIPIA for the voiceprint, and you confirm the operator’s own retention and training terms rather than assume the global FAQ. Second, where you want independence from the managed service, a licensed in-country speech provider (a domestic STT/TTS service, data kept in the mainland) or a self-hosted open model you run in-country — Amazon Transcribe and Polly themselves offer no on-prem or self-hosted build, but an open speech-to-text model such as Whisper can be self-hosted on in-mainland infrastructure. Third, minimize and obtain consent: capture the Article 28 separate consent, run the PIPIA, disable training-retention (the AWS Organizations opt-out on the global service), and for any China-facing synthesized voice meet the deep-synthesis, generative-AI filing and content-labeling duties. What none of this is: a tunnel that ships the audio offshore anyway and calls it local.
This is a risk map, not a verdict that Amazon Polly or Transcribe is “blocked” or “illegal.” Which of these obligations bite depends on your entity, whose voiceprints and what spoken content your audio carries, your role under Chinese law, and who your users are — worth settling the specifics with counsel before your speech pipeline depends on it.
The lawful path — map, localize, deliver
There is a compliant way to transcribe and synthesize speech for a China-facing product, and it has a shape.
First, map. Our China compliance team charts what audio actually flows to Amazon Transcribe and Amazon Polly today — whose voiceprints and what spoken content it carries, which jobs reach a non-China region, whether the audio is retained or used to train models, whether any feature synthesizes or clones a voice, and where you lack a lawful basis for the cross-border leg, the Article 28 consent and PIPIA, or a deep-synthesis filing. We build the technical picture; the legal conclusions are settled with your counsel.
Then localize. Where China-origin audio has to be processed on a China footing, we keep its processing in-country — on AWS’s own China partition (Beijing/Ningxia, the data held in the mainland under Sinnet or NWCD), a licensed domestic speech service, or a self-hosted open model you run in-country — obtain the Article 28 separate consent, run the PIPIA, disable training-retention, and, for a China-facing synthesized voice, meet the deep-synthesis and generative-AI filing and labeling duties. Localize means your China users’ voiceprints stop leaving the country by default — not a tunnel that ships the audio offshore anyway. You keep Amazon Polly and Transcribe for the markets and content where they already serve you.
Then deliver. The China-facing site or app that uses these speech features is itself a public service in the mainland, so it carries an ICP filing duty and needs compliant, in-country delivery. 21YunBox delivers it in-country — the 21YunBox Optimizer — set in front of the origin you already run, with no rebuild and no re-platform. The result is a China-facing product whose speech processing and delivery both run legally and compliantly for your users in China. 21YunBox never uses or suggests circumvention of any kind: we keep what must stay in-country on a consented, in-country path, deliver in-country, and never move personal information out of China by stealth. We are a compliant overlay and a partner to AWS, not a competitor and not a migration.
Related reading:
- Cross-border data transfers under PIPL
- China’s Cybersecurity Law (data localization, Article 39)
- China’s data-export security assessment measures
- China’s generative-AI measures (deep synthesis & filing)
- How to get an ICP filing for China
