Does Speechmatics Work in China? PIPL Cross-Border, Voiceprint Biometrics & Data Residency
The audio you send Speechmatics to transcribe carries the speaker's voiceprint — Article 28 sensitive biometric personal information you can't anonymize — and its SaaS has no mainland-China region, so China audio is processed offshore (EU, US, or Australia): a PIPL cross-border transfer. A compliance-first look at the lawful in-country path, Article 28 consent, and CSL Article 39 residency.
Does Speechmatics work in China?
Speechmatics is speech-to-text, and the audio you send it is the speaker's voiceprint — sensitive biometric personal information under PIPL Article 28 that you cannot anonymize — processed on a SaaS with no mainland-China region, so a China speaker's audio and words are shipped offshore to the EU, US, or Australia.
That makes each transcription call a PIPL cross-border transfer of sensitive biometric personal information (Articles 38–40, plus the Article 28 separate-consent and PIPIA bar), with an in-country storage duty under Cybersecurity Law Article 39 (formerly Article 37) for a CIIO or high-volume handler. The honest lever is not to make the offshore API reachable — it is to keep China audio's processing in-country: Speechmatics ships an on-premises container you can run on infrastructure inside the mainland (or use a licensed domestic speech service), paired with the Article 28 consent.
This is a risk map, not a verdict — settle the specifics with counsel. Our China team can map your exposure →
What Speechmatics's own documentation says about China
| Fact | Primary source |
|---|---|
| Speechmatics ships an on-premises, self-hosted container you run in your own environment. Its deployment docs state, "Deploy Speechmatics services in your own environment using containers," and that "This option provides maximum control over your deployment and data," with CPU, GPU, and Kubernetes builds that run "on your own hardware." Run that container on infrastructure inside the mainland and the audio — and the speaker's voiceprint — never leaves China. | Speechmatics Documentation — Deployments (docs.speechmatics.com), retrieved 2026-10-10 |
| The Speechmatics SaaS has no mainland-China region. Its regions documentation lists only EU1 (Europe), US1 (United States), and AU1 (Australia), states that "A region is the location where your audio is processed," and that "the region you use is determined by the endpoint you call." A China speaker's audio on the SaaS is therefore processed offshore. | Speechmatics Documentation — Regions (docs.speechmatics.com), retrieved 2026-10-10 |
| A voice recording is sensitive biometric personal information you cannot anonymize. Under PIPL Article 28, a voiceprint is biometric data in the sensitive category, requiring a separate specific consent and a prior personal-information protection impact assessment (PIPIA) before it is processed or transferred — and stripping identifiers from a transcript does not de-identify the audio itself. | Personal Information Protection Law of the PRC, Article 28 (cac.gov.cn), retrieved 2026-10-10 |
| Sending China-origin audio to an offshore endpoint is a PIPL cross-border transfer. PIPL Articles 38–40 require notice, a separate consent, and a transfer mechanism (a CAC security assessment, the standard contract, or certification); for a critical information infrastructure operator or high-volume handler, Cybersecurity Law Article 39 (formerly Article 37) adds an in-country storage duty. | PIPL Articles 38–40; PRC Cybersecurity Law Article 39 (formerly Article 37) (cac.gov.cn), retrieved 2026-10-10 |
Sources verified by the 21YunBox compliance team on 2026-10-10.
For a mainland-China audience, the question to settle about Speechmatics — the speech-to-text engine that turns recorded or live audio into a transcript — is not whether its API can be reached from inside the country. It is what happens to the audio you send it. A recording of someone speaking is not just words: it is that person’s voiceprint, a biometric identifier that PIPL Article 28 treats as sensitive personal information, and one you cannot anonymize — redact the transcript all you like, the audio itself still identifies the speaker. On the Speechmatics SaaS, that audio is processed in the EU, the US, or Australia; there is no mainland-China region, so every call ships a China speaker’s voiceprint and whatever they said offshore. That raises four questions at once: a cross-border transfer of personal information, a sensitive-biometric bar, content residency, and — for the app that captures the audio — an ICP filing duty.
Speechmatics in China at a glance
| What decides it | In Speechmatics' own terms — and China's law |
|---|---|
| What you send | Audio — a call, a meeting, a support conversation, a voice message, or a dictated note — plus, inseparable from it, the speaker's voiceprint (a biometric identifier) and the spoken content, which is often sensitive in its own right (health, finance, identity, legal matters). Speechmatics is speech-to-text: audio in, transcript out. It recognizes speech; it does not synthesize or clone voices. |
| Where it goes | The hosted SaaS processes audio in EU1 (Europe), US1 (United States), or AU1 (Australia) — there is no mainland-China region. Sending a China user's audio to any of them is a cross-border transfer of personal information under PIPL Articles 38–40 (数据出境): notice, a separate consent, and a transfer mechanism. |
| Why voice is different | A voiceprint is sensitive biometric personal information — PIPL Article 28 — requiring a separate specific consent and a prior personal-information protection impact assessment (PIPIA). You cannot anonymize it: redacting a transcript does not de-identify the recording, which still identifies the speaker. |
| Retention, training & residency | On the SaaS, audio and transcripts are retained under Speechmatics' data-handling terms, with an opt-in that lets anonymized audio improve its models. For a CII operator or high-volume handler, in-country storage applies — Cybersecurity Law Article 39 (formerly Article 37). Speechmatics also ships an on-prem/self-hosted container — the in-country lever. |
| What actually decides it | Not whether the API is reachable. The lawful path is to keep processing in-country — run Speechmatics' own on-prem container inside the mainland, or a licensed domestic speech service, with Article 28 consent — and to deliver the China-facing app that captures the audio on ICP-filed infrastructure. |
What you actually send — your speakers’ voiceprints and words
Speechmatics is a speech-to-text engine: audio goes in — a call recording, a support conversation, a voice message, a dictated note, a meeting — and a transcript comes out. On the hosted SaaS, that audio is processed in one of three regions. Speechmatics’ own documentation lists EU1 (Europe), US1 (United States), and AU1 (Australia), states plainly that “A region is the location where your audio is processed,” and that “the region you use is determined by the endpoint you call.” None of them is in the mainland, and the word China does not appear on that page. So when a person in China speaks into a Speechmatics-powered feature, the SaaS ships their audio out of the country on every call.
What leaves is not only what they said. The recording is the speaker’s voiceprint — a biometric identifier unique to them — and the spoken content is frequently sensitive in its own right (health, finance, identity, legal matters). Speechmatics recognizes speech; it does not synthesize or clone voices, so the deep-synthesis and generative-AI filing door is not the question here. The exposure is simpler and sharper: the live cross-border transfer of the voiceprint and the words, on every request, plus whatever is retained afterward and any opt-in that lets anonymized audio improve the models.
It’s a cross-border transfer of sensitive biometric data — under PIPL
Routing a Chinese user’s audio to a Speechmatics region outside the mainland is a cross-border transfer of personal information under China’s Personal Information Protection Law. The duty sits on the handler — you, not the vendor. PIPL Articles 38–40 require notice, a separate consent distinct from the user’s agreement to use the feature, and one transfer mechanism: a CAC security assessment, the CAC standard contract, or certification.
Voice raises the bar higher than ordinary text. A voiceprint is biometric data, and PIPL Article 28 places biometric data in the sensitive category — which demands a separate, specific consent and a prior personal-information protection impact assessment (PIPIA) before the data is processed or transferred, on top of a showing of necessity. The catch that makes audio different from a document is that you cannot anonymize it: you can redact a name from a transcript, but the recording itself still identifies the speaker, so the usual de-identification escape hatch is closed.
Underneath that sits residency. A critical information infrastructure operator or a large-volume handler must store personal information collected in China inside the mainland — PIPL Article 40 together with the Cybersecurity Law. (The 2025 Cybersecurity Law amendment, in force January 1, 2026, renumbered the data-localization article from 37 to 39; the substance is unchanged, so this is Cybersecurity Law Article 39, formerly Article 37.) An offshore region structurally cannot meet that storage duty, and where volumes or sensitivity cross the thresholds, the export itself needs a CAC data-export security assessment before it may proceed.
Reaching the API isn’t the question — keeping the audio in-country is
That the API can be reached from China does not make the transfer lawful — reachability is not the axis. The axis is whether the voiceprint and the spoken content of people in China leave the country without a lawful basis. So the lever is not to make an offshore endpoint faster or easier to reach; it is to stop exporting the audio in the first place.
For Speechmatics, that lever is unusually clean, because the vendor itself offers a self-hosted build. Speechmatics’ deployment documentation describes an on-prem option — “Deploy Speechmatics services in your own environment using containers” — in CPU, GPU, and Kubernetes forms that run “on your own hardware,” and notes that this “provides maximum control over your deployment and data.” Run that container on infrastructure inside the mainland and the recognition happens in-country: the audio, and the speaker’s voiceprint, never cross the border. Where self-hosting the vendor’s own engine is not practical, a licensed domestic speech-to-text service that keeps the data in the mainland is the alternative. Either way you still obtain the Article 28 separate consent, run the PIPIA, minimize what you collect, and disable any model-training retention — localization keeps the processing in-country; it does not remove the consent and impact-assessment duties.
This is a risk map, not a verdict: which prongs bite, and how hard, depends on your volumes, your sector, whether you are treated as a CII operator, and exactly what your feature does with the audio — settle the specifics with counsel.
The lawful path — map, localize, deliver
There is a lawful way to run Speechmatics for users in China, and it has a clear shape: the audio is handled from inside the mainland on a footing that keeps the voiceprint resident, and the China-facing app that captures it is itself filed and delivered in-country. The part 21YunBox owns is that footing, and it is more than advice. Our China team does three things. We map the exposure — what audio flows to the speech service, whose voiceprints and what spoken content it carries, where it is processed and whether it is retained or used to train models, and where you lack a lawful basis (the Article 28 separate consent and PIPIA; a transfer mechanism). We localize the transcription onto an in-country path — Speechmatics’ own on-prem container run on infrastructure inside the mainland, or a licensed domestic speech-to-text service — so China audio’s processing stays in the country instead of being shipped offshore, paired with the Article 28 consent. And we deliver the China-facing site or app that captures the audio in-country on ICP-filed infrastructure — the 21YunBox Optimizer — in front of the stack you already run, with no rebuild and no re-platform.
The result is a speech feature that runs legally and compliantly for your users in China. 21YunBox never uses or suggests circumvention of any kind: this is a lawful, data-resident in-country deployment, built on the vendor’s own self-hosted engine or a licensed domestic equivalent — never a path that ships the audio offshore anyway.
Related reading:
- Cross-border data transfers under PIPL
- China’s Cybersecurity Law — Article 39 (formerly Article 37) and data localization
- China’s data-export security assessment measures
- How to get an ICP filing for China
