ElevenLabs Voice Agents in Thai, Vietnamese and Indonesian: What the Model Docs Actually Allow in 2026
A clinic in Bangkok, a tour operator in Da Nang, a furniture seller in Surabaya — all three lose the same thing after 6pm. The phone rings, nobody answers, and the caller books with whoever picks up. Voice agents are sold as the fix, and the vendor pages all lead with a language count: 74 languages, 70+ languages, 99 for transcription.
The count is not the number that decides whether this works on your line.
ElevenLabs does not run every language on every model, and its models have very different latency characteristics. A language that only exists on the expressive, higher-latency model is a different operational proposition from one that runs on the model built for real-time calls — even though both appear in the same "74 languages" headline. For Southeast Asia, that distinction lands unevenly, and it lands hardest on Thai.
This is a single-product operational read, not a shootout. Everything below is checked against ElevenLabs' own documentation, with the pages named so you can re-check them — their docs move, and their own affiliate material warns that its PDFs go stale.
The short answer, per language
| Language | On the real-time model (Flash v2.5)? | Documented latency to plan against | Practical read |
|---|---|---|---|
| Indonesian | Yes | ~75 ms | Deployable |
| Malay | Yes | ~75 ms | Deployable |
| Filipino | Yes | ~75 ms | Deployable |
| Vietnamese | Yes | ~75 ms | Deployable |
| Thai | No | None published | Deployable only via Eleven v3 Conversational, unquantified |
That table is the whole post. The rest is why it is true and what it costs you.
Which model carries which language
From ElevenLabs' language-support documentation, the per-model lists differ substantially:
- Eleven v3 — 74 languages. Includes Thai, Vietnamese, Indonesian, Malay, Filipino, Javanese.
- Flash v2.5 — 32 languages. Includes Vietnamese, Indonesian, Malay, Filipino. No Thai.
- Multilingual v2 — 29 languages. Includes Indonesian, Malay, Filipino. No Thai, no Vietnamese.
- Flash v2 — English only.
So Thai is on exactly one model family. Vietnamese is on two. Indonesian, Malay and Filipino are on three. Every one of those languages is inside the "74 languages" headline, and the headline tells you none of this.
Why the model matters more than the language count
ElevenLabs documents Flash v2.5 at "Ultra-low latency (~75ms†)" — the dagger is theirs, excluding application and network latency — and describes it as "perfect for real-time voice agents and chatbots". Multilingual v2 is documented as having "higher latency & cost per character than Flash models". Eleven v3 publishes no latency figure in the model table at all.
ElevenLabs also publishes its own targets for what conversational latency needs to be. On their engineering blog: below 700 ms end-to-end is "a good target for production-ready agents", agents above 1,200 ms show "high abandonment rates", and under 500 ms is where conversation "flows smoothly". They recommend watching P95, not the average — which is the correct instinct on a mobile network in Jakarta or Hanoi.
That is the gap. Voice synthesis is one leg of a pipeline that also includes speech recognition, the language model, and the network. If your TTS leg has a documented ~75 ms budget, you have roughly 600 ms to spend on everything else and still land inside ElevenLabs' own "production-ready" threshold. If your TTS leg has no published figure, you cannot do that arithmetic before you build.
Thai is not blocked — it is unquantified
The useful nuance, and the thing that makes this less bad than it first reads: ElevenLabs ships Eleven v3 Conversational, described in the agents documentation as "an ultra-low-latency version of Eleven v3, optimized for live, back-and-forth dialogue" that "maintains conversational context across turns and adapts delivery to match the tone and intent of each exchange". It covers the full 70+ language set, so Thai is included, and Expressive Mode is on by default when an agent selects it.
So a Thai voice agent is buildable today. What you do not get is a number. "Ultra-low-latency" is a description; ~75 ms is a specification. For a Vietnamese or Indonesian deployment you can budget the pipeline on paper before committing. For Thai you have to build a prototype and measure it on your own network, against your own telephony, before you know whether it clears 700 ms.
That is a real difference in project risk, and it falls on the one SEA market where the language has no fallback model.
The agents documentation also states limits on v3 Conversational that apply to every language using it: it "does not preserve Professional Voice Clone characteristics well", expressive tag effects last "approximately 4-5 words before returning to normal delivery", and "expressiveness may vary across voices and languages". That last clause is the vendor conceding the thing this post is about — the language list is not a quality guarantee.
What we could not verify, and are not going to assert
No published per-language quality benchmark exists for Thai, Vietnamese or Indonesian output. ElevenLabs does not provide one, and we have not run a structured listening test with native speakers. So this post makes no claim about how natural Thai output sounds, whether Vietnamese tones survive synthesis, or how Indonesian handles code-switched English product names — all three of which matter enormously in practice and none of which are answerable from documentation.
What we can say is that the documentation puts Thai on a different footing from its neighbours, and that a buyer should price that in. A listening test with native speakers on your own scripts is the missing evidence, and it is worth running before you commit a support line to any of these.
Data residency: Singapore exists, on one tier
Worth correcting a common assumption: ElevenLabs does offer a Southeast Asian data residency region. The documentation lists three endpoints — EU (api.eu.residency.elevenlabs.io), India (api.in.residency.elevenlabs.io) and Singapore (api.sg.residency.elevenlabs.io).
Two qualifications matter for a SEA buyer:
- It is Enterprise-only. The docs state plainly that "data residency is an exclusive feature available to ElevenLabs' Enterprise customers." The self-serve tiers below do not have it.
- Residency governs storage, not all processing. The documentation says customer data is "hosted/stored" in the selected location, but that "processing may nevertheless occur outside of the selected location, including by ElevenLabs' international affiliates and subprocessors, for support purposes, and for content moderation purposes." Only the EU region documents a path to restrict processing as well, via Zero Retention Mode and the API.
If you are handling health, financial or government-adjacent voice data under a Thai, Indonesian or Vietnamese regime that treats processing location as material, read that second point carefully rather than treating "Singapore region" as settled.
What it costs at the tier a SEA SME actually buys
Published pricing, verified 2026-08-06, all on a shared credit pool across products:
| Plan | Monthly | Credits |
|---|---|---|
| Free | $0 | 10k |
| Starter | $6 | 30k |
| Creator | $22 | 121k |
| Pro | $99 | 600k |
| Scale | $299 | 1.8M |
| Business | $990 | 6M |
The free tier is the honest place to settle the question this post cannot: start on ElevenLabs at $0, run your own Thai or Vietnamese script through it, and listen. Credits are consumed differently per feature — text-to-speech runs at roughly 1 credit per character, speech-to-text at around 330 credits per minute. A clinic answering a few dozen after-hours calls a month lives in Creator or Pro; a support line handling real volume is a Scale conversation. Nothing in the self-serve range includes data residency.
Who this fits, and who should wait
Reasonable to deploy now: an Indonesian, Malay, Filipino or Vietnamese phone or chat agent, where Flash v2.5 gives you a documented latency budget and the language has a fallback model if v3 behaves unexpectedly.
Prototype before you commit: anything Thai-facing. The capability exists via Eleven v3 Conversational, but you are measuring it yourself rather than planning against a published figure, and there is no second model to fall back to if the result disappoints.
Do not assume it is solved: regulated Thai, Indonesian or Vietnamese workloads where processing location is a compliance question, and any deployment whose quality bar depends on how the language actually sounds to a native ear. Neither is answerable from a language count.
The honest summary is that ElevenLabs' SEA language coverage is genuinely broad and unevenly documented, and the unevenness maps onto exactly the markets a directory like this one covers. Check elevenlabs.io/docs/overview/models and the language-support page before you build — they changed during the writing of this post, and they will change again.
Facts in this article were verified against ElevenLabs' public documentation on 6 August 2026: the models overview, the language-support page, the agents Expressive Mode documentation, the data residency documentation, the conversational-latency engineering blog, and the pricing page. Vendor marketing material was not used as a source.