ElevenLabs Voice Agents in Thai, Vietnamese and Indonesian: What the Model Docs Actually Allow in 2026
A clinic in Bangkok, a tour operator in Da Nang, a furniture seller in Surabaya — all three lose the same thing after 6pm. The phone rings, nobody answers, and the caller books with whoever picks up. Voice agents get sold as the fix, and the vendor pages lead with a language count: 74 languages, 70+ languages, 99 for transcription.
The count is not the number that decides whether this works on your line.
ElevenLabs does not run every language on every model, and its models have very different latency characteristics. A language stuck on the expressive, higher-latency model is a different bet than one on the model built for real-time calls — even though both sit inside the same "74 languages" headline. For Southeast Asia, that lands hardest on Thai.
This is a single-product read, not a shootout, checked against ElevenLabs' own docs with the pages named so you can re-check them.
The short answer, per language
| Language | On the real-time model (Flash v2.5)? | Documented latency to plan against | Practical read |
|---|---|---|---|
| Indonesian | Yes | ~75 ms | Deployable |
| Malay | Yes | ~75 ms | Deployable |
| Filipino | Yes | ~75 ms | Deployable |
| Vietnamese | Yes | ~75 ms | Deployable |
| Thai | No | None published | Deployable only via Eleven v3 Conversational, unquantified |
That table is the whole post. The rest is why, and what it costs.
Which model carries which language
From ElevenLabs' docs, the per-model lists differ a lot:
- Eleven v3 — 74 languages. Includes Thai, Vietnamese, Indonesian, Malay, Filipino, Javanese.
- Flash v2.5 — 32 languages. Includes Vietnamese, Indonesian, Malay, Filipino. No Thai.
- Multilingual v2 — 29 languages. Includes Indonesian, Malay, Filipino. No Thai, no Vietnamese.
- Flash v2 — English only.
Thai sits on exactly one model family. Vietnamese is on two. Indonesian, Malay and Filipino are on three. All of it hides inside the "74 languages" headline.
Why the model matters more than the language count
ElevenLabs documents Flash v2.5 at "Ultra-low latency (~75ms†)" — the dagger is theirs, excluding application and network latency — and calls it "perfect for real-time voice agents and chatbots." Multilingual v2 gets "higher latency & cost per character than Flash models." Eleven v3 publishes no latency figure in the model table at all.
ElevenLabs publishes its own latency targets too: below 700 ms end-to-end is "a good target for production-ready agents," above 1,200 ms shows "high abandonment rates," and under 500 ms is where conversation flows smoothly. They recommend watching P95, not the average — the right instinct on a mobile network in Jakarta or Hanoi.
That is the gap. Voice synthesis is one leg of a pipeline that also runs speech recognition, the language model, and the network. A ~75 ms TTS budget leaves roughly 600 ms for everything else and still lands inside the "production-ready" threshold. No published TTS figure, no arithmetic before you build.
Thai is not blocked — it is unquantified
Here's the nuance: ElevenLabs ships Eleven v3 Conversational, an ultra-low-latency version of v3 built for live, back-and-forth dialogue. It maintains context across turns and covers the full 70+ language set — Thai included, Expressive Mode on by default.
So a Thai voice agent is buildable today. What you don't get is a number. "Ultra-low-latency" is a description; ~75 ms is a specification. Plan a Vietnamese or Indonesian deployment on paper. For Thai, build a prototype and measure it on your own network and telephony before you know whether it clears 700 ms — real project risk, on the one SEA market with no fallback model.
A few limits apply to every language on v3 Conversational: it doesn't preserve Professional Voice Clone characteristics well, and expressive tags last "approximately 4-5 words" before resetting. Expressiveness itself "may vary across voices and languages" — the vendor conceding the point of this post.
What we could not verify
No published per-language quality benchmark exists for Thai, Vietnamese or Indonesian output, and we haven't run a listening test with native speakers ourselves. We won't claim how natural the Thai sounds, or whether Vietnamese tones survive synthesis. Only testing your own scripts answers that.
Data residency: Singapore exists, on one tier
Worth correcting a common assumption: ElevenLabs does offer a Southeast Asian residency region — three endpoints total, EU, India, and Singapore (api.sg.residency.elevenlabs.io).
Two catches. It's Enterprise-only, so self-serve tiers don't get it. And residency governs storage, not all processing. ElevenLabs' docs admit processing "may nevertheless occur outside" the selected location for support and content-moderation purposes — only the EU region offers a way to restrict that too, via Zero Retention Mode. Handling health, financial or government-adjacent voice data under a regime that treats processing location as material? Read that second point twice.
What it costs at the tier a SEA SME actually buys
Published pricing, verified 2026-08-06, all on a shared credit pool:
| Plan | Monthly | Credits |
|---|---|---|
| Free | $0 | 10k |
| Starter | $6 | 30k |
| Creator | $22 | 121k |
| Pro | $99 | 600k |
| Scale | $299 | 1.8M |
| Business | $990 | 6M |
That Creator tier, $22 a month, is roughly THB 790 or IDR 357,000 — the range a Bangkok clinic or a Surabaya furniture seller actually budgets against, not the enterprise price further down. Credits burn per feature: about 1 credit per character for text-to-speech, 330 credits per minute for speech-to-text. A clinic fielding a few dozen after-hours calls a month lives in Creator or Pro; real volume means Scale. Data residency isn't in the self-serve range at all.
The free tier is the honest place to settle the question the docs can't: run your own Thai or Vietnamese script through it at $0, and listen for yourself.
Who this fits, and who should wait
Deploy now: Indonesian, Malay, Filipino or Vietnamese agents, where Flash v2.5 gives a documented latency budget and a fallback if v3 misbehaves.
Prototype first: anything Thai-facing. The capability exists via Eleven v3 Conversational, but you're measuring it yourself against no published figure and no fallback model.
Don't assume it's solved: regulated workloads where processing location is a compliance question, or anything whose quality bar depends on how the language sounds to a native ear. Neither is answerable from a language count.
ElevenLabs' SEA language coverage is genuinely broad and unevenly documented, and the unevenness maps onto exactly the markets a directory like this one covers. Check elevenlabs.io/docs/overview/models before you build — the docs changed while this post was being written, and they'll change again.
Facts in this article were verified against ElevenLabs' public documentation on 6 August 2026: the models overview, the language-support page, the agents Expressive Mode documentation, the data residency documentation, the conversational-latency engineering blog, and the pricing page. Vendor marketing material was not used as a source.