ElevenLabs now offers two v4 paths. One prioritizes expressive multi-speaker dialogue. The other trades some flexibility for a faster streaming response.
ElevenLabs announced Eleven v4 on September 28 and updated the post on October 1. The company says the models are available in ElevenAgents, ElevenCreative, and the API. The public model identifiers are eleven_v4 and eleven_v4_turbo.
Start with the endpoint, not just the model name.
The standard v4 model is documented for the Text to Dialogue API. It supports multi-speaker output and up to 10,000 characters per request. The model documentation describes v4 Turbo as the lower-latency option, while the WebSocket guide explains the streaming Text to Dialogue path.
This distinction matters during migration. Replacing a model string without checking the endpoint can leave a workflow using the wrong request shape or streaming behavior. Pin both the model and the API route in your test plan.
The latency number is a vendor target.
ElevenLabs says v4 Turbo can reach roughly 100 to 150 milliseconds of latency. Treat that as a company-reported figure, not a guarantee for every region, voice, network, and script. Measure time to first audio and time to complete audio from the same client environment your product will use.
For an interactive agent, a fast first chunk can matter more than total generation time. For a produced video or audiobook, consistency, pronunciation, and emotional control may matter more than the first packet. Define the use case before choosing the model.
Languages and request size need real scripts.
ElevenLabs lists support for more than 90 languages. A language count doesn’t guarantee equal quality across accents, names, code-switching, or long-form narration. Test the languages and voices your product supports, including difficult proper nouns and punctuation.
The 10,000-character v4 limit is a per-request ceiling for the documented Text to Dialogue path. Longer material should be split at natural boundaries and checked for voice continuity. Keep the original text segments so you can regenerate one section without paying to recreate an entire production.
Credits are the useful budgeting unit.
ElevenLabs offers free and paid entry points, but usage is metered through credits. Before setting a monthly estimate, run a representative sample for each content type and record characters, output duration, retries, and discarded takes. A polished minute can require more than one generation.
Our earlier ElevenLabs Studio 4.0 review covers a timeline-based production tool. The v4 API decision is different: it concerns the model, endpoint, and operating cost inside your own workflow.
A migration test that catches the expensive failures
- Pin the model identifier and endpoint in a non-production environment.
- Test representative voices, languages, and multi-speaker turns.
- Measure first-audio latency, total latency, and failed requests.
- Track credits for accepted output, not generated output alone.
- Retain the previous model path until the new output passes review.
The best v4 choice is not universal. Use standard v4 when expressive dialogue and production control are the priority. Evaluate v4 Turbo when responsiveness is central, then verify that the speed gain does not reduce the qualities your users notice.