Skip to main content
POST
Eleven v4
eleven-v4 runs ElevenLabs Eleven v4 text to speech on the OpenAI-compatible speech endpoint. Send up to 10,000 characters with one of 21 premade voices and receive the audio file in the response. Use it with the official OpenAI SDKs by changing only the base URL, the model, and the voice. Confirm live rates on the SnapGen models page. Failed requests are not charged.

Choose a speech model

All four speech models share the endpoint, the fields, and the voices. They differ in the upstream model, the character limit, and the price. For several speakers in one file, use eleven-v4-dialogue.
string
required
Set to eleven-v4.
string
required
The text to speak, up to 10,000 characters. Newlines are allowed; blank text and other control characters are rejected. Each character is billed.
string
required
A voice ID or first name from the voice list, for example JBFqnCBsd6RMkjVDRZzb or George. Names are case-insensitive.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.
number
Speaking speed, from 0.7 to 1.2.
string
An ISO 639 language code of two or three lowercase letters, such as en or de. The provider rejects a code the model doesn’t support.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
number
From 0 to 1. How closely the output follows the original voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the original speaker.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
Up to 5,000 characters that come before input, for continuity across requests. Not billed.
string
Up to 5,000 characters that follow input, for continuity across requests. Not billed.
string
auto, on, or off: whether numbers, dates, and similar text are spelled out before speaking.

Request

The response body is the audio file, with a Content-Type that matches response_format (audio/mpeg for MP3).

Longer text

A request takes up to 10,000 characters. For a longer script, split it at sentence or paragraph boundaries and send one request per part. Pass the neighboring text as previous_text and next_text so each part flows into the next; that context isn’t billed. Keep voice, seed, and the voice settings the same for every part.

Price examples

Synchronous audio endpoints reject Idempotency-Key. If a request times out, check your Console request logs before you retry. See Text to speech for the response headers and errors.