Skip to main content
POST
Eleven v3
eleven-v3 runs ElevenLabs Eleven v3 text to speech on the OpenAI-compatible speech endpoint. Send up to 5,000 characters with one of 21 premade voices and receive the audio file in the response. Eleven v3 follows audio tags in square brackets, such as [excited], [whispers], or [laughs], so you can direct the delivery line by line. Confirm live rates on the SnapGen models page. Failed requests are not charged. For a conversation between several voices in one file, use eleven-v3-dialogue. For longer text per request, use eleven-v4 (10,000 characters) or eleven-flash-v2.5 (40,000 characters).
string
required
Set to eleven-v3.
string
required
The text to speak, up to 5,000 characters, including audio tags. Newlines are allowed; blank text and other control characters are rejected. Each character is billed.
string
required
A voice ID or first name from the voice list, for example JBFqnCBsd6RMkjVDRZzb or George. Names are case-insensitive.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.
number
Speaking speed, from 0.7 to 1.2.
string
An ISO 639 language code of two or three lowercase letters, such as en or de. The provider rejects a code the model doesn’t support.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
number
From 0 to 1. How closely the output follows the original voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the original speaker.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
Up to 5,000 characters that come before input, for continuity across requests. Not billed.
string
Up to 5,000 characters that follow input, for continuity across requests. Not billed.
string
auto, on, or off: whether numbers, dates, and similar text are spelled out before speaking.

Request

The response body is the audio file, with a Content-Type that matches response_format (audio/mpeg for MP3).

Direct the delivery with audio tags

Put a tag before the words it should shape. Tags are part of input, so they count toward the 5,000-character limit and are billed like any other text.

Price examples

Synchronous audio endpoints reject Idempotency-Key. If a request times out, check your Console request logs before you retry. See Text to speech for the response headers and errors.