Skip to main content
POST
POST /v1/audio/speech turns text into speech and returns the audio file as the response body. It accepts the OpenAI speech request (model, input, voice, response_format, and speed) and adds ElevenLabs voice settings. Choose a model from Audio Series > ElevenLabs Speech. For a conversation between several voices in one file, use Text to dialogue.
string
required
One of the speech models above.
string
required
The text to speak, up to the model’s character limit. Newlines are allowed; blank text and other control characters are rejected. Each character is billed.
string
required
A voice ID or first name from the voice list, for example JBFqnCBsd6RMkjVDRZzb or George. Names are case-insensitive. OpenAI voice names such as alloy aren’t available.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.
number
Speaking speed, from 0.7 to 1.2.
string
An ISO 639 language code of two or three lowercase letters, such as en or de. The provider rejects a code the model doesn’t support.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent from one generation to the next.
number
From 0 to 1. How closely the output follows the original voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the original speaker.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
Up to 5,000 characters that come before input. Send it when you split long text across requests so the speech flows across the join. Not billed.
string
Up to 5,000 characters that follow input, for the same purpose. Not billed.
string
auto, on, or off: whether numbers, dates, and similar text are spelled out before they’re spoken.
The body is validated strictly. A field that isn’t listed here, including OpenAI’s instructions, returns invalid_request.

Response

A successful request returns HTTP 200 with the audio as the body. The gateway streams it to you as the provider sends it; save the body to a file.

Output formats

Voices

Every speech, dialogue, and voice changer model uses the same 21 ElevenLabs premade voices. Send the voice ID or the first name as voice. Other ElevenLabs voices, such as cloned or library voices, aren’t available.

Use the OpenAI SDK

The official OpenAI SDKs work with the SnapGen base URL. Pass a voice from the list above. In Python, send ElevenLabs fields such as stability through extra_body.

Billing

The request is billed per character of input, counted in Unicode code points: Héllo 👋 is 7 characters. previous_text and next_text are free. The gateway prices the request before it calls the provider, so the charge is known up front, and a failed request isn’t charged. For example, 1,000 characters cost $0.12 on eleven-v4 and $0.06 on eleven-flash-v2.5.
This endpoint is synchronous and rejects Idempotency-Key. If a request times out, the audio may already have been generated: check your Console request logs before you send it again.

Errors

Errors use the standard error envelope.