Skip to main content
POST
Eleven Flash v2.5
eleven-flash-v2.5 runs ElevenLabs Flash v2.5 text to speech. It costs half as much per character as the other speech models and accepts the most text: up to 40,000 characters per request. Confirm live rates on the SnapGen models page. Failed requests are not charged.
string
required
Set to eleven-flash-v2.5.
string
required
The text to speak, up to 40,000 characters for this model. Newlines are allowed; blank text and other control characters are rejected. Each character is billed.
string
required
A voice ID or first name from the voice list, for example JBFqnCBsd6RMkjVDRZzb or George. Names are case-insensitive.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.
number
Speaking speed, from 0.7 to 1.2.
string
An ISO 639 language code of two or three lowercase letters, such as en or de. The provider rejects a code the model doesn’t support.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
number
From 0 to 1. How closely the output follows the original voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the original speaker.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
Up to 5,000 characters that come before input, for continuity across requests. Not billed.
string
Up to 5,000 characters that follow input, for continuity across requests. Not billed.
string
auto, on, or off: whether numbers, dates, and similar text are spelled out before speaking.

Request

The response body is the audio file, with a Content-Type that matches response_format.

Price examples

Synchronous audio endpoints reject Idempotency-Key. If a request times out, check your Console request logs before you retry. See Text to speech for every field, the response, and errors.