Skip to main content
POST
Eleven English STS v2
eleven-english-sts-v2 is the English speech-to-speech voice changer: it re-voices an English recording with one of the premade voices and returns the new audio in the response. Use eleven-multilingual-sts-v2 for other languages. Confirm live rates on the SnapGen models page. Failed requests are not charged.
string
required
Set to eleven-english-sts-v2.
file
The audio file, sent as multipart/form-data, up to 50 MB. Send file or audio_url, not both.
string
Public http or https URL of the audio. The gateway downloads it (up to 50 MB). Send audio_url or file, not both.
string
required
A voice ID or first name from the voice list, for example JBFqnCBsd6RMkjVDRZzb or George. Names are case-insensitive.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
number
From 0 to 1. How closely the output follows the original voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the original speaker.
boolean
Removes background noise from the input before changing the voice.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.

Request

Upload a local file as multipart/form-data:
Or send JSON with a public audio_url:
The response body is the audio file, with a Content-Type that matches response_format.

Price examples

Synchronous audio endpoints reject Idempotency-Key. If a request times out, check your Console request logs before you retry. See Voice changer for every field, the response, and errors.