Skip to main content
POST
Eleven Multilingual STS v2
eleven-multilingual-sts-v2 is the multilingual speech-to-speech voice changer: it re-voices a recording with one of the premade voices and returns the new audio in the response. Confirm live rates on the SnapGen models page. Failed requests are not charged.
string
required
Set to eleven-multilingual-sts-v2.
file
The audio file, sent as multipart/form-data, up to 50 MB. Send file or audio_url, not both.
string
Public http or https URL of the audio. The gateway downloads it (up to 50 MB). Send audio_url or file, not both.
string
required
A voice ID or first name from the voice list, for example JBFqnCBsd6RMkjVDRZzb or George. Names are case-insensitive.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
number
From 0 to 1. How closely the output follows the original voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the original speaker.
boolean
Removes background noise from the input before changing the voice.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.

Request

Upload a local file as multipart/form-data:
Or send JSON with a public audio_url:
The response body is the audio file, with a Content-Type that matches response_format.

Price examples

Synchronous audio endpoints reject Idempotency-Key. If a request times out, check your Console request logs before you retry. See Voice changer for every field, the response, and errors.