Skip to main content
POST
POST /v1/audio/voice-changer takes a speech recording and re-voices it with one of the premade voices (speech to speech). Send the recording as a multipart upload or a public audio_url; the response body is the new audio file.
string
required
eleven-multilingual-sts-v2 or eleven-english-sts-v2.
file
The recording, in a multipart field named file, up to 50 MB. Send file or audio_url, not both.
string
Public http or https URL of the recording. The gateway downloads it, up to 50 MB. Send audio_url or file, not both.
string
required
The target voice: an ID or first name from the voice list, for example Sarah.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
number
From 0 to 1. How closely the output follows the target voice.
number
From 0 to 1. Exaggerates the voice’s speaking style.
boolean
Boosts similarity to the target speaker.
boolean
Removes background noise from the recording before changing the voice.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.
In a multipart upload, send one model field, one file in file, and the other fields as text form fields. The gateway accepts WAV, AIFF, MP3, AAC, MP4, M4A, MOV, FLAC, Ogg, and WebM files whose length can be read.

Response

A successful request returns HTTP 200 with the audio as the body. The Content-Type matches response_format, and the response carries x-gateway-execution-id.

Billing

The gateway bills the measured length of the input recording at $0.003 per second, rounded up to whole seconds after a 0.1-second allowance, with a minimum of 1 second. A one-minute recording costs $0.18; the 300-second maximum costs $0.90. Failed requests aren’t charged. This endpoint rejects Idempotency-Key; if a request times out, check your Console request logs before you retry.

Errors

See Errors for rate limits and provider failures.