Skip to main content
POST
POST /v1/audio/dialogue turns an ordered list of lines into one audio file of a conversation. Each line names its own voice, so you can script an interview, a podcast intro, or a scene without stitching clips together. The response body is the audio file.
string
required
eleven-v4-dialogue or eleven-v3-dialogue.
object[]
required
1 to 50 lines in speaking order. The text of all lines totals at most 2,000 characters, and one dialogue uses at most 10 different voices.
string
default:"mp3"
mp3, opus, pcm, or wav. See Output formats.
string
An ISO 639 language code of two or three lowercase letters, such as en. The provider rejects a code the model doesn’t support.
number
From 0 to 1. Lower values give a broader emotional range; higher values sound more consistent.
integer
From 0 to 4294967295. Makes repeated requests more repeatable, without guaranteeing identical audio.
Fields that aren’t listed here return invalid_request.

Direct the delivery

Lines can carry audio tags in square brackets, such as [excited] or [laughs], to shape how they’re spoken. Tags are part of text, so they count toward the 2,000-character total and are billed.

Response

A successful request returns HTTP 200 with the audio as the body. The Content-Type matches response_format, for example audio/mpeg, and the response carries x-gateway-execution-id.

Billing

The request is billed per character across every line’s text, counted in Unicode code points. A full 2,000-character dialogue costs $0.24. The price is known before the provider is called, and failed requests aren’t charged. This endpoint rejects Idempotency-Key; if a request times out, check your Console request logs before you retry.

Errors

See Errors for rate limits and provider failures.