Skip to main content
POST
Scribe v2
scribe-v2 transcribes speech with ElevenLabs Scribe v2 on the OpenAI-compatible transcription endpoint. Upload a file or pass a public URL, and receive JSON with the transcript and, on request, word timings, speaker labels, tagged sounds, and SRT subtitles. One request takes up to two hours of audio. Confirm live rates on the SnapGen models page. Failed requests are not charged.
string
required
Set to scribe-v2.
file
The audio file, in a multipart field named file, up to 50 MB. Send file or audio_url, not both.
string
Public http or https URL of the audio. The gateway downloads it, up to 50 MB. Send audio_url or file, not both.
string
Language of the audio as an ISO 639 code of two or three lowercase letters, such as en. Omit it to detect the language automatically.
boolean
Labels who speaks each word in words[].speaker.
integer
The most speakers in the audio, from 1 to 32. Helps speaker labeling.
boolean
Tags sounds such as laughter or footsteps in the transcript.
string
none, word, or character. word and character add words with word-level timings.
boolean
true adds SubRip subtitles in srt. The gateway turns on speaker labels and word timings at the provider, which subtitles need.
number
Randomness, from 0 to 2. Higher values give more varied output.
string
default:"json"
json or verbose_json. verbose_json adds words unless timestamps is none. Both return the SnapGen JSON shape, not OpenAI’s verbose schema.

Request

Upload a local file as multipart/form-data:
Or send JSON with a public audio_url:
Response
The OpenAI SDKs send the same multipart upload, so a basic transcription works unchanged. Read text from the result; the extra fields in this response don’t follow OpenAI’s verbose schema.

Words, speakers, and subtitles

Price examples

The gateway measures the audio from the file itself and rounds up to whole seconds after a 0.1-second allowance, with a minimum of 1 second. Options such as subtitles or diarize don’t change the price.
Synchronous audio endpoints reject Idempotency-Key. If a request times out, check your Console request logs before you retry. See Speech to text for every response field and error.