Skip to main content
POST
POST /v1/audio/transcriptions transcribes speech with scribe-v2. Upload the file as multipart/form-data, the same way as the OpenAI transcription API, or send JSON with a public audio_url that the gateway downloads. The response is JSON with the transcript and, on request, word timings, speaker labels, and SRT subtitles.
string
required
Set to scribe-v2.
file
The audio file, in a multipart field named file, up to 50 MB. Send file or audio_url, not both.
string
Public http or https URL of the audio. The gateway downloads it over a public-only connection, up to 50 MB. Send audio_url or file, not both.
string
Language of the audio as an ISO 639 code of two or three lowercase letters, such as en. Omit it to detect the language automatically.
boolean
Labels who speaks each word in words[].speaker.
integer
The most speakers in the audio, from 1 to 32. Helps speaker labeling.
boolean
Tags sounds such as laughter or footsteps in the transcript.
string
none, word, or character. word and character add words to the response with word-level timings; per-character timings aren’t returned.
boolean
true adds SubRip (SRT) subtitles in srt. Subtitles need speaker labels and word timings, so the gateway turns both on at the provider. To also get words in the response, send timestamps: "word".
number
Randomness, from 0 to 2. Higher values give more varied output.
string
default:"json"
json or verbose_json. verbose_json adds words unless timestamps is none. Both return the SnapGen JSON shape below, not OpenAI’s verbose schema.

Send the audio

  • Multipart upload: send exactly one model field and one file in the file field. Send the other fields as text form fields; true, false, and numbers are converted for you.
  • JSON: send Content-Type: application/json with audio_url and the other fields.
The gateway reads the length from the audio itself; a declared duration is never used. It accepts WAV, AIFF, MP3, AAC, MP4, M4A, MOV, FLAC, Ogg, and WebM files whose length can be read. Anything else returns unsupported_audio_format before any balance is reserved.

Response

string
The full transcript.
string | null
The language code the provider detected or used, for example eng.
number | null
Confidence in the detected language, from 0 to 1.
number
Length of the audio in seconds.
object[]
Present when timestamps is word or character, or with response_format: "verbose_json".
string
SubRip subtitles. Present when subtitles is true.
When the charge has settled, the response carries it in micro-USD in the x-gateway-charge-microusd header.

Billing

The gateway bills the measured length of the audio at $0.000092 per second. It rounds up to whole seconds after a 0.1-second allowance for container padding, with a minimum of 1 second: a 12.34-second file bills 13 seconds ($0.001196). A one-hour recording costs $0.3312. Options such as subtitles and diarize don’t change the price, and failed requests aren’t charged. This endpoint rejects Idempotency-Key. If a request times out, check your Console request logs before you retry.

Errors