> ## Documentation Index
> Fetch the complete documentation index at: https://docs.snapgen.org/llms.txt
> Use this file to discover all available pages before exploring further.

# Scribe v2

> ElevenLabs Scribe v2 speech to text with word timings, speaker labels, and SRT subtitles. $0.3312 per hour of audio.

`scribe-v2` transcribes speech with ElevenLabs Scribe v2 on the OpenAI-compatible transcription endpoint. Upload a file or pass a public URL, and receive JSON with the transcript and, on request, word timings, speaker labels, tagged sounds, and SRT subtitles. One request takes up to two hours of audio.

| Property | Value |
| - | - |
| Model ID | `scribe-v2` |
| Endpoint | `POST /v1/audio/transcriptions` (multipart file or JSON `audio_url`) |
| Output | JSON: `text`, `language`, `language_probability`, `duration_seconds`, optional `words` and `srt` |
| Limits | Up to 7,200 seconds (2 hours) and 50 MB per request |
| Input formats | WAV, AIFF, MP3, AAC, MP4, M4A, MOV, FLAC, Ogg, and WebM with a readable length |
| Billed on | Audio length measured by SnapGen, in whole seconds |
| Price | \$0.000092 per second (\$0.00552 per minute, \$0.3312 per hour) |

Confirm live rates on the [SnapGen models page](https://snapgen.org/models). Failed requests are not charged.

<ParamField body="model" type="string" required>
  Set to `scribe-v2`.
</ParamField>

<ParamField body="file" type="file">
  The audio file, in a multipart field named `file`, up to 50 MB. Send `file`
  or `audio_url`, not both.
</ParamField>

<ParamField body="audio_url" type="string">
  Public `http` or `https` URL of the audio. The gateway downloads it, up to
  50 MB. Send `audio_url` or `file`, not both.
</ParamField>

<ParamField body="language" type="string">
  Language of the audio as an ISO 639 code of two or three lowercase letters,
  such as `en`. Omit it to detect the language automatically.
</ParamField>

<ParamField body="diarize" type="boolean">
  Labels who speaks each word in `words[].speaker`.
</ParamField>

<ParamField body="num_speakers" type="integer">
  The most speakers in the audio, from `1` to `32`. Helps speaker labeling.
</ParamField>

<ParamField body="tag_audio_events" type="boolean">
  Tags sounds such as laughter or footsteps in the transcript.
</ParamField>

<ParamField body="timestamps" type="string">
  `none`, `word`, or `character`. `word` and `character` add `words` with
  word-level timings.
</ParamField>

<ParamField body="subtitles" type="boolean">
  `true` adds SubRip subtitles in `srt`. The gateway turns on speaker labels
  and word timings at the provider, which subtitles need.
</ParamField>

<ParamField body="temperature" type="number">
  Randomness, from `0` to `2`. Higher values give more varied output.
</ParamField>

<ParamField body="response_format" type="string" default="json">
  `json` or `verbose_json`. `verbose_json` adds `words` unless `timestamps` is
  `none`. Both return the SnapGen JSON shape, not OpenAI's verbose schema.
</ParamField>

## Request

Upload a local file as `multipart/form-data`:

```bash theme={null}
curl https://api.snapgen.org/v1/audio/transcriptions \
  -H "Authorization: Bearer $SNAPGEN_API_KEY" \
  -F model=scribe-v2 \
  -F file=@interview.mp3 \
  -F diarize=true \
  -F subtitles=true
```

Or send JSON with a public `audio_url`:

```bash theme={null}
curl https://api.snapgen.org/v1/audio/transcriptions \
  -H "Authorization: Bearer $SNAPGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "scribe-v2",
    "audio_url": "https://example.com/interview.mp3",
    "diarize": true,
    "subtitles": true
  }'
```

```json Response theme={null}
{
  "text": "Welcome to the show. Thanks for having me.",
  "language": "eng",
  "language_probability": 0.99,
  "duration_seconds": 12.34,
  "srt": "1\n00:00:00,120 --> 00:00:01,480\nWelcome to the show.\n\n2\n00:00:01,900 --> 00:00:03,020\nThanks for having me.\n"
}
```

The OpenAI SDKs send the same multipart upload, so a basic transcription works
unchanged. Read `text` from the result; the extra fields in this response
don't follow OpenAI's verbose schema.

<CodeGroup>
  ```python Python theme={null}
  import os

  from openai import OpenAI

  client = OpenAI(
      api_key=os.environ["SNAPGEN_API_KEY"],
      base_url="https://api.snapgen.org/v1",
  )

  with open("interview.mp3", "rb") as audio:
      transcript = client.audio.transcriptions.create(model="scribe-v2", file=audio)

  print(transcript.text)
  ```

  ```javascript Node.js theme={null}
  const fs = require("node:fs");
  const OpenAI = require("openai");

  async function main() {
    const client = new OpenAI({
      apiKey: process.env.SNAPGEN_API_KEY,
      baseURL: "https://api.snapgen.org/v1",
    });

    const transcript = await client.audio.transcriptions.create({
      model: "scribe-v2",
      file: fs.createReadStream("interview.mp3"),
    });

    console.log(transcript.text);
  }

  main();
  ```
</CodeGroup>

## Words, speakers, and subtitles

| You want | Send | You get |
| - | - | - |
| Plain text | Nothing extra | `text`, `language`, `duration_seconds` |
| Word timings | `timestamps: "word"` | `words[]` with `start` and `end` in seconds |
| Who said what | `diarize: true` and `timestamps: "word"` | `words[].speaker`, such as `speaker_0` |
| Subtitles | `subtitles: true` | `srt`, ready to save as a `.srt` file |
| Subtitles and word timings | `subtitles: true` and `timestamps: "word"` | `srt` and `words[]` |
| Sound tags | `tag_audio_events: true` | Tagged sounds in `text` and `words[]` |

## Price examples

| Length | Price |
| - | - |
| 1 minute | \$0.00552 |
| 10 minutes | \$0.0552 |
| 1 hour | \$0.3312 |
| 2 hours | \$0.6624 |

The gateway measures the audio from the file itself and rounds up to whole
seconds after a 0.1-second allowance, with a minimum of 1 second. Options such
as `subtitles` or `diarize` don't change the price.

<Note>
  Synchronous audio endpoints reject `Idempotency-Key`. If a request times
  out, check your Console request logs before you retry. See
  [Speech to text](/api-reference/audio-transcriptions) for every response
  field and error.
</Note>
