curl https://api.snapgen.org/v1/audio/alignment \
-H "Authorization: Bearer $SNAPGEN_API_KEY" \
-F model=eleven-forced-alignment \
-F file=@narration.mp3 \
-F text="The quick brown fox jumps over the lazy dog."
curl https://api.snapgen.org/v1/audio/alignment \
-H "Authorization: Bearer $SNAPGEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "eleven-forced-alignment",
"audio_url": "https://example.com/narration.mp3",
"text": "The quick brown fox jumps over the lazy dog."
}'
import os
import requests
with open("narration.mp3", "rb") as audio:
response = requests.post(
"https://api.snapgen.org/v1/audio/alignment",
headers={"Authorization": f"Bearer {os.environ['SNAPGEN_API_KEY']}"},
data={
"model": "eleven-forced-alignment",
"text": "The quick brown fox jumps over the lazy dog.",
},
files={"file": ("narration.mp3", audio, "audio/mpeg")},
timeout=600,
)
response.raise_for_status()
for word in response.json()["words"]:
print(word["start"], word["end"], word["text"])
import { openAsBlob } from "node:fs";
const form = new FormData();
form.append("model", "eleven-forced-alignment");
form.append("file", await openAsBlob("narration.mp3"), "narration.mp3");
form.append("text", "The quick brown fox jumps over the lazy dog.");
const response = await fetch("https://api.snapgen.org/v1/audio/alignment", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.SNAPGEN_API_KEY}` },
body: form,
});
if (!response.ok) throw new Error(await response.text());
const { words } = await response.json();
for (const word of words) console.log(word.start, word.end, word.text);
{
"words": [
{ "text": "The", "start": 0.08, "end": 0.21, "loss": 0.12 },
{ "text": "quick", "start": 0.21, "end": 0.48, "loss": 0.09 },
{ "text": "brown", "start": 0.48, "end": 0.77, "loss": 0.11 }
],
"loss": 0.1
}
{
"error": {
"message": "The audio must be at most 3600 seconds long",
"type": "invalid_request_error",
"param": null,
"code": "audio_too_long"
}
}
Endpoint Reference
Forced alignment
Align a transcript you already have to its audio and get the start and end time of every word.
POST
/
v1
/
audio
/
alignment
curl https://api.snapgen.org/v1/audio/alignment \
-H "Authorization: Bearer $SNAPGEN_API_KEY" \
-F model=eleven-forced-alignment \
-F file=@narration.mp3 \
-F text="The quick brown fox jumps over the lazy dog."
curl https://api.snapgen.org/v1/audio/alignment \
-H "Authorization: Bearer $SNAPGEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "eleven-forced-alignment",
"audio_url": "https://example.com/narration.mp3",
"text": "The quick brown fox jumps over the lazy dog."
}'
import os
import requests
with open("narration.mp3", "rb") as audio:
response = requests.post(
"https://api.snapgen.org/v1/audio/alignment",
headers={"Authorization": f"Bearer {os.environ['SNAPGEN_API_KEY']}"},
data={
"model": "eleven-forced-alignment",
"text": "The quick brown fox jumps over the lazy dog.",
},
files={"file": ("narration.mp3", audio, "audio/mpeg")},
timeout=600,
)
response.raise_for_status()
for word in response.json()["words"]:
print(word["start"], word["end"], word["text"])
import { openAsBlob } from "node:fs";
const form = new FormData();
form.append("model", "eleven-forced-alignment");
form.append("file", await openAsBlob("narration.mp3"), "narration.mp3");
form.append("text", "The quick brown fox jumps over the lazy dog.");
const response = await fetch("https://api.snapgen.org/v1/audio/alignment", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.SNAPGEN_API_KEY}` },
body: form,
});
if (!response.ok) throw new Error(await response.text());
const { words } = await response.json();
for (const word of words) console.log(word.start, word.end, word.text);
{
"words": [
{ "text": "The", "start": 0.08, "end": 0.21, "loss": 0.12 },
{ "text": "quick", "start": 0.21, "end": 0.48, "loss": 0.09 },
{ "text": "brown", "start": 0.48, "end": 0.77, "loss": 0.11 }
],
"loss": 0.1
}
{
"error": {
"message": "The audio must be at most 3600 seconds long",
"type": "invalid_request_error",
"param": null,
"code": "audio_too_long"
}
}
POST /v1/audio/alignment matches a known transcript to its audio and returns the start and end time of each word. Use it to time captions, karaoke lyrics, or highlighted read-along text when you already have the exact words. To transcribe audio you don’t have text for, use Speech to text.
| Model | Input | Price |
|---|---|---|
eleven-forced-alignment | Up to 3,600 seconds (1 hour) and 50 MB | $0.000092 per second ($0.3312 per hour) |
string
required
Set to
eleven-forced-alignment.file
The audio file, in a multipart field named
file, up to 50 MB. Send file
or audio_url, not both.string
Public
http or https URL of the audio. The gateway downloads it, up to
50 MB. Send audio_url or file, not both.string
required
The transcript to align, up to 100,000 characters. The text isn’t billed.
model field, one file in file,
and text as a form field) or as JSON with audio_url. The gateway accepts
WAV, AIFF, MP3, AAC, MP4, M4A, MOV, FLAC, Ogg, and WebM files whose length can
be read.
Response
object[]
number | null
The provider’s alignment loss for the whole text.
x-gateway-charge-microusd header.
Billing
The gateway bills the measured length of the audio at $0.000092 per second, rounded up to whole seconds after a 0.1-second allowance, with a minimum of 1 second. A 10-minute recording costs $0.0552. Failed requests aren’t charged. This endpoint rejectsIdempotency-Key; if a request times out, check your
Console request logs before you retry.
Errors
| Status | error.code | Cause |
|---|---|---|
| 400 | audio_input_required | Neither file nor audio_url was sent. |
| 400 | invalid_request | text is missing, blank, or too long, a field is unknown, or both file and audio_url were sent. |
| 400 | invalid_multipart | The form has no model field, more than one file, a repeated field, or the file isn’t in the file field. |
| 400 | unsupported_audio_format | The format isn’t supported, or its length can’t be read. |
| 400 | audio_too_long | The audio is longer than 3,600 seconds. |
| 400 | audio_download_failed | The gateway couldn’t download audio_url. |
| 400 | provider_rejected_request | The provider refused the audio or text. |
| 402 | insufficient_funds | Your balance can’t cover the request. |
| 413 | audio_too_large, request_too_large | The downloaded file or the upload is larger than 50 MB. |
curl https://api.snapgen.org/v1/audio/alignment \
-H "Authorization: Bearer $SNAPGEN_API_KEY" \
-F model=eleven-forced-alignment \
-F file=@narration.mp3 \
-F text="The quick brown fox jumps over the lazy dog."
curl https://api.snapgen.org/v1/audio/alignment \
-H "Authorization: Bearer $SNAPGEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "eleven-forced-alignment",
"audio_url": "https://example.com/narration.mp3",
"text": "The quick brown fox jumps over the lazy dog."
}'
import os
import requests
with open("narration.mp3", "rb") as audio:
response = requests.post(
"https://api.snapgen.org/v1/audio/alignment",
headers={"Authorization": f"Bearer {os.environ['SNAPGEN_API_KEY']}"},
data={
"model": "eleven-forced-alignment",
"text": "The quick brown fox jumps over the lazy dog.",
},
files={"file": ("narration.mp3", audio, "audio/mpeg")},
timeout=600,
)
response.raise_for_status()
for word in response.json()["words"]:
print(word["start"], word["end"], word["text"])
import { openAsBlob } from "node:fs";
const form = new FormData();
form.append("model", "eleven-forced-alignment");
form.append("file", await openAsBlob("narration.mp3"), "narration.mp3");
form.append("text", "The quick brown fox jumps over the lazy dog.");
const response = await fetch("https://api.snapgen.org/v1/audio/alignment", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.SNAPGEN_API_KEY}` },
body: form,
});
if (!response.ok) throw new Error(await response.text());
const { words } = await response.json();
for (const word of words) console.log(word.start, word.end, word.text);
{
"words": [
{ "text": "The", "start": 0.08, "end": 0.21, "loss": 0.12 },
{ "text": "quick", "start": 0.21, "end": 0.48, "loss": 0.09 },
{ "text": "brown", "start": 0.48, "end": 0.77, "loss": 0.11 }
],
"loss": 0.1
}
{
"error": {
"message": "The audio must be at most 3600 seconds long",
"type": "invalid_request_error",
"param": null,
"code": "audio_too_long"
}
}