Skip to main content
POST
POST /v1/audio/alignment matches a known transcript to its audio and returns the start and end time of each word. Use it to time captions, karaoke lyrics, or highlighted read-along text when you already have the exact words. To transcribe audio you don’t have text for, use Speech to text.
string
required
Set to eleven-forced-alignment.
file
The audio file, in a multipart field named file, up to 50 MB. Send file or audio_url, not both.
string
Public http or https URL of the audio. The gateway downloads it, up to 50 MB. Send audio_url or file, not both.
string
required
The transcript to align, up to 100,000 characters. The text isn’t billed.
Send the audio as a multipart upload (one model field, one file in file, and text as a form field) or as JSON with audio_url. The gateway accepts WAV, AIFF, MP3, AAC, MP4, M4A, MOV, FLAC, Ogg, and WebM files whose length can be read.

Response

object[]
One entry per word, in order.
number | null
The provider’s alignment loss for the whole text.
When the charge has settled, the response carries it in micro-USD in the x-gateway-charge-microusd header.

Billing

The gateway bills the measured length of the audio at $0.000092 per second, rounded up to whole seconds after a 0.1-second allowance, with a minimum of 1 second. A 10-minute recording costs $0.0552. Failed requests aren’t charged. This endpoint rejects Idempotency-Key; if a request times out, check your Console request logs before you retry.

Errors

See Errors for rate limits and provider failures.