Transcribe
AvailableAudio or video to text with timestamps. 1.6 credits per started minute (2.08 with keyterms).
What it does
Transcribe an audio or video file (https URL) into text with word timings and speaker labels. Models: scribe-v2 (default): transcripts with word timings and speakers. Returns the text plus an SRT subtitle file and a JSON file with every word's timing. Priced per started minute of the file, whose length is read before any charge; keyterms (names, jargon) add 30%. Call list_models for prices and describe_model for a model's options.
Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| audio_url | string | Required | — | The audio or video file to transcribe (https) |
| language | string | Optional | — | ISO 639 language code (en, es, …); detected when omitted2–8 characters |
| speakers | boolean | Optional | true | Label who is speaking |
| audio_events | boolean | Optional | true | Tag sounds like (laughter) or (applause) |
| keyterms | string[] | Optional | — | Names or jargon to spell right, up to 100; adds 30% to the price |
| model | scribe-v2 | Optional | scribe-v2 | scribe-v2 (default): transcripts with word timings and speakers. Call list_models for what each model costs. |
| extras | object | Optional | — | Model-specific parameters passed through (see describe_model). Priced or top-level parameters are not accepted here. |
| confirm | boolean | Optional | false | Set true to accept a quote above 500 credits |
Which model should I use?
| Model | Best for | Price | |
|---|---|---|---|
| ElevenLabs Scribe v2 | transcripts with word timings and speakers | 1.6 credits per started minute |
Examples
- Transcribe this interview recording with speaker labels
- Get a transcript of this video with word-by-word timestamps
- Transcribe this call and spell these names and terms correctly: Sakaira, fal.ai
Use cases
- Turning a recorded interview or meeting into a written transcript
- Getting word-level timing to build subtitles for a video
- Making spoken content searchable and easy to quote from
How to use it
Connect Sakaira
One click or one URL. Sign in with Google. No API keys.
Ask in plain language
"Transcribe this interview recording with speaker labels"
Check the cost and balance in the reply
Every result tells you what it cost and what you have left.
FAQ
Which model does it use?
scribe-v2, the only transcription model in the catalog right now, with word timings and speaker labels across 90+ languages.
How much does it cost?
1.6 credits per started minute (2.08 with keyterms). The reply always states the exact cost and your remaining balance.
What do I get back?
The transcript text, plus an SRT subtitle file and a JSON file with every word's timing.
Ready to try transcribe?
Get 100 free credits