Sakaira

Transcribe

Available

Audio or video to text with timestamps. 1.6 credits per started minute (2.08 with keyterms).

What it does

Transcribe an audio or video file (https URL) into text with word timings and speaker labels. Models: scribe-v2 (default): transcripts with word timings and speakers. Returns the text plus an SRT subtitle file and a JSON file with every word's timing. Priced per started minute of the file, whose length is read before any charge; keyterms (names, jargon) add 30%. Call list_models for prices and describe_model for a model's options.

Parameters

Parameters of transcribe
ParameterTypeRequiredDefaultDescription
audio_urlstringRequired—The audio or video file to transcribe (https)
languagestringOptional—ISO 639 language code (en, es, …); detected when omitted2–8 characters
speakersbooleanOptionaltrueLabel who is speaking
audio_eventsbooleanOptionaltrueTag sounds like (laughter) or (applause)
keytermsstring[]Optional—Names or jargon to spell right, up to 100; adds 30% to the price
modelscribe-v2Optionalscribe-v2scribe-v2 (default): transcripts with word timings and speakers. Call list_models for what each model costs.
extrasobjectOptional—Model-specific parameters passed through (see describe_model). Priced or top-level parameters are not accepted here.
confirmbooleanOptionalfalseSet true to accept a quote above 500 credits

Which model should I use?

ModelBest forPrice 
ElevenLabs Scribe v2transcripts with word timings and speakers1.6 credits per started minute

Examples

  • Transcribe this interview recording with speaker labels
  • Get a transcript of this video with word-by-word timestamps
  • Transcribe this call and spell these names and terms correctly: Sakaira, fal.ai

Use cases

  • Turning a recorded interview or meeting into a written transcript
  • Getting word-level timing to build subtitles for a video
  • Making spoken content searchable and easy to quote from

How to use it

01

Connect Sakaira

One click or one URL. Sign in with Google. No API keys.

02

Ask in plain language

"Transcribe this interview recording with speaker labels"

03

Check the cost and balance in the reply

Every result tells you what it cost and what you have left.

FAQ

Which model does it use?

scribe-v2, the only transcription model in the catalog right now, with word timings and speaker labels across 90+ languages.

How much does it cost?

1.6 credits per started minute (2.08 with keyterms). The reply always states the exact cost and your remaining balance.

What do I get back?

The transcript text, plus an SRT subtitle file and a JSON file with every word's timing.

Ready to try transcribe?

Get 100 free credits