Create voiceover
AvailableText to speech. From 2 to 20 credits per 1,000 characters, depending on the model.
What it does
Turn text into speech: voiceovers, narration, ads. Models: gemini-3.8-flash-tts (default): natural voices directed in plain English; gemini-3.8-flash-lite-tts: the same voices, cheaper; elevenlabs-v3: expressive voices with emotion tags and word timings; elevenlabs-multilingual-v2: steady narration with speed control; minimax-speech-2.8-hd: emotional voices for long texts; minimax-speech-2.8-turbo: MiniMax voices at a lower price; xai-tts: cheap voices with laughs and whispers; inworld-tts-1.5: low-cost voices, 73 in English; kokoro: fast English drafts; chatterbox-hd: dramatic voices with adjustable intensity; dia: two-speaker dialogue; orpheus: open-source English narration. Pick a voice with voice (describe_model lists each model's voices); style gives the Gemini voices directions in plain English; speed where the model has it. Priced per character of text. No voice cloning. Returns a permanent audio URL plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options.
Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| text | string | Required | — | The words to speak, verbatim. Each model has its own length cap (describe_model).Up to 15,000 characters |
| voice | string | Optional | — | Voice name from the model's list (describe_model); the model's default when omittedUp to 100 characters |
| style | string | Optional | — | How to say it, in plain English, e.g. 'warm and slow, like a bedtime story' (Gemini models)Up to 1,000 characters |
| speed | number | Optional | — | Speaking speed, 1 is normal (ElevenLabs Multilingual v2 0.7–1.2, MiniMax and Kokoro 0.5–2)0.5–2 |
| model | gemini-3.8-flash-tts | gemini-3.8-flash-lite-tts | elevenlabs-v3 | elevenlabs-multilingual-v2 | minimax-speech-2.8-hd | minimax-speech-2.8-turbo | xai-tts | inworld-tts-1.5 | kokoro | chatterbox-hd | dia | orpheus | Optional | gemini-3.8-flash-tts | gemini-3.8-flash-tts (default): natural voices directed in plain English; gemini-3.8-flash-lite-tts: the same voices, cheaper; elevenlabs-v3: expressive voices with emotion tags and word timings; elevenlabs-multilingual-v2: steady narration with speed control; minimax-speech-2.8-hd: emotional voices for long texts; minimax-speech-2.8-turbo: MiniMax voices at a lower price; xai-tts: cheap voices with laughs and whispers; inworld-tts-1.5: low-cost voices, 73 in English; kokoro: fast English drafts; chatterbox-hd: dramatic voices with adjustable intensity; dia: two-speaker dialogue; orpheus: open-source English narration. Call list_models for what each model costs. describe_model lists each model's voices. |
| extras | object | Optional | — | Model-specific parameters passed through (see describe_model). Priced or top-level parameters are not accepted here. |
| confirm | boolean | Optional | false | Set true to accept a quote above 500 credits |
Which model should I use?
| Model | Best for | Price | |
|---|---|---|---|
| Gemini 3.8 Flash TTS | natural voices directed in plain English | 9 credits per 1,000 characters | |
| Gemini 3.8 Flash-Lite TTS | the same voices, cheaper | 6 credits per 1,000 characters | |
| ElevenLabs v3 | expressive voices with emotion tags and word timings | 20 credits per 1,000 characters | |
| ElevenLabs Multilingual v2 | steady narration with speed control | 20 credits per 1,000 characters | |
| MiniMax Speech 2.8 HD | emotional voices for long texts | 20 credits per 1,000 characters | |
| MiniMax Speech 2.8 Turbo | MiniMax voices at a lower price | 12 credits per 1,000 characters | |
| xAI TTS | cheap voices with laughs and whispers | 3 credits per 1,000 characters | |
| Inworld TTS 1.5 Max | low-cost voices, 73 in English | 2 credits per 1,000 characters | |
| Kokoro | fast English drafts | 4 credits per 1,000 characters | |
| Chatterbox HD | dramatic voices with adjustable intensity | 8 credits per 1,000 characters | |
| Dia | two-speaker dialogue | 8 credits per 1,000 characters | |
| Orpheus | open-source English narration | 10 credits per 1,000 characters |
Examples
- Turn this script into a voiceover: "Welcome to our store, where quality meets..."
- Read this paragraph in a warm, slow voice, like a bedtime story
- Narrate this product description at a slightly faster pace
Use cases
- Recording a voiceover for a product video or ad without a studio session
- Turning a blog post or script into narration for a podcast intro
- Prototyping a few voice styles before picking one for a final recording
How to use it
Connect Sakaira
One click or one URL. Sign in with Google. No API keys.
Ask in plain language
"Turn this script into a voiceover: "Welcome to our store, where quality meets...""
Check the cost and balance in the reply
Every result tells you what it cost and what you have left.
FAQ
Which model should I use?
gemini-3.8-flash-tts is the default and takes plain-English style directions. Call list_models to see every voice model available to you and what each one costs.
How much does it cost?
From 2 to 20 credits per 1,000 characters, depending on the model. The reply always states the exact cost and your remaining balance.
Can I clone my own voice?
Not yet — every model speaks with its own built-in voices; pick one with voice (describe_model lists each model's voices).
Ready to try create voiceover?
Get 100 free credits