# Sakaira > Sakaira is a remote MCP server that gives Claude, ChatGPT and other AI assistants access to the top AI image, video and audio models. MCP endpoint (Streamable HTTP, OAuth 2.1): https://mcp.sakaira.com/mcp MCP reference (sign-in, every tool's parameters, errors): https://sakaira.com/mcp (https://sakaira.com/mcp.md) Website: https://sakaira.com Connect guides: https://sakaira.com/connect Pricing: https://sakaira.com/pricing (https://sakaira.com/pricing.md) ## Image ### create_image https://sakaira.com/mcp#create_image https://sakaira.com/tools/create-image Generate 1–4 images from a prompt, or edit images by passing image_url (more references in extras.image_urls). Models: nano-banana-2 (default): fast all-round images and edits, 0.5K to 4K; nano-banana-pro: harder edits and detailed scenes, up to 4K; nano-banana-2-lite: cheap 1K drafts; gpt-image-2.5-sunburst: accurate text and edits; gpt-image-2.5-flare: the same quality, faster; muse-image: high quality at a low price; grok-imagine-image-2: posters and designs with legible text; seedream-5-pro: consistent characters from many references; seedream-5-lite: cheap, consistent sets of images; flux-2-pro: photoreal images at exact sizes; flux-2-max: FLUX's highest detail; flux-2-flex: typography and fine control; flux-2-klein: very fast, very cheap drafts; qwen-image-3: long, detailed prompts; ideogram-4: logos and typography; recraft-4.1: brand illustrations; recraft-4.1-vector: editable SVG icons and logos; mai-image-2.5: realistic scenes and portraits; krea-2: stylized images from style references; luma-uni-1: concept art and moody frames. Set quality (low, medium, high) where a model offers it. Call list_models for prices and describe_model for a model's options. Returns fixed URLs, each kept until its expires_at, plus the cost and your remaining balance. Price: From 2 credits per image (model-dependent). remove_background adds 2 per image. ### upscale_image https://sakaira.com/mcp#upscale_image https://sakaira.com/tools/upscale-image Make an image 1–4× bigger (scale, 2 by default) from its https URL. Models: topaz-upscale (default): faithful upscales up to 4×; recraft-crisp-upscale: cheap, sharp upscales of PNGs; seedvr2-upscale: very cheap large upscales; clarity-upscaler: creative upscales that add detail. recraft-crisp-upscale takes PNG only. The price follows the output size, read from the image before any charge. Returns a fixed URL, kept until its expires_at, plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options. Price: Topaz (the default): 16 credits per started 24 megapixels of output. The others: from 1 credit per image (Recraft Crisp, SeedVR2) to 6 credits per output megapixel (Clarity). ### remove_background https://sakaira.com/mcp#remove_background https://sakaira.com/tools/remove-background Cut out the subject of an image (https URL; PNG, JPEG or WebP) and return a transparent PNG. Models: ideogram-remove-background (default): transparent PNG cut-outs; bria-rmbg-2: cut-outs from a model trained on licensed data. Returns a fixed URL, kept until its expires_at, plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options. Price: 2 credits per image (Bria: 4). ### restore_image https://sakaira.com/mcp#restore_image https://sakaira.com/tools/restore-image Repair an old or damaged photo from its https URL: remove scratches, fix or add colors (colorize) and raise the resolution. Models: photo-restoration (default): repairing and colorizing old photos. Returns a fixed URL, kept until its expires_at, plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options. Price: 8 credits per image. ## Video ### create_video https://sakaira.com/mcp#create_video https://sakaira.com/tools/create-video Generate a video from a prompt, animate an image (image_url is the first frame, end_image_url the last), or keep subjects consistent with reference_image_urls. Models: gemini-omni-flash (default): video with sound; veo-3.1: Veo at up to 4K; veo-3.1-fast: Veo quality at a lower price, up to 4K; veo-3.1-lite: cheap clips with sound, or silent b-roll; wan-3: single shots up to 30 s at native 1080p; happyhorse-1.1: talking characters with lip-sync; minimax-h3-max: fast image-to-video with sound; minimax-h3-max-turbo: the cheapest H3 Max clips; minimax-h3: the open-weights H3, up to 4K; seedance-2.5: ads up to 30 s with many reference images; seedance-2: cinematic multi-shot clips; seedance-2-fast: Seedance 2.0 at a lower price; seedance-2-mini: cheap Seedance drafts; kling-3-pro: cinematic motion; kling-3-turbo: faster Kling with sound; kling-o3: consistent characters from reference images; grok-imagine-video-1.5: short clips with sound from text or images; flux-3: cinematic clips from text or a first and last frame; ltx-2.5: open-weights clips up to 20 s and 4K; pixverse-6: cheap social clips with optional sound; vidu-q3: clips up to 16 s with start and end frames; luma-ray-3.2: cinematic silent shots; creatify-boreal: product, UGC and presenter ads from a script. Duration, resolution and aspect ratio are fitted to what the model makes, and any change is listed in warnings. Returns a job_id at once: videos take about 1–5 minutes, so call get_job with that job_id to get the fixed URL, kept until its expires_at. Call list_models for prices and describe_model for each model's options. Quotes above 500 credits need confirm: true. Price: From 2 credits per second, depending on the model, resolution and sound. ## Audio ### create_voiceover https://sakaira.com/mcp#create_voiceover https://sakaira.com/tools/create-voiceover Turn text into speech: voiceovers, narration, ads. Models: gemini-3.8-flash-tts (default): natural voices directed in plain English; gemini-3.8-flash-lite-tts: the same voices, cheaper; elevenlabs-v3: expressive voices with emotion tags and word timings; elevenlabs-multilingual-v2: steady narration with speed control; minimax-speech-2.8-hd: emotional voices for long texts; minimax-speech-2.8-turbo: MiniMax voices at a lower price; xai-tts: cheap voices with laughs and whispers; inworld-tts-1.5: low-cost voices, 73 in English; kokoro: fast English drafts; chatterbox-hd: dramatic voices with adjustable intensity; dia: two-speaker dialogue; orpheus: open-source English narration. Pick a voice with voice (describe_model lists each model's voices); style gives the Gemini voices directions in plain English; speed where the model has it. Priced per character of text. No voice cloning. Returns a fixed audio URL, kept until its expires_at, plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options. Price: From 2 to 20 credits per 1,000 characters, depending on the model. ### create_music https://sakaira.com/mcp#create_music https://sakaira.com/tools/create-music Generate a music track from a prompt (genre, mood, instruments, tempo), with lyrics or as an instrumental. Models: lyria-3.5 (default): songs or instrumental tracks up to about 3 minutes; elevenlabs-music-2.5: tracks with an exact length for ads and videos; minimax-music-2.6: songs with lyrics or instrumental tracks; stable-audio-3: instrumental beds and ambience with an exact length. duration_seconds is exact on elevenlabs-music-2.5 and stable-audio-3 (30 s by default there) and a hint on lyria-3.5. Returns a fixed MP3 URL, kept until its expires_at, plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options. Price: From 8 credits per track. ### transcribe https://sakaira.com/mcp#transcribe https://sakaira.com/tools/transcribe Transcribe an audio or video file (https URL) into text with word timings and speaker labels. Models: scribe-v2 (default): transcripts with word timings and speakers. Returns the text plus an SRT subtitle file and a JSON file with every word's timing. Priced per started minute of the file, whose length is read before any charge; keyterms (names, jargon) add 30%. Call list_models for prices and describe_model for a model's options. Price: 1.6 credits per started minute (2.08 with keyterms). ### enhance_audio https://sakaira.com/mcp#enhance_audio https://sakaira.com/tools/enhance-audio Clean a voice recording from an audio or video file (https URL): elevenlabs-audio-isolation removes background noise, music and reverb; deepfilternet-3 removes noise from audio files. Models: elevenlabs-audio-isolation (default): removing noise and music from a voice; deepfilternet-3: cheap noise removal. Priced per minute of the file, whose length is read before any charge. Returns a fixed audio URL, kept until its expires_at, plus the cost and your remaining balance. Call list_models for prices and describe_model for a model's options. Price: 20 credits per started minute (DeepFilterNet: 12 per minute). ## Studio: brand kit and files ### save_brand_kit https://sakaira.com/mcp#save_brand_kit https://sakaira.com/tools/brand-kit Create or update the account's brand kit (one per account). Call it when the user wants to save or change their brand. Send only what changes: a field left out or null keeps its saved value. The fields: name (needed the first time), up to 6 colors as #RRGGBB, up to 3 fonts by name (heading, body or accent), tone (up to 500 characters), notes (up to 1,000) and voice (a create_voiceover model and one of its voices). colors and fonts replace the whole list, and [] clears one; "" clears tone or notes; clear_voice: true removes the voice. With a paid plan the kit also takes images by https URL (PNG, JPEG or WebP, up to 20 MB each): one logo, which replaces the current one, and up to 10 reference images; remove_image_ids removes saved ones. Returns the kit and edit_url, the Brand kit page where the user can drop files, plus, with a paid plan, upload: a ready command to upload a file from this computer. Free. Price: Free. ### get_brand_kit https://sakaira.com/mcp#get_brand_kit https://sakaira.com/tools/brand-kit Return the account's brand kit: name, colors, fonts, tone, notes, voice, and its logo and reference images with their ids and URLs. Call it before making anything for the user's brand (an image, a video, a voiceover, a post) and apply it: put its colors, fonts and tone into the prompt; pass the logo or reference image URLs as reference images only when the piece should show them; never ask a model to redraw the logo from memory; use the kit's voice for voiceovers unless the user picks another. The kit is null when there is none yet (create one with save_brand_kit). Also returns edit_url, the Brand kit page in the user's Library where they can drop files, how_to_use and, with a paid plan, upload: a ready command to upload a file from this computer. Free. Price: Free. ### delete_brand_kit https://sakaira.com/mcp#delete_brand_kit https://sakaira.com/tools/brand-kit Delete the account's brand kit for good, with its logo and reference images: their URLs stop working. Call it only when the user asks to delete their brand kit; to change it, call save_brand_kit. Files in the user's library are not touched. Free. Price: Free. ### delete_asset https://sakaira.com/mcp#delete_asset Delete a file for good: it leaves your library and its URL stops working. Pass asset_id or the file's url. Free; the credits it cost are not refunded. Price: Free. ## Account and discovery ### list_models https://sakaira.com/mcp#list_models List the available models with vendor, category, price in credits and which tools use them. Filter by category (image, video, voice, music, audio, utility) or free text. Free. Price: Free. ### describe_model https://sakaira.com/mcp#describe_model Return a model’s parameters (name, type, default), pricing rule and the tools that use it. Free. Price: Free. ### check_balance https://sakaira.com/mcp#check_balance Return your plan and credits: the subscription's plan with its renewal or end date, your balance, and each lot of credits with its expiry (a plan's credits at the end of their month, a pack's 90 days after purchase, the free ones 30 days after signup), and without a subscription the date your next file is deleted. Optionally the deterministic cost of a tool call given its input. Free. Price: Free. ### buy_credits https://sakaira.com/mcp#buy_credits Credit packs are for subscribers: create a secure Stripe Checkout link for a pack of 700 to 7,000 credits, which expire 90 days after purchase. Without a subscription it returns the plans link instead, where the user subscribes; plans are changed there too. The checkout page shows the price. Free. Price: Free. ### get_job https://sakaira.com/mcp#get_job Return the status (queued, running, succeeded, failed) and, when done, the fixed URLs (each kept until its expires_at) and cost of a job created by create_video or any other asynchronous tool. Without job_id, list your 10 most recent jobs, newest first, with their status, cost and file URLs: use it to recover a result whose reply was lost or timed out, instead of generating again. Free. Price: Free. ### spend_report https://sakaira.com/mcp#spend_report Summarize credits spent today, in the last 7 or 30 days, or all time, broken down by tool. Free. Price: Free. ### send_feedback https://sakaira.com/mcp#send_feedback Send feedback to the Sakaira team from inside the assistant: kind (bug, missing, confusing, praise) and a message. Free. Price: Free. ## Models Every model and its price: https://sakaira.com/models ### Nano Banana (Google) Nano Banana 2 & Pro: fast, versatile images and edits with up to 14 reference images, from drafts to 4K. - Nano Banana 2 (nano-banana-2) · available · 12 credits per image (0.5K); 16 credits per image (1K); 24 credits per image (2K); 32 credits per image (4K) · used by create_image Parameters: - prompt (string): What to generate or how to edit the reference image - aspect_ratio (enum) default auto: auto, 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 - resolution (enum) default 1K: 0.5K, 1K, 2K, 4K - num_images (integer) default 1: 1–4 - output_format (enum) default png: png, jpeg, webp - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: up to 14 reference images - Nano Banana Pro (nano-banana-pro) · available · 30 credits per image (1K); 30 credits per image (2K); 60 credits per image (4K) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - aspect_ratio (enum) default auto: auto, 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 - resolution (enum) default 1K: 1K, 2K, 4K - num_images (integer) default 1: 1–4 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg, webp - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: up to 14 reference images - system_prompt (string): A system instruction that steers style across the request - enable_web_search (boolean): Web search grounding - Nano Banana 2 Lite (nano-banana-2-lite) · available · 9 credits per image; +0.1 credits per reference image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - aspect_ratio (enum) default auto: auto, 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16, 4:1, 1:4, 8:1, 1:8 - num_images (integer) default 1: 1–4 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg, webp - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: up to 14 reference images - system_prompt (string): A system instruction that steers style across the request - thinking_level (enum): minimal, high ### GPT Image (OpenAI) GPT Image 2.5: accurate text in images and precise edits. - GPT Image 2.5 Sunburst (gpt-image-2.5-sunburst) · available · 2 credits per image (1024×1024 low); 4 credits per image (1024×1024 medium); 13 credits per image (1024×1024 high); 3 credits per image (3840×2160 low); 7 credits per image (3840×2160 medium); 25 credits per image (3840×2160 high); +7.2 credits per reference image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 32,000 characters) - image_size (object): {width, height}: sides in multiples of 16, up to 3840 px, ratio up to 3:1; 1024×1024 by default - quality (enum) default high: low, medium, high - num_images (integer) default 1: 1–10 on fal; create_image asks for 1–4 - background (enum) default auto: auto, transparent, opaque - output_format (enum) default png: png, jpeg, webp - output_compression (integer): 0–100, for jpeg and webp - image_urls (string[]): Edit mode: up to 16 reference images; each one adds to the price - mask_url (string): Edit mode: a mask of the area to change - GPT Image 2.5 Flare (gpt-image-2.5-flare) · available · 2 credits per image (1024×1024 low); 4 credits per image (1024×1024 medium); 13 credits per image (1024×1024 high); 3 credits per image (3840×2160 low); 7 credits per image (3840×2160 medium); 25 credits per image (3840×2160 high); +7.2 credits per reference image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 32,000 characters) - image_size (object): {width, height}: sides in multiples of 16, up to 3840 px, ratio up to 3:1; 1024×1024 by default - quality (enum) default high: low, medium, high - num_images (integer) default 1: 1–10 on fal; create_image asks for 1–4 - background (enum) default auto: auto, transparent, opaque - output_format (enum) default png: png, jpeg, webp - output_compression (integer): 0–100, for jpeg and webp - image_urls (string[]): Edit mode: up to 16 reference images; each one adds to the price - mask_url (string): Edit mode: a mask of the area to change ### Muse Image (Meta) Muse Image: high image quality at a low price. - Muse Image (muse-image) · available · 2 credits per image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - aspect_ratio (enum): any w:h from 1:16 to 16:1; the output is about 2.5 MP - num_images (integer) default 1: 1–10 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg, webp - image_urls (string[]): Edit mode: 1–10 reference images ### Grok Imagine (xAI) Grok Imagine: posters and designs with legible text, plus short videos with sound. - Grok Imagine Image 2.0 (grok-imagine-image-2) · available · 8 credits per image (1K low); 12 credits per image (1K medium); 12 credits per image (2K low); 16 credits per image (2K medium); +2 credits per reference image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 8,000 characters) - aspect_ratio (enum) default 1:1: 2:1, 20:9, 19.5:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:19.5, 9:20, 1:2 - resolution (enum) default 1k: 1k, 2k - quality (enum) default medium: low, medium - num_images (integer) default 1: 1–4 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg, webp - image_urls (string[]): Edit mode: up to 5 images; each one adds to the price - Grok Imagine Video 1.5 (grok-imagine-video-1.5) · available · 16 credits per second (480p); 28 credits per second (720p); 50 credits per second (1080p); 5 s (720p) = 140 credits; +2 credits per reference image · used by create_video Parameters: - prompt (string): What happens in the video; up to 4,096 characters - duration (integer) default 5: Seconds: 1–15 s; other values are fitted - resolution (string) default 720p: 480p, 720p or 1080p; other values are fitted - aspect_ratio (string): 16:9, 4:3, 3:2, 1:1, 2:3, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - reference_image_urls (string[]): Up to 7 images of subjects or products to keep consistent; not together with image_url or end_image_url ### Seedream (ByteDance) Seedream 5.0: consistent characters across many reference images and dense layouts. - Seedream 5.0 Pro (seedream-5-pro) · available · 14 credits per image (up to 1536×1536); 27 credits per image (up to 2048×2048); +0.9 credits per reference image after the first · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: 1024×1024 to 2048×2048 pixels in total, ratio 1:16 to 16:1; 1536×1536 by default - num_images (integer) default 1: 1–6 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg - image_urls (string[]): Edit mode: up to 10 reference images; the first is free, each more adds to the price - Seedream 5.0 Lite (seedream-5-lite) · available · 7 credits per image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: 2560×1440 to 4096×4096 pixels in total; 2048×2048 by default - num_images (integer) default 1: 1–6 on fal; create_image asks for 1–4 - max_images (integer): Several images from one generation - image_urls (string[]): Edit mode: up to 10 reference images ### FLUX (Black Forest Labs) FLUX.2 & FLUX 3: photoreal images at exact sizes, and FLUX 3 video guided by keyframes. - FLUX.2 [pro] (flux-2-pro) · available · 6 credits for the first megapixel, +3 credits per extra megapixel (input images count); each reference image counts as 5 megapixels (+15 credits) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: up to 4 MP; 1024×768 by default - output_format (enum) default png: png, jpeg - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: up to 4 reference images; each counts as 5 megapixels - safety_tolerance (enum): 1–5 - FLUX.2 [max] (flux-2-max) · available · 14 credits for the first megapixel, +6 credits per extra megapixel (input images count); each reference image counts as 5 megapixels (+30 credits) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: up to 4 MP; 1024×768 by default - output_format (enum) default png: png, jpeg - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: up to 4 reference images; each counts as 5 megapixels - safety_tolerance (enum): 1–5 - FLUX.2 [flex] (flux-2-flex) · available · 10 credits for the first megapixel, +10 credits per extra megapixel (input images count); each reference image counts as 5 megapixels (+50 credits) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: up to 4 MP; 1024×768 by default - output_format (enum) default png: png, jpeg - seed (integer): Reproducible generations - guidance_scale (number) default 3.5: 1.5–10 - num_inference_steps (integer) default 28: 2–50 - image_urls (string[]): Edit mode: up to 4 reference images; each counts as 5 megapixels - safety_tolerance (enum): 1–5 - FLUX.2 [klein] (flux-2-klein) · available · 1.2 credits per megapixel; edits: 3 credits for the first megapixel, +2.2 credits per extra megapixel (input images count); edits: each reference image counts as 1 megapixel (+2.2 credits) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: up to 4 MP; 1024×768 by default - num_images (integer) default 1: 1–4 on fal; create_image asks for 1–4 - num_inference_steps (integer) default 4: 4–8 - output_format (enum) default png: png, jpeg, webp - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: up to 4 images, each resized to 1 MP and billed as input - FLUX 3 (flux-3) · available · 12 credits per second (720p draft); 34 credits per second (720p); 58 credits per second (1080p); 5 s (720p) = 170 credits · used by create_video Parameters: - prompt (string): What happens in the video - duration (integer) default 5: Seconds: 5–20 s; other values are fitted - resolution (string) default 720p: 720p or 1080p; other values are fitted - aspect_ratio (string): 21:9, 2:1, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url ### Qwen Image (Alibaba) Qwen Image 3: long, detailed prompts of up to 5,000 characters. - Qwen Image 3 (qwen-image-3) · available · 8 credits per image (1K); 15 credits per image (2K) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 5,000 characters) - image_size (object): {width, height}: 512×512 to 2048×2048 pixels in total; 1024×1024 by default - num_images (integer) default 1: 1–6 on fal; create_image asks for 1–4 - negative_prompt (string): What to avoid (up to 500 characters) - enable_prompt_expansion (boolean) default true: Rewrites the prompt with an LLM first - output_format (enum) default png: png, jpeg, webp - seed (integer): Reproducible generations - image_urls (string[]): Edit mode: 1–3 reference images, named image 1, 2, 3 in the prompt ### Ideogram (Ideogram) Ideogram 4: typography and logos, plus one-step background removal. - Ideogram 4 (ideogram-4) · available · 1.5 credits per megapixel (turbo); 3 credits per megapixel (balanced); 5 credits per megapixel (quality) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images - image_size (object): {width, height}: up to 4096 px a side; 1024×1024 by default - rendering_speed (enum) default BALANCED: TURBO, BALANCED, QUALITY - num_images (integer) default 1: 1–4 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg - seed (integer): Reproducible generations - image_url (string): Edit mode: the one source image - strength (number) default 0.8: 0–1: how far to move from the source image - acceleration (enum) default none: none, low, regular, high - expansion_model (enum): None, Medium, Large - Ideogram Remove Background (ideogram-remove-background) · available · 2 credits per image · used by remove_background Parameters: - image_url (string): The image URL whose background needs to be removed. The foreground subject is preserved against a transparent background. JPEG, PNG and WebP formats are supported (maximum file size 10MB). - sync_mode (boolean) default false: If `True`, the media will be returned as a data URI and the output data won't be available in the request history. ### Recraft (Recraft) Recraft V4.1: brand illustrations and editable SVG vectors, plus cheap upscaling. - Recraft V4.1 (recraft-4.1) · available · 7 credits per image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 10,000 characters) - image_size (object): {width, height}: custom sizes up to 4096 px a side; 1024×1024 by default - colors (object[]): Preferred colors as [{r, g, b}] - background_color (object): A preferred background as {r, g, b} - Recraft V4.1 Vector (recraft-4.1-vector) · available · 16 credits per image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 10,000 characters) - image_size (object): {width, height}: custom sizes up to 4096 px a side; 1024×1024 by default - colors (object[]): Preferred colors as [{r, g, b}] - background_color (object): A preferred background as {r, g, b} - Recraft Crisp Upscale (recraft-crisp-upscale) · available · 1 credit per image · used by upscale_image Parameters: - image_url (string): The URL of the image to be upscaled. Must be in PNG format. - sync_mode (boolean) default false: If `True`, the media will be returned as a data URI and the output data won't be available in the request history. - enable_safety_checker (boolean) default false: If set to true, the safety checker will be enabled. Disabling it requires account authorization; unauthorized requests are always checked. ### MAI-Image (Microsoft) MAI-Image-2.5: realistic scenes and portraits. - MAI-Image-2.5 (mai-image-2.5) · available · 11 credits per image; +2.2 credits per reference image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 5,000 characters) - aspect_ratio (enum) default auto: auto, 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3 - num_images (integer) default 1: 1–4 on fal; create_image asks for 1–4 - output_format (enum) default png: png, jpeg, webp - image_urls (string[]): Edit mode: exactly 1 image; it adds to the price ### Krea (Krea) Krea 2: stylized images guided by style reference images. - Krea 2 (krea-2) · available · 12 credits per image (standard); 13 credits per image (with style references) · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 5,000 characters) - aspect_ratio (enum) default 1:1: 1:1, 4:3, 3:2, 16:9, 2.35:1, 4:5, 2:3, 9:16 - creativity (enum) default medium: raw, low, medium, high - styles (object[]): Up to 10 Krea styles as [{id, strength}] - moodboards (object[]): One Krea moodboard as [{id, strength}] - seed (integer): Reproducible generations - image_style_references (object[]): Style reference images ### Luma (Luma AI) Luma Uni-1 & Ray 3.2: concept-art images and cinematic silent shots guided by keyframes. - Luma Uni-1 (luma-uni-1) · available · 9 credits per image; +0.6 credits per reference image · used by create_image Parameters: - prompt (string): What to generate, or how to edit the reference images (up to 6,000 characters) - aspect_ratio (enum): 3:1, 2:1, 16:9, 3:2, 1:1, 2:3, 9:16, 1:2, 1:3 - style (enum) default auto: auto, manga - output_format (enum) default png: png, jpeg - image_url (string): Edit mode: the source image - reference_image_urls (string[]): Up to 9 reference images - Luma Ray 3.2 (luma-ray-3.2) · available · 100 credits per clip (5 s 540p); 200 credits per clip (5 s 720p); 400 credits per clip (5 s 1080p); 200 credits per clip (10 s 540p); 400 credits per clip (10 s 720p); 800 credits per clip (10 s 1080p); from an image: 30 credits per clip (5 s 540p); from an image: 60 credits per clip (5 s 720p); from an image: 240 credits per clip (5 s 1080p) · used by create_video Parameters: - prompt (string): What happens in the video; up to 6,000 characters - duration (integer) default 5: Seconds: 5 or 10 s; other values are fitted - resolution (string) default 720p: 540p, 720p or 1080p; other values are fitted - aspect_ratio (string): 3:4, 4:3, 1:1, 9:16, 16:9 or 21:9; other ratios are fitted - audio (boolean): Makes silent video - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - loop (boolean) default false: In extras: a seamless loop (5 s, no last frame) ### Gemini Omni (Google) Gemini Omni Flash: text-to-video with native sound, from 360p drafts to 4K. - Gemini Omni Flash 1.1 (gemini-omni-flash) · available · 6 credits per second (360p); 20 credits per second (720p); 30 credits per second (1080p); 60 credits per second (4K); 5 s (720p) = 100 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 20,000 characters - duration (integer) default 5: Seconds: 3–10 s; other values are fitted - resolution (string) default 720p: 360p, 720p, 1080p or 4k; other values are fitted - aspect_ratio (string): 16:9 or 9:16; other ratios are fitted - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 10 images of subjects or products to keep consistent; not together with image_url or end_image_url ### Veo (Google) Veo 3.1: clips with native sound, from cheap Lite drafts to 4K with Fast and standard. - Veo 3.1 (veo-3.1) · available · 40 credits per second (720p silent); 80 credits per second (720p with audio); 40 credits per second (1080p silent); 80 credits per second (1080p with audio); 80 credits per second (4K silent); 120 credits per second (4K with audio); 5 s (720p with audio) = 400 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 20,000 characters - duration (integer) default 6: Seconds: 4, 6 or 8 s; other values are fitted - resolution (string) default 720p: 720p, 1080p or 4k; other values are fitted - aspect_ratio (string): 16:9 or 9:16; other ratios are fitted - audio (boolean): Sound on (default) or off; off costs less - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 3 images of subjects or products to keep consistent; not together with image_url or end_image_url - seed (integer): Same seed and prompt give a similar video - negative_prompt (string): In extras: what to keep out of the video - auto_fix (boolean): In extras: let Veo rewrite a prompt its policy refuses - Veo 3.1 Fast (veo-3.1-fast) · available · 20 credits per second (720p silent); 30 credits per second (720p with audio); 20 credits per second (1080p silent); 30 credits per second (1080p with audio); 60 credits per second (4K silent); 70 credits per second (4K with audio); 5 s (720p with audio) = 150 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 20,000 characters - duration (integer) default 6: Seconds: 4, 6 or 8 s; other values are fitted - resolution (string) default 720p: 720p, 1080p or 4k; other values are fitted - aspect_ratio (string): 16:9 or 9:16; other ratios are fitted - audio (boolean): Sound on (default) or off; off costs less - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - seed (integer): Same seed and prompt give a similar video - negative_prompt (string): In extras: what to keep out of the video - auto_fix (boolean): In extras: let Veo rewrite a prompt its policy refuses - Veo 3.1 Lite (veo-3.1-lite) · available · 6 credits per second (720p silent); 10 credits per second (720p with audio); 10 credits per second (1080p silent); 16 credits per second (1080p with audio); 5 s (720p with audio) = 50 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 20,000 characters - duration (integer) default 6: Seconds: 4, 6 or 8 s; other values are fitted - resolution (string) default 720p: 720p or 1080p; other values are fitted - aspect_ratio (string): 16:9 or 9:16; other ratios are fitted - audio (boolean): Sound on (default) or off; off costs less - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - seed (integer): Same seed and prompt give a similar video - negative_prompt (string): In extras: what to keep out of the video - auto_fix (boolean): In extras: let Veo rewrite a prompt its policy refuses ### Wan (Alibaba) Wan 3.0: single shots up to 30 seconds at native 1080p, from text, images or a web page. - Wan 3.0 (wan-3) · available · 10 credits per second (480p); 20 credits per second (720p); 40 credits per second (1080p); 5 s (720p) = 100 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 20,000 characters - duration (integer) default 5: Seconds: 2–30 s; other values are fitted - resolution (string) default 720p: 480p, 720p or 1080p; other values are fitted - aspect_ratio (string): 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 10 images of subjects or products to keep consistent; not together with image_url or end_image_url - seed (integer): Same seed and prompt give a similar video - enable_prompt_expansion (boolean) default true: In extras: rewrite the prompt first (off saves 20–60 s) - enable_thinking (boolean) default false: In extras: reason before generating; needed for file_url and web_url - file_url (string): In extras: with reference_image_urls and enable_thinking, a document to base the video on - web_url (string): In extras: with reference_image_urls and enable_thinking, a public web page to base the video on ### HappyHorse (Alibaba) HappyHorse 1.1: talking characters with multilingual lip-sync. - HappyHorse 1.1 (happyhorse-1.1) · available · 28 credits per second (720p); 36 credits per second (1080p); 5 s (720p) = 140 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 2,500 characters - duration (integer) default 5: Seconds: 3–15 s; other values are fitted - resolution (string) default 720p: 720p or 1080p; other values are fitted - aspect_ratio (string): 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4 or 4:5; other ratios are fitted; with image_url the video follows the image - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - reference_image_urls (string[]): Up to 9 images of subjects or products to keep consistent; not together with image_url or end_image_url - seed (integer): Same seed and prompt give a similar video ### MiniMax (MiniMax) MiniMax H3: fast video with sound (H3), expressive voices (Speech 2.8) and songs (Music 2.6). - MiniMax H3 Max (minimax-h3-max) · available · 10 credits per second (480p); 16 credits per second (768p); 32 credits per second (1080p); 5 s (768p) = 80 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 50,000 characters - duration (integer) default 5: Seconds: 5–15 s; other values are fitted - resolution (string) default 768p: 480p, 768p or 1080p; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - end_image_url (string): Last frame; may come without image_url - seed (integer): Same seed and prompt give a similar video - prompt_expansion_mode (string) default balanced: In extras: disabled, balanced or quality (quality spends up to about 30 s) - target_audio_url (string): In extras: an audio clip (2 s or longer, up to 15 MB) to use as the soundtrack - MiniMax H3 Max Turbo (minimax-h3-max-turbo) · available · 5 credits per second (480p); 8 credits per second (768p); 16 credits per second (1080p); 5 s (768p) = 40 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 50,000 characters - duration (integer) default 5: Seconds: 5–15 s; other values are fitted - resolution (string) default 768p: 480p, 768p or 1080p; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - end_image_url (string): Last frame; may come without image_url - seed (integer): Same seed and prompt give a similar video - prompt_expansion_mode (string) default balanced: In extras: disabled, balanced or quality (quality spends up to about 30 s) - target_audio_url (string): In extras: an audio clip (2 s or longer, up to 15 MB) to use as the soundtrack - MiniMax H3 (minimax-h3) · available · 10 credits per second (480p); 12 credits per second (768p); 26 credits per second (2K); 32 credits per second (4K); 5 s (768p) = 60 credits; +16 credits per reference image after the first 5 · used by create_video Parameters: - prompt (string): What happens in the video; up to 50,000 characters - duration (integer) default 5: Seconds: 5–15 s; other values are fitted - resolution (string) default 768p: 480p, 768p, 2k or 4k; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - end_image_url (string): Last frame; may come without image_url - reference_image_urls (string[]): Up to 9 images of subjects or products to keep consistent; not together with image_url or end_image_url - seed (integer): Same seed and prompt give a similar video - prompt_expansion_mode (string) default balanced: In extras: disabled, balanced or quality (quality spends up to about 30 s) - target_audio_url (string): In extras: an audio clip (2 s or longer, up to 15 MB) to use as the soundtrack - MiniMax Speech 2.8 HD (minimax-speech-2.8-hd) · available · 20 credits per 1,000 characters · used by create_voiceover Parameters: - prompt (string): Text to convert to speech. Use `<#x#>` for pauses (x = 0.01-99.99 seconds). - voice_setting (object) default {"pitch":0,"speed":1,"voice_id":"Wise_Woman","vol":1,"english_normalization":false}: Voice configuration settings. Fields: pitch, emotion (happy/sad/angry/fearful/disgusted/surprised/neutral), speed, voice_id, vol, english_normalization. voice_id examples: Wise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl. - audio_setting (object): Audio configuration settings. Fields: format (mp3/pcm/flac), bitrate (32000/64000/128000/256000), channel (1/2), sample_rate (8000/16000/22050/24000/32000/44100). - language_boost (enum): Enhance recognition of specified languages and dialects One of: Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Slovak, Swedish, Croatian, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Afrikaans, auto. - output_format (enum) default hex: Format of the output content (non-streaming only) One of: url, hex. - pronunciation_dict (object): Custom pronunciation dictionary for text replacement. Fields: tone_list. - normalization_setting (object): Loudness normalization settings for the audio. Fields: target_loudness, target_peak, enabled, target_range. - voice_modify (object): Voice modification settings to adjust pitch, intensity, and timbre. Fields: pitch, timbre, intensity. - MiniMax Speech 2.8 Turbo (minimax-speech-2.8-turbo) · available · 12 credits per 1,000 characters · used by create_voiceover Parameters: - prompt (string): Text to convert to speech. Use `<#x#>` for pauses (x = 0.01-99.99 seconds). - voice_setting (object) default {"pitch":0,"speed":1,"voice_id":"Wise_Woman","vol":1,"english_normalization":false}: Voice configuration settings. Fields: pitch, emotion (happy/sad/angry/fearful/disgusted/surprised/neutral), speed, voice_id, vol, english_normalization. voice_id examples: Wise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl. - audio_setting (object): Audio configuration settings. Fields: format (mp3/pcm/flac), bitrate (32000/64000/128000/256000), channel (1/2), sample_rate (8000/16000/22050/24000/32000/44100). - language_boost (enum): Enhance recognition of specified languages and dialects One of: Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Slovak, Swedish, Croatian, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Afrikaans, auto. - output_format (enum) default hex: Format of the output content (non-streaming only) One of: url, hex. - pronunciation_dict (object): Custom pronunciation dictionary for text replacement. Fields: tone_list. - normalization_setting (object): Loudness normalization settings for the audio. Fields: target_loudness, target_peak, enabled, target_range. - voice_modify (object): Voice modification settings to adjust pitch, intensity, and timbre. Fields: pitch, timbre, intensity. - MiniMax Music 2.6 (minimax-music-2.6) · available · 30 credits per track · used by create_music Parameters: - prompt (string): A description of the music style, mood, genre, and scenario. 10-2000 characters. - lyrics (string): Lyrics of the song. Use \n to separate lines. Supports structure tags: [Intro], [Verse], [Pre Chorus], [Chorus], [Post Chorus], [Hook], [Bridge], [Interlude], [Transition], [Build Up], [Break],… - lyrics_optimizer (boolean) default false: When true and lyrics is empty, auto-generates lyrics from the prompt. - is_instrumental (boolean) default false: When true, generates vocal-free instrumental music. - audio_setting (object): Audio configuration settings. Fields: sample_rate (16000/24000/32000/44100), format (mp3/wav/pcm), bitrate (32000/64000/128000/256000). ### Seedance (ByteDance) Seedance 2.5: multi-shot videos up to 30 seconds guided by dozens of references. - Seedance 2.5 (seedance-2.5) · available · about 44.14 credits per second (480p, 880×500); about 94.63 credits per second (720p, 1280×737); about 232.92 credits per second (1080p, 1920×1106); 5 s (720p, 1280×737) = 474 credits · used by create_video Parameters: - prompt (string): What happens in the video - duration (integer) default 5: Seconds: 4–30 s; other values are fitted - resolution (string) default 720p: 480p, 720p or 1080p; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 30 images of subjects or products to keep consistent; not together with image_url or end_image_url - seed (integer): Same seed and prompt give a similar video - codec (string) default auto: In extras: auto, H264 or H265 - bitrate_mode (string) default standard: In extras: standard or high (a larger file) - Seedance 2.0 (seedance-2) · available · about 28.88 credits per second (480p, 880×500); about 61.91 credits per second (720p, 1280×737); about 139.36 credits per second (1080p, 1920×1106); about 311.04 credits per second (4K, 3840×2160); 5 s (720p, 1280×737) = 310 credits · used by create_video Parameters: - prompt (string): What happens in the video - duration (integer) default 5: Seconds: 4–15 s; other values are fitted - resolution (string) default 720p: 480p, 720p, 1080p or 4k; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 9 images of subjects or products to keep consistent; not together with image_url or end_image_url - codec (string) default auto: In extras: auto, H264 or H265 - bitrate_mode (string) default standard: In extras: standard or high (a larger file) - Seedance 2.0 Fast (seedance-2-fast) · available · about 23.1 credits per second (480p, 880×500); about 49.53 credits per second (720p, 1280×737); 5 s (720p, 1280×737) = 248 credits · used by create_video Parameters: - prompt (string): What happens in the video - duration (integer) default 5: Seconds: 4–15 s; other values are fitted - resolution (string) default 720p: 480p or 720p; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 9 images of subjects or products to keep consistent; not together with image_url or end_image_url - codec (string) default auto: In extras: auto, H264 or H265 - bitrate_mode (string) default standard: In extras: standard or high (a larger file) - Seedance 2.0 Mini (seedance-2-mini) · available · about 14.44 credits per second (480p, 880×500); about 30.95 credits per second (720p, 1280×737); 5 s (720p, 1280×737) = 155 credits · used by create_video Parameters: - prompt (string): What happens in the video - duration (integer) default 5: Seconds: 4–15 s; other values are fitted - resolution (string) default 720p: 480p or 720p; other values are fitted - aspect_ratio (string): 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 9 images of subjects or products to keep consistent; not together with image_url or end_image_url - codec (string) default auto: In extras: auto, H264 or H265 ### Kling (Kuaishou) Kling 3.0: cinematic motion and multi-shot storyboards. - Kling 3.0 Pro (kling-3-pro) · available · 22.4 credits per second (silent); 33.6 credits per second (with audio); 39.2 credits per second (with audio and voice control); 5 s (with audio) = 168 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 2,500 characters - duration (integer) default 5: Seconds: 3–15 s; other values are fitted - aspect_ratio (string): 16:9, 9:16 or 1:1; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off; off costs less - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - negative_prompt (string): In extras: what to keep out of the video - cfg_scale (number) default 0.5: In extras: 0–1, how closely to follow the prompt - Kling 3.0 Turbo (kling-3-turbo) · available · 28 credits per second; 5 s (1080p) = 140 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 3,072 characters - duration (integer) default 5: Seconds: 3–15 s; other values are fitted - aspect_ratio (string): 16:9, 9:16 or 1:1; other ratios are fitted; with image_url the video follows the image - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - Kling O3 (kling-o3) · available · 22.4 credits per second (silent); 28 credits per second (with audio); 5 s (with audio) = 140 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 2,500 characters - duration (integer) default 5: Seconds: 3–15 s; other values are fitted - aspect_ratio (string): 16:9, 9:16 or 1:1; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off; off costs less - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 4 images of subjects or products to keep consistent; can come with image_url and end_image_url ### LTX (Lightricks) LTX-2.5: open-weights video up to 20 seconds and 4K. - LTX-2.5 (ltx-2.5) · available · 18 credits per second (720p); 26 credits per second (1080p); 38 credits per second (1440p); 60 credits per second (4K); 5 s (720p) = 90 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 5,000 characters - duration (integer) default 6: Seconds: 6, 8, 10, 12, 14, 16, 18 or 20 s; other values are fitted - resolution (string) default 720p: 720p, 1080p, 1440p or 2160p; other values are fitted - aspect_ratio (string): 16:9 or 9:16; other ratios are fitted - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - camera_motion (string): In extras: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static or focus_shift - fps (integer) default 25: In extras: 24, 25, 48 or 50; 48 and 50 allow at most 10 s ### PixVerse (PixVerse) PixVerse V6: cheap social clips with optional sound. - PixVerse V6 (pixverse-6) · available · 12 credits per second (720p with audio); from 5 credits (360p silent) to 23 credits (1080p with audio); 5 s (720p with audio) = 60 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 2,048 bytes (UTF-8) - duration (integer) default 5: Seconds: 1–15 s; other values are fitted - resolution (string) default 720p: 360p, 540p, 720p or 1080p; other values are fitted - aspect_ratio (string): 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2 or 21:9; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off; off costs less - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - seed (integer): Same seed and prompt give a similar video - negative_prompt (string): In extras: what to keep out of the video - style (string): In extras: anime, 3d_animation, clay, comic or cyberpunk - thinking_type (string): In extras: enabled, disabled or auto prompt optimization - generate_multi_clip_switch (boolean) default false: In extras: several clips with camera changes in one video ### Vidu (Shengshu) Vidu Q3: clips up to 16 seconds with start and end frames. - Vidu Q3 (vidu-q3) · available · 14 credits per second (360p); 14 credits per second (540p); 30.8 credits per second (720p); 30.8 credits per second (1080p); 5 s (720p) = 154 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 2,000 characters - duration (integer) default 5: Seconds: 1–16 s; other values are fitted - resolution (string) default 720p: 360p, 540p, 720p or 1080p; other values are fitted - aspect_ratio (string): 16:9, 9:16, 4:3, 3:4 or 1:1; other ratios are fitted; with image_url the video follows the image - audio (boolean): Sound on (default) or off - image_url (string): First frame to animate - end_image_url (string): Last frame; needs image_url - reference_image_urls (string[]): Up to 4 images of subjects or products to keep consistent; not together with image_url or end_image_url - seed (integer): Same seed and prompt give a similar video ### Creatify (Creatify) Creatify Boreal: product, UGC and presenter ads written as a script. - Creatify Boreal (creatify-boreal) · available · 2 credits per second (720p); 6 credits per second (1080p); 24 credits per second (2K); 5 s (720p) = 10 credits · used by create_video Parameters: - prompt (string): What happens in the video; up to 5,000 characters - duration (integer) default 5: Seconds: 1–20 s; other values are fitted - resolution (string) default 720p: 720p, 1080p or 2k; other values are fitted - aspect_ratio (string): 16:9, 9:16, 1:1, 4:3 or 3:4; other ratios are fitted - audio (boolean): Always makes sound; audio: false changes nothing - image_url (string): First frame to animate - negative_prompt (string): In extras: what to keep out of the video - audio_url (string): In extras: narration or a soundtrack to keep in the video - manifest_disclosure (boolean) default false: In extras: add a visible "AI-generated" label ### Gemini TTS (Google) Gemini 3.8 TTS: natural voices directed with plain-English style notes. - Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) · available · 9 credits per 1,000 characters · used by create_voiceover Parameters: - prompt (string): Verbatim text for single-speaker speech. Put delivery directions in style_instructions; inline vocal events may use or . - style_instructions (string): Delivery style, separate from the spoken transcript. Applies to all turns unless overridden. - voice (enum) default Kore: Prebuilt voice for single-speaker speech. One of: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, Zubenelgenubi. - speakers (array of object): Exactly two distinct speaker aliases and their prebuilt voices for dialogue. - turns (array of object): Ordered dialogue turns, each identifying a configured speaker. - Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) · available · 6 credits per 1,000 characters · used by create_voiceover Parameters: - prompt (string): Verbatim text for single-speaker speech. Put delivery directions in style_instructions; inline vocal events may use or . - style_instructions (string): Delivery style, separate from the spoken transcript. Applies to all turns unless overridden. - voice (enum) default Kore: Prebuilt voice for single-speaker speech. One of: Achernar, Achird, Algenib, Algieba, Alnilam, Aoede, Autonoe, Callirrhoe, Charon, Despina, Enceladus, Erinome, Fenrir, Gacrux, Iapetus, Kore, Laomedeia, Leda, Orus, Pulcherrima, Puck, Rasalgethi, Sadachbia, Sadaltager, Schedar, Sulafat, Umbriel, Vindemiatrix, Zephyr, Zubenelgenubi. - speakers (array of object): Exactly two distinct speaker aliases and their prebuilt voices for dialogue. - turns (array of object): Ordered dialogue turns, each identifying a configured speaker. ### ElevenLabs (ElevenLabs) ElevenLabs v3: the best-known voices with emotion tags and word timings, plus music, transcription and voice cleanup. - ElevenLabs v3 (elevenlabs-v3) · available · 20 credits per 1,000 characters · used by create_voiceover Parameters: - text (string): The text to convert to speech - voice (string) default Rachel: The voice to use for speech generation. Examples: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill. - stability (number) default 0.5: Voice stability (0-1) - timestamps (boolean) default false: Whether to return timestamps for each word in the generated speech - language_code (string): Language code (ISO 639-1) used to enforce a language for the model. - apply_text_normalization (enum) default auto: This parameter controls text normalization with three modes: 'auto', 'on', and 'off'. One of: auto, on, off. - ElevenLabs Multilingual v2 (elevenlabs-multilingual-v2) · available · 20 credits per 1,000 characters · used by create_voiceover Parameters: - text (string): The text to convert to speech - voice (string) default Rachel: The voice to use for speech generation. Examples: Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill. - stability (number) default 0.5: Voice stability (0-1) - similarity_boost (number) default 0.75: Similarity boost (0-1) - style (number) default 0: Style exaggeration (0-1) - speed (number) default 1: Speech speed (0.7-1.2). Values below 1.0 slow down the speech, above 1.0 speed it up. Extreme values may affect quality. - timestamps (boolean) default false: Whether to return timestamps for each word in the generated speech - previous_text (string): The text that came before the text of the current request. Can be used to improve the speech's continuity when concatenating together multiple generations or to influence the speech's continuity in… - next_text (string): The text that comes after the text of the current request. Can be used to improve the speech's continuity when concatenating together multiple generations or to influence the speech's continuity in… - language_code (string): Language code (ISO 639-1) used to enforce a language for the model. An error will be returned if language code is not supported by the model. - apply_text_normalization (enum) default auto: This parameter controls text normalization with three modes: 'auto', 'on', and 'off'. One of: auto, on, off. - ElevenLabs Music v2.5 (elevenlabs-music-2.5) · available · 120 credits per started minute · used by create_music Parameters: - prompt (string): The text prompt describing the music to generate - composition_plan (object): The chunk-based composition plan for the music. Fields: chunks. - music_length_ms (integer): The length of the song to generate in milliseconds. Used only in conjunction with prompt. Must be between 3000ms and 600000ms. - force_instrumental (boolean) default false: If true, guarantees that the generated song will be instrumental. If false, the song may or may not be instrumental depending on the prompt. Can only be used with prompt. - seed (integer): Random seed to initialize the music generation process. Can only be used with composition_plan. - output_format (enum) default mp3_48000_192: Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. One of: mp3_22050_32, mp3_24000_48, mp3_44100_32, mp3_44100_64, mp3_44100_96, mp3_44100_128, mp3_44100_192, mp3_48000_128, mp3_48000_192, mp3_48000_240, mp3_48000_320, pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_32000, pcm_44100, pcm_48000, ulaw_8000, alaw_8000, opus_48000_32, opus_48000_64, opus_48000_96, opus_48000_128, opus_48000_192. - ElevenLabs Scribe v2 (scribe-v2) · available · 1.6 credits per started minute · used by transcribe Parameters: - audio_url (string): URL of the audio file to transcribe - language_code (string): Language code of the audio - tag_audio_events (boolean) default true: Tag audio events like laughter, applause, etc. - diarize (boolean) default true: Whether to annotate who is speaking - keyterms (array of string): Words or sentences to bias the model towards transcribing. Up to 100 keyterms, max 50 characters each. Adds 30% premium over base transcription price. - ElevenLabs Audio Isolation (elevenlabs-audio-isolation) · available · 20 credits per started minute · used by enhance_audio Parameters: - audio_url (string): URL of the audio file to isolate voice from - video_url (string): Video file to use for audio isolation. Either `audio_url` or `video_url` must be provided. ### xAI TTS (xAI) xAI TTS: very cheap voices with laughs, whispers and pauses. - xAI TTS (xai-tts) · available · 3 credits per 1,000 characters · used by create_voiceover Parameters: - text (string): The text to convert to speech. Maximum 15,000 characters. Supports speech tags for expressive delivery: inline tags like [laugh], [pause], [sigh] and wrapping tags like text,… - voice (enum) default eve: Built-in xAI voice to use for synthesis. One of: carina, zagan, helix, orion, luna, iris, altair, zenith, perseus, helios, lux, kepler, rigel, cosmo, celeste, ursa, sirius, lumen, castor, naksh, atlas, aurora, liora, ara, eve, leo, rex, sal. - language (enum) default auto: BCP-47 language code or 'auto' for automatic detection. Supported: en, zh, fr, de, hi, id, it, ja, ko, pt-BR, pt-PT, ru, es-MX, es-ES, tr, vi, bn, ar-EG, ar-SA, ar-AE. One of: auto, en, ar-EG, ar-SA, ar-AE, bn, zh, fr, de, hi, id, it, ja, ko, pt-BR, pt-PT, ru, es-MX, es-ES, tr, vi. - output_format (object): Output format configuration. Defaults to MP3 at 24 kHz / 128 kbps. Fields: sample_rate (8000/16000/22050/24000/44100/48000), codec (mp3/wav/pcm/mulaw/alaw), bit_rate (32000/64000/96000/128000/192000). ### Inworld (Inworld) Inworld TTS 1.5: low-cost voices, with 73 in English. - Inworld TTS 1.5 Max (inworld-tts-1.5) · available · 2 credits per 1,000 characters · used by create_voiceover Parameters: - text (string): The text to synthesize into speech. Maximum 2000 characters. - voice (enum) default Craig (en): The voice to use for synthesis. One of 113 values: Loretta (en), Darlene (en), Marlene (en), Hank (en), Evelyn (en), Celeste (en), Pippa (en), Tessa (en), Liam (en), Callum (en), Hamish (en), Abby (en), Graham (en), Rupert (en), Mortimer (en), Snik (en), Anjali (en), Saanvi (en), Arjun (en), Claire (en), Oliver (en), Simon (en), Elliot (en), James (en), Serena (en), Gareth (en), Vinny (en), Lauren (en), Jessica (en), Ethan (en), Tyler (en), Jason (en), Chloe (en), Veronica (en), Victoria (en), Miranda (en), Sebastian (en), Victor (en), Malcolm (en), Kayla (en), Nate (en), Jake (en), Brian (en), Amina (en), Kelsey (en), Derek (en), Grant (en), Evan (en), Alex (en), Ashley (en), Craig (en), Deborah (en), Dennis (en), Edward (en), Elizabeth (en), Hades (en), Julia (en), Pixie (en), Mark (en), Olivia (en), Priya (en), Ronald (en), Sarah (en), Shaun (en), Theodore (en), Timothy (en), Wendy (en), Dominus (en), Hana (en), Clive (en), Carter (en), Blake (en), Luna (en), Yichen (zh), Xiaoyin (zh), Xinyi (zh), Jing (zh), Erik (nl), Katrien (nl), Lennart (nl), Lore (nl), Alain (fr), Hélène (fr), Mathieu (fr), Étienne (fr), Johanna (de), Josef (de), Gianni (it), Orietta (it), Asuka (ja), Satoshi (ja), Hyunwoo (ko), Minji (ko), Seojun (ko), Yoona (ko), Szymon (pl), Wojciech (pl), Heitor (pt), Maitê (pt), Diego (es), Lupita (es), Miguel (es), Rafael (es), Svetlana (ru), Elena (ru), Dmitry (ru), Nikolai (ru), Riya (hi), Manoj (hi), Yael (he), Oren (he), Nour (ar), Omar (ar). - sample_rate_hertz (enum) default 48000: The sample rate in Hz for the output audio. One of: 8000, 16000, 24000, 32000, 40000, 48000. ### Kokoro (Hexgrad) Kokoro: quick, very cheap English drafts. - Kokoro (kokoro) · available · 4 credits per 1,000 characters · used by create_voiceover Parameters: - prompt (string): Prompt - voice (enum) default af_heart: Voice ID for the desired voice. One of: af_heart, af_alloy, af_aoede, af_bella, af_jessica, af_kore, af_nicole, af_nova, af_river, af_sarah, af_sky, am_adam, am_echo, am_eric, am_fenrir, am_liam, am_michael, am_onyx, am_puck, am_santa. - speed (number) default 1: Speed of the generated audio. Default is 1.0. ### Chatterbox (Resemble AI) Chatterbox HD: dramatic voices with adjustable intensity. - Chatterbox HD (chatterbox-hd) · available · 8 credits per 1,000 characters · used by create_voiceover Parameters: - text (string) default My name is Maximus Decimus Meridius, commander of the Armies of the North, General of the Felix Legions and loyal servant to the true emperor, Marcus Aurelius. Father to a murdered son, husband to a murdered wife. And I will have my vengeance, in this life or the next.: Text to synthesize into speech. - voice (enum): The voice to use for the TTS request. If neither voice nor audio are provided, a random voice will be used. One of: Aurora, Blade, Britney, Carl, Cliff, Richard, Rico, Siobhan, Vicky. - audio_url (string): URL to the audio sample to use as a voice prompt for zero-shot TTS voice cloning. Providing a audio sample will override the voice setting. - exaggeration (number) default 0.5: Controls emotion exaggeration. Range typically 0.25 to 2.0. - cfg (number) default 0.5: Classifier-free guidance scale (CFG) controls the conditioning factor. Range typically 0.2 to 1.0. For expressive or dramatic speech, try lower cfg values (e.g. - high_quality_audio (boolean) default false: If True, the generated audio will be upscaled to 48kHz. The generation of the audio will take longer, but the quality will be higher. If False, the generated audio will be 24kHz. - seed (integer) default 0: Useful to control the reproducibility of the generated audio. Assuming all other properties didn't change, a fixed seed should always generate the exact same audio file. Set to 0 for random seed. - temperature (number) default 0.8: Controls the randomness of generation. Range typically 0.05 to 5. ### Dia (Nari Labs) Dia: two-speaker dialogue with laughs and other nonverbal sounds. - Dia (dia) · available · 8 credits per 1,000 characters · used by create_voiceover Parameters: - text (string): The text to be converted to speech. ### Orpheus (Canopy Labs) Orpheus: open-source English narration. - Orpheus (orpheus) · available · 10 credits per 1,000 characters · used by create_voiceover Parameters: - text (string): The text to be converted to speech. You can additionally add the following emotive tags: , , , , , , , - voice (enum) default tara: Voice ID for the desired voice. One of: tara, leah, jess, leo, dan, mia, zac, zoe. - temperature (number) default 0.7: Temperature for generation (higher = more creative). - repetition_penalty (number) default 1.2: Repetition penalty (>= 1.1 required for stable generations). ### Lyria (Google) Lyria 3.5: full songs or instrumental background tracks of up to about 3 minutes. - Lyria 3.5 (lyria-3.5) · available · 20 credits per track · used by create_music Parameters: - prompt (string): The text prompt describing the music you want to generate. Include genre, mood, instrumentation, tempo, vocals, and structure for best results. - negative_prompt (string): Negative prompting is not supported by Lyria 3.5. - image_url (string): Optional image URL to use as visual inspiration for music generation. The model will create music that matches the mood and theme of the image. ### Stable Audio (Stability AI) Stable Audio 3: instrumental tracks and ambience with an exact length, trained on licensed data. - Stable Audio 3 (stable-audio-3) · available · 8 credits per track · used by create_music Parameters: - negative_prompt (string): Text description of qualities to avoid in the output. - duration (number) default 30: Duration of the generated audio in seconds. The medium model supports up to 380 seconds (~6m20s). - num_inference_steps (integer) default 8: Number of sampling steps. Post-trained (distilled) checkpoints look good with the default 8 and gain little from going higher. - guidance_scale (number) default 1: Classifier-free guidance scale. Higher values follow the prompt more strictly. Only effective on base (non-distilled) checkpoints. - seed (integer): Random seed for reproducible outputs. Omit for a random seed. - enable_prompt_expansion (boolean) default false: If True, the prompt will be expanded using an LLM for more detailed and higher quality results. - enable_safety_checker (boolean) default true: Enable NSFW content safety checking. - sync_mode (boolean) default false: If True, the audio is returned inline as a data URI and the result is not saved to the request history. - output_format (enum) default mp3: Container format for the generated audio output. One of: mp3, wav, flac, ogg, opus, m4a, aac. - bitrate (string) default 192k: Audio bitrate for compressed output formats (e.g., mp3, aac, opus). Format e.g. '192k' or '320k'. Ignored for lossless formats (wav, flac). - prompt (string): Text description of the audio to generate. ### Utility models - DeepFilterNet 3 (deepfilternet-3) · available · 12 credits per minute of audio · https://sakaira.com/tools/enhance-audio Parameters: - audio_url (string): The URL of the audio to enhance. - sync_mode (boolean) default false: If `True`, the media will be returned as a data URI and the output data won't be available in the request history. - audio_format (enum) default mp3: The format for the output audio. One of: mp3, aac, m4a, ogg, opus, flac, wav. - bitrate (string) default 192k: The bitrate of the output audio. - Bria RMBG 2.0 (bria-rmbg-2) · available · 4 credits per image · https://sakaira.com/tools/remove-background Parameters: - image_url (string): Input Image to erase from - sync_mode (boolean) default false: If `True`, the media will be returned as a data URI and the output data won't be available in the request history. - Topaz Upscale Precision (topaz-upscale) · available · 16 credits per started 24 megapixels of output · https://sakaira.com/tools/upscale-image Parameters: - image_url (string): Url of the image to be upscaled - model (enum) default Standard V2: Precision upscaling model. Standard V2 fits most photos; High Fidelity V3/V2 preserve detail in professional shots; Low Resolution V2 recovers compressed sources; CGI targets art and rendered… One of: Standard V2, High Fidelity V3, High Fidelity V2, Low Resolution V2, CGI, Text Refine, Faces. - upscale_factor (number) default 2: Factor to upscale the image by (e.g. 2.0 doubles width and height) - crop_to_fill (boolean) default false: Crop To Fill - output_format (enum) default jpeg: Output format of the upscaled image. One of: jpeg, png. - subject_detection (enum) default All: Subject detection mode for the image enhancement. One of: All, Foreground, Background. - face_enhancement (boolean) default true: Whether to apply face enhancement to the image. - face_enhancement_creativity (number) default 0: Creativity level for face enhancement. 0.0 means no creativity, 1.0 means maximum creativity. Ignored if face enhancement is disabled. - face_enhancement_strength (number) default 0.8: Strength of the face enhancement. 0.0 means no enhancement, 1.0 means maximum enhancement. Ignored if face enhancement is disabled. - sharpen (number): Sharpening level (0.0-1.0). Default varies by model. - denoise (number): Denoising level (0.0-1.0). Default varies by model. - fix_compression (number): Compression artifact removal level (0.0-1.0). Not supported by the CGI model. - strength (number): Enhancement strength (0.01-1.0). Applies to Text Refine model only. - SeedVR2 Upscaler (seedvr2-upscale) · available · 0.2 credits per megapixel · https://sakaira.com/tools/upscale-image Parameters: - image_url (string): The input image to be processed - upscale_mode (enum) default factor: The mode to use for the upscale. If 'target', the upscale factor will be calculated based on the target resolution. If 'factor', the upscale factor will be used directly. One of: target, factor. - upscale_factor (number) default 2: Upscaling factor to be used. Will multiply the dimensions with this factor when `upscale_mode` is `factor`. - target_resolution (enum) default 1080p: The target resolution to upscale to when `upscale_mode` is `target`. One of: 720p, 1080p, 1440p, 2160p. - seed (integer): The random seed used for the generation process. - noise_scale (number) default 0.1: The noise scale to use for the generation process. - output_format (enum) default jpg: The format of the output image. One of: png, jpg, webp. - sync_mode (boolean) default false: If `True`, the media will be returned as a data URI and the output data won't be available in the request history. - Clarity Upscaler (clarity-upscaler) · available · 6 credits per megapixel · https://sakaira.com/tools/upscale-image Parameters: - image_url (string): The URL of the image to upscale. - prompt (string) default masterpiece, best quality, highres: The prompt to use for generating the image. Be as descriptive as possible for best results. - upscale_factor (number) default 2: The upscale factor - negative_prompt (string) default (worst quality, low quality, normal quality:2): The negative prompt to use. Use it to address details that you don't want in the image. - creativity (number) default 0.35: The creativity of the model. The higher the creativity, the more the model will deviate from the prompt. Refers to the denoise strength of the sampling. - resemblance (number) default 0.6: The resemblance of the upscaled image to the original image. The higher the resemblance, the more the model will try to keep the original image. Refers to the strength of the ControlNet. - guidance_scale (number) default 4: The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt when looking for a related image to show you. - num_inference_steps (integer) default 18: The number of inference steps to perform. - seed (integer): The same seed and the same prompt given to the same version of Stable Diffusion will output the same image every time. - enable_safety_checker (boolean) default true: If set to false, the safety checker will be disabled. - Photo Restoration (photo-restoration) · available · 8 credits per image · https://sakaira.com/tools/restore-image Parameters: - image_url (string): Old or damaged photo URL to restore - enhance_resolution (boolean) default true: Enhance Resolution - fix_colors (boolean) default true: Fix Colors - remove_scratches (boolean) default true: Remove Scratches - aspect_ratio (object): Aspect ratio for 4K output (default: 4:3 for classic photos). Fields: ratio (1:1/16:9/9:16/4:3/3:4). - ByteDance Video Upscaler (bytedance-video-upscaler) · coming soon · 1.44 credits per second (1080p 30 fps); 2.88 credits per second (2K 30 fps); 5.76 credits per second (4K 30 fps); 2.88 credits per second (1080p 60 fps); 5.76 credits per second (2K 60 fps); 11.52 credits per second (4K 60 fps); 5 s (1080p 30 fps) = 8 credits ## Sakaira pricing Sakaira is sold by subscription: three plans, billed monthly or yearly (yearly saves 20%). A plan's credits arrive every month and expire when their month ends, a month after they arrive. A generation that fails gives its credits back, except credits from a payment that was later refunded or disputed. Start free: new accounts get 100 credits (no card), valid for 30 days. ### Plans | Plan | Credits a month | Monthly | Yearly | | --- | --- | --- | --- | | Starter | 1,800 | $19 | $182.40 | | Pro | 4,650 | $49 | $470.40 | | Studio (9,400 credits) | 9,400 | $99 | $950.40 | | Studio (18,900 credits) | 18,900 | $199 | $1,910.40 | | Studio (37,900 credits) | 37,900 | $399 | $3,830.40 | ### Credit packs For subscribers who run short in a month. A pack's credits expire 90 days after purchase. | Price | Credits | | --- | --- | | $10 | 700 | | $25 | 1,750 | | $50 | 3,500 | | $100 | 7,000 | Taxes are added at checkout. Payments are processed by Stripe. ### Files Files are kept 30 days on free accounts, 6 months while subscribed, and up to 90 days after a subscription ends, never past 6 months from when the file was made. ### What things cost | Example | Credits | | --- | --- | | 1 image, Nano Banana 2 (1K) | 16 | | 1 image, GPT Image 2.5 (1024×1024, high quality) | 13 | | 1 image, Muse Image | 2 | | 5 s of video, Gemini Omni Flash (720p, with sound) | 100 | | 6 s of video, Veo 3.1 Lite (720p, with sound) | 60 | | 5 s of video, Kling 3.0 Pro (with sound) | 168 | | 1,000 characters of voice, ElevenLabs v3 (~1 minute) | 20 | | 1,000 characters of voice, Gemini 3.8 Flash TTS | 9 | | 1 music track, Lyria 3.5 (up to ~3 minutes) | 20 | | 1 minute transcribed, ElevenLabs Scribe v2 | 2 | | Remove a background, Ideogram | 2 | ### Models Every model with its price in credits. Models marked "Coming soon" are listed but not available yet. #### Image | Model | Id | Price | Status | | --- | --- | --- | --- | | Nano Banana 2 (Google) | nano-banana-2 | 12 credits per image (0.5K); 16 credits per image (1K); 24 credits per image (2K); 32 credits per image (4K) | Available | | Nano Banana Pro (Google) | nano-banana-pro | 30 credits per image (1K); 30 credits per image (2K); 60 credits per image (4K) | Available | | Nano Banana 2 Lite (Google) | nano-banana-2-lite | 9 credits per image; +0.1 credits per reference image | Available | | GPT Image 2.5 Sunburst (OpenAI) | gpt-image-2.5-sunburst | 2 credits per image (1024×1024 low); 4 credits per image (1024×1024 medium); 13 credits per image (1024×1024 high); 3 credits per image (3840×2160 low); 7 credits per image (3840×2160 medium); 25 credits per image (3840×2160 high); +7.2 credits per reference image | Available | | GPT Image 2.5 Flare (OpenAI) | gpt-image-2.5-flare | 2 credits per image (1024×1024 low); 4 credits per image (1024×1024 medium); 13 credits per image (1024×1024 high); 3 credits per image (3840×2160 low); 7 credits per image (3840×2160 medium); 25 credits per image (3840×2160 high); +7.2 credits per reference image | Available | | Muse Image (Meta) | muse-image | 2 credits per image | Available | | Grok Imagine Image 2.0 (xAI) | grok-imagine-image-2 | 8 credits per image (1K low); 12 credits per image (1K medium); 12 credits per image (2K low); 16 credits per image (2K medium); +2 credits per reference image | Available | | Seedream 5.0 Pro (ByteDance) | seedream-5-pro | 14 credits per image (up to 1536×1536); 27 credits per image (up to 2048×2048); +0.9 credits per reference image after the first | Available | | Seedream 5.0 Lite (ByteDance) | seedream-5-lite | 7 credits per image | Available | | FLUX.2 [pro] (Black Forest Labs) | flux-2-pro | 6 credits for the first megapixel, +3 credits per extra megapixel (input images count); each reference image counts as 5 megapixels (+15 credits) | Available | | FLUX.2 [max] (Black Forest Labs) | flux-2-max | 14 credits for the first megapixel, +6 credits per extra megapixel (input images count); each reference image counts as 5 megapixels (+30 credits) | Available | | FLUX.2 [flex] (Black Forest Labs) | flux-2-flex | 10 credits for the first megapixel, +10 credits per extra megapixel (input images count); each reference image counts as 5 megapixels (+50 credits) | Available | | FLUX.2 [klein] (Black Forest Labs) | flux-2-klein | 1.2 credits per megapixel; edits: 3 credits for the first megapixel, +2.2 credits per extra megapixel (input images count); edits: each reference image counts as 1 megapixel (+2.2 credits) | Available | | Qwen Image 3 (Alibaba) | qwen-image-3 | 8 credits per image (1K); 15 credits per image (2K) | Available | | Ideogram 4 (Ideogram) | ideogram-4 | 1.5 credits per megapixel (turbo); 3 credits per megapixel (balanced); 5 credits per megapixel (quality) | Available | | Recraft V4.1 (Recraft) | recraft-4.1 | 7 credits per image | Available | | Recraft V4.1 Vector (Recraft) | recraft-4.1-vector | 16 credits per image | Available | | MAI-Image-2.5 (Microsoft) | mai-image-2.5 | 11 credits per image; +2.2 credits per reference image | Available | | Krea 2 (Krea) | krea-2 | 12 credits per image (standard); 13 credits per image (with style references) | Available | | Luma Uni-1 (Luma AI) | luma-uni-1 | 9 credits per image; +0.6 credits per reference image | Available | #### Video | Model | Id | Price | Status | | --- | --- | --- | --- | | Gemini Omni Flash 1.1 (Google) | gemini-omni-flash | 6 credits per second (360p); 20 credits per second (720p); 30 credits per second (1080p); 60 credits per second (4K); 5 s (720p) = 100 credits | Available | | Veo 3.1 (Google) | veo-3.1 | 40 credits per second (720p silent); 80 credits per second (720p with audio); 40 credits per second (1080p silent); 80 credits per second (1080p with audio); 80 credits per second (4K silent); 120 credits per second (4K with audio); 5 s (720p with audio) = 400 credits | Available | | Veo 3.1 Fast (Google) | veo-3.1-fast | 20 credits per second (720p silent); 30 credits per second (720p with audio); 20 credits per second (1080p silent); 30 credits per second (1080p with audio); 60 credits per second (4K silent); 70 credits per second (4K with audio); 5 s (720p with audio) = 150 credits | Available | | Veo 3.1 Lite (Google) | veo-3.1-lite | 6 credits per second (720p silent); 10 credits per second (720p with audio); 10 credits per second (1080p silent); 16 credits per second (1080p with audio); 5 s (720p with audio) = 50 credits | Available | | Wan 3.0 (Alibaba) | wan-3 | 10 credits per second (480p); 20 credits per second (720p); 40 credits per second (1080p); 5 s (720p) = 100 credits | Available | | HappyHorse 1.1 (Alibaba) | happyhorse-1.1 | 28 credits per second (720p); 36 credits per second (1080p); 5 s (720p) = 140 credits | Available | | MiniMax H3 Max (MiniMax) | minimax-h3-max | 10 credits per second (480p); 16 credits per second (768p); 32 credits per second (1080p); 5 s (768p) = 80 credits | Available | | MiniMax H3 Max Turbo (MiniMax) | minimax-h3-max-turbo | 5 credits per second (480p); 8 credits per second (768p); 16 credits per second (1080p); 5 s (768p) = 40 credits | Available | | MiniMax H3 (MiniMax) | minimax-h3 | 10 credits per second (480p); 12 credits per second (768p); 26 credits per second (2K); 32 credits per second (4K); 5 s (768p) = 60 credits; +16 credits per reference image after the first 5 | Available | | Seedance 2.5 (ByteDance) | seedance-2.5 | about 44.14 credits per second (480p, 880×500); about 94.63 credits per second (720p, 1280×737); about 232.92 credits per second (1080p, 1920×1106); 5 s (720p, 1280×737) = 474 credits | Available | | Seedance 2.0 (ByteDance) | seedance-2 | about 28.88 credits per second (480p, 880×500); about 61.91 credits per second (720p, 1280×737); about 139.36 credits per second (1080p, 1920×1106); about 311.04 credits per second (4K, 3840×2160); 5 s (720p, 1280×737) = 310 credits | Available | | Seedance 2.0 Fast (ByteDance) | seedance-2-fast | about 23.1 credits per second (480p, 880×500); about 49.53 credits per second (720p, 1280×737); 5 s (720p, 1280×737) = 248 credits | Available | | Seedance 2.0 Mini (ByteDance) | seedance-2-mini | about 14.44 credits per second (480p, 880×500); about 30.95 credits per second (720p, 1280×737); 5 s (720p, 1280×737) = 155 credits | Available | | Kling 3.0 Pro (Kuaishou) | kling-3-pro | 22.4 credits per second (silent); 33.6 credits per second (with audio); 39.2 credits per second (with audio and voice control); 5 s (with audio) = 168 credits | Available | | Kling 3.0 Turbo (Kuaishou) | kling-3-turbo | 28 credits per second; 5 s (1080p) = 140 credits | Available | | Kling O3 (Kuaishou) | kling-o3 | 22.4 credits per second (silent); 28 credits per second (with audio); 5 s (with audio) = 140 credits | Available | | Grok Imagine Video 1.5 (xAI) | grok-imagine-video-1.5 | 16 credits per second (480p); 28 credits per second (720p); 50 credits per second (1080p); 5 s (720p) = 140 credits; +2 credits per reference image | Available | | FLUX 3 (Black Forest Labs) | flux-3 | 12 credits per second (720p draft); 34 credits per second (720p); 58 credits per second (1080p); 5 s (720p) = 170 credits | Available | | LTX-2.5 (Lightricks) | ltx-2.5 | 18 credits per second (720p); 26 credits per second (1080p); 38 credits per second (1440p); 60 credits per second (4K); 5 s (720p) = 90 credits | Available | | PixVerse V6 (PixVerse) | pixverse-6 | 12 credits per second (720p with audio); from 5 credits (360p silent) to 23 credits (1080p with audio); 5 s (720p with audio) = 60 credits | Available | | Vidu Q3 (Shengshu) | vidu-q3 | 14 credits per second (360p); 14 credits per second (540p); 30.8 credits per second (720p); 30.8 credits per second (1080p); 5 s (720p) = 154 credits | Available | | Luma Ray 3.2 (Luma AI) | luma-ray-3.2 | 100 credits per clip (5 s 540p); 200 credits per clip (5 s 720p); 400 credits per clip (5 s 1080p); 200 credits per clip (10 s 540p); 400 credits per clip (10 s 720p); 800 credits per clip (10 s 1080p); from an image: 30 credits per clip (5 s 540p); from an image: 60 credits per clip (5 s 720p); from an image: 240 credits per clip (5 s 1080p) | Available | | Creatify Boreal (Creatify) | creatify-boreal | 2 credits per second (720p); 6 credits per second (1080p); 24 credits per second (2K); 5 s (720p) = 10 credits | Available | #### Voice | Model | Id | Price | Status | | --- | --- | --- | --- | | Gemini 3.8 Flash TTS (Google) | gemini-3.8-flash-tts | 9 credits per 1,000 characters | Available | | Gemini 3.8 Flash-Lite TTS (Google) | gemini-3.8-flash-lite-tts | 6 credits per 1,000 characters | Available | | ElevenLabs v3 (ElevenLabs) | elevenlabs-v3 | 20 credits per 1,000 characters | Available | | ElevenLabs Multilingual v2 (ElevenLabs) | elevenlabs-multilingual-v2 | 20 credits per 1,000 characters | Available | | MiniMax Speech 2.8 HD (MiniMax) | minimax-speech-2.8-hd | 20 credits per 1,000 characters | Available | | MiniMax Speech 2.8 Turbo (MiniMax) | minimax-speech-2.8-turbo | 12 credits per 1,000 characters | Available | | xAI TTS (xAI) | xai-tts | 3 credits per 1,000 characters | Available | | Inworld TTS 1.5 Max (Inworld) | inworld-tts-1.5 | 2 credits per 1,000 characters | Available | | Kokoro (Hexgrad) | kokoro | 4 credits per 1,000 characters | Available | | Chatterbox HD (Resemble AI) | chatterbox-hd | 8 credits per 1,000 characters | Available | | Dia (Nari Labs) | dia | 8 credits per 1,000 characters | Available | | Orpheus (Canopy Labs) | orpheus | 10 credits per 1,000 characters | Available | #### Music | Model | Id | Price | Status | | --- | --- | --- | --- | | Lyria 3.5 (Google) | lyria-3.5 | 20 credits per track | Available | | ElevenLabs Music v2.5 (ElevenLabs) | elevenlabs-music-2.5 | 120 credits per started minute | Available | | MiniMax Music 2.6 (MiniMax) | minimax-music-2.6 | 30 credits per track | Available | | Stable Audio 3 (Stability AI) | stable-audio-3 | 8 credits per track | Available | #### Audio | Model | Id | Price | Status | | --- | --- | --- | --- | | ElevenLabs Scribe v2 (ElevenLabs) | scribe-v2 | 1.6 credits per started minute | Available | | ElevenLabs Audio Isolation (ElevenLabs) | elevenlabs-audio-isolation | 20 credits per started minute | Available | | DeepFilterNet 3 (DeepFilterNet) | deepfilternet-3 | 12 credits per minute of audio | Available | #### Utilities | Model | Id | Price | Status | | --- | --- | --- | --- | | Ideogram Remove Background (Ideogram) | ideogram-remove-background | 2 credits per image | Available | | Bria RMBG 2.0 (Bria AI) | bria-rmbg-2 | 4 credits per image | Available | | Recraft Crisp Upscale (Recraft) | recraft-crisp-upscale | 1 credit per image | Available | | Topaz Upscale Precision (Topaz Labs) | topaz-upscale | 16 credits per started 24 megapixels of output | Available | | SeedVR2 Upscaler (ByteDance) | seedvr2-upscale | 0.2 credits per megapixel | Available | | Clarity Upscaler (Clarity AI) | clarity-upscaler | 6 credits per megapixel | Available | | Photo Restoration (fal) | photo-restoration | 8 credits per image | Available | | ByteDance Video Upscaler (ByteDance) | bytedance-video-upscaler | 1.44 credits per second (1080p 30 fps); 2.88 credits per second (2K 30 fps); 5.76 credits per second (4K 30 fps); 2.88 credits per second (1080p 60 fps); 5.76 credits per second (2K 60 fps); 11.52 credits per second (4K 60 fps); 5 s (1080p 30 fps) = 8 credits | Coming soon | ### Tools | Tool | Price | | --- | --- | | create_image | From 2 credits per image (model-dependent). remove_background adds 2 per image. | | upscale_image | Topaz (the default): 16 credits per started 24 megapixels of output. The others: from 1 credit per image (Recraft Crisp, SeedVR2) to 6 credits per output megapixel (Clarity). | | remove_background | 2 credits per image (Bria: 4). | | restore_image | 8 credits per image. | | create_video | From 2 credits per second, depending on the model, resolution and sound. | | create_voiceover | From 2 to 20 credits per 1,000 characters, depending on the model. | | create_music | From 8 credits per track. | | transcribe | 1.6 credits per started minute (2.08 with keyterms). | | enhance_audio | 20 credits per started minute (DeepFilterNet: 12 per minute). | | save_brand_kit | Free. | | get_brand_kit | Free. | | delete_brand_kit | Free. | | delete_asset | Free. | | list_models | Free. | | describe_model | Free. | | check_balance | Free. | | buy_credits | Free. | | get_job | Free. | | spend_report | Free. | | send_feedback | Free. | Jobs above 500 credits return a quote first and run only with confirm: true.