image-to-videoVeo 3.1 Fast
fal-ai/veo3.1/fast/image-to-videoGenerate videos from your image prompts using Veo 3.1 fast.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videofal-ai/veo3.1/fast/image-to-videoGenerate videos from your image prompts using Veo 3.1 fast.
text-to-audiofal-ai/elevenlabs/tts/multilingual-v2Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.
image-to-imagefal-ai/sam-3/imageSAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.
fal-ai/elevenlabs/sound-effects/v2Generate sound effects using ElevenLabs advanced sound effects model.
image-to-videominimax/h3-max/camera-controlsH3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space
text-to-speechfal-ai/gemini-3.1-flash-ttsNewest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.
image-to-imagefal-ai/clarity-upscalerClarity upscaler for upscaling images with high very fidelity.
speech-to-textfal-ai/elevenlabs/speech-to-text/scribe-v2Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!
text-to-audiofal-ai/elevenlabs/musicGenerate high quality, realistic music with fine controls using Elevenlabs Music!
image-to-imagefal-ai/birefnetbilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
image-to-imagegoogle/nano-banana-lite/editNano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
text-to-audiominimax/music-3MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long
image-to-videofal-ai/bytedance/seedance/v1.5/pro/image-to-videoGenerate videos with audio with Seedance 1.5 (supports start & end frame)
text-to-imagefal-ai/flux-2Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
image-to-videofal-ai/veo3.1/image-to-videoVeo 3.1 is the latest state-of-the art video generation model from Google DeepMind
image-to-videoxai/grok-imagine-video/image-to-videoGenerate videos from images with audio using xAI's Grok Imagine Video model.
image-to-imagefal-ai/gemini-3-pro-image-preview/editGemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model
fal-ai/elevenlabs/tts/turbo-v2.5Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.
text-to-imagexai/grok-imagine-imageGenerate highly aesthetic images with xAI's Grok Imagine Image generation model.
image-to-imagefal-ai/gemini-25-flash-image/editGoogle's famous original image generation and editing model, a.k.a Nano Banana
image-to-imagefal-ai/flux-2/klein/9b/editImage-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
image-to-imagemeta/muse-image/editMeta's Muse Image model does precise edits that change only what you ask, stay coherent across turns, and compose from multiple reference images.
image-to-imagexai/grok-imagine-image/v2.0/editEdit images with xAi's Grok Imagine 2.0 model.
image-to-videoxai/grok-imagine-video/v1.5/image-to-videoGenerate videos from images with audio using xAI's Grok Imagine 1.5 Video model.
image-to-videofal-ai/veo3.1/lite/image-to-videoVeo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video
image-to-3dfal-ai/hunyuan-3d/v3.1/pro/image-to-3dGenerate 3D models from images with Hunyuan 3D Pro
image-to-imagefal-ai/gpt-image-1.5/editGPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.
image-to-videofal-ai/kling-video/v2.1/standard/image-to-videoKling 2.1 Standard is a cost-efficient endpoint for the Kling 2.1 model, delivering high-quality image-to-video generation