text-to-videoKling Video v2.6 Text to Video
fal-ai/kling-video/v2.6/pro/text-to-videoKling 2.6 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-videofal-ai/kling-video/v2.6/pro/text-to-videoKling 2.6 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation.
text-to-videofal-ai/veo3.1/liteVeo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video
video-to-videofal-ai/mmaudio-v2MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.
text-to-imagefal-ai/flux-pro/kontext/text-to-imageThe FLUX.1 Kontext [pro] text-to-image delivers state-of-the-art image generation results with unprecedented prompt following, photorealistic rendering, and flawless typography.
image-to-imagefal-ai/qwen-image-editEndpoint for Qwen's Image Editing model. Has superior text editing capabilities.
text-to-videofal-ai/bytedance/seedance/v1.5/pro/text-to-videoGenerate videos with audio with Seedance 1.5
image-to-imagefal-ai/image-preprocessors/depth-anything/v2Depth Anything v2 preprocessor.
image-to-videoxai/grok-imagine-video/reference-to-videoGenerate videos using multiple reference images with xAI's Grok Imagine video model
image-to-videofal-ai/kling-video/v3/4k/image-to-videoKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
text-to-imagexai/grok-imagine-image/quality/text-to-imageGrok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.
text-to-videogoogle/gemini-omni-flashCreates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.
speech-to-textfal-ai/elevenlabs/forced-alignmentAlign the transcript and your audio recording using Elevenlab's forced alignment feature!
video-to-videoveed/lipsyncGenerate realistic lipsync from any audio using VEED's model.
image-to-videofal-ai/ltx-2.3/image-to-video/fastLTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.
text-to-audiofal-ai/lyria2Lyria 2 is Google's latest music generation model, you can generate any type of music with this model.
text-to-audio
image-to-imagefal-ai/flux-2-pro/outpaintOutpainting generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.
image-to-imagefal-ai/evf-samEVF-SAM2 combines natural language understanding with advanced segmentation capabilities, allowing you to precisely mask image regions using intuitive positive and negative text prompts.
text-to-imagebytedance/seedream/v5/lite/text-to-imageText to Image endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent text-to-image generation.
image-to-3dfal-ai/hyper3d/rodin/v2.5Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
text-to-videoalibaba/wan-3.0-prime/text-to-videoWan 3.0 Prime Text-to-Video transforms written prompts into polished videos with accelerated generation, fluid motion, strong scene fidelity, and coherent visual storytelling. Built for fast creative iteration, it brings complex ideas to life while preserving visual detail and cinematic consistency throughout each shot.
video-to-videofal-ai/kling-video/o3/standard/video-to-video/editEdit videos using Kling O3 from Kling Team!
text-to-imagegoogle/nano-banana-liteNano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
text-to-audiofal-ai/stable-audio-3/medium/text-to-audioStable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.
text-to-imagefal-ai/recraft/v4.1/text-to-imageRecraft V4.1 builds on the design-first foundation of V4 with sharper prompt control and cleaner composition. Tuned for brand systems and editorial work, it delivers production-ready raster images that hold up next to a designer's hand.
text-to-imagefal-ai/flux-2-flexText-to-image generation with FLUX.2 [flex] from Black Forest Labs. Features adjustable inference steps and guidance scale for fine-tuned control. Enhanced typography and text rendering capabilities.
video-to-videoveed/subtitlesVEED’s Subtitles API transforms raw footage into polished, publish-ready content with professional burned-in subtitles starting at a base rate of $0.10 per minute.
video-to-videoveed/lipsync/v2Generate production-quality lipsync from any audio using VEED's most advanced model yet.