image-to-imageFlux 2 Flex
fal-ai/flux-2-flex/editImage editing with FLUX.2 [flex] from Black Forest Labs. Supports multi-reference editing with customizable inference steps and enhanced text rendering.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/flux-2-flex/editImage editing with FLUX.2 [flex] from Black Forest Labs. Supports multi-reference editing with customizable inference steps and enhanced text rendering.
visionfal-ai/video-understandingA video understanding model to analyze video content and answer questions about what's happening in the video based on user prompts.
image-to-videofal-ai/sync-lipsync/v3/image-to-videosync-3 image to video turns a single still into a talking character, and works with any illustration or animated frame paired with a voice track
image-to-imagefal-ai/kling-image/o3/image-to-imageKling Omni 3: Top-tier image-to-image with flawless consistency.
image-to-videoxai/grok-imagine-video/v1.5/reference-to-videoGenerate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.
text-to-imagefal-ai/flux-2/loraText-to-image generation with LoRA support for FLUX.2 [dev] from Black Forest Labs. Custom style adaptation and fine-tuned model variations.
text-to-imagefal-ai/gpt-image-1/text-to-imageOpenAI's latest image generation and editing model: gpt-1-image.
text-to-imagefal-ai/krea-2/turbo/loraGenerate high-fidelity images from text with Krea 2 using a custom-trained LoRA. Apply your LoRA weights to carry a learned subject, character, or style into new generations, with aspect ratio, creativity, and seed controls.
video-to-videofal-ai/sync-lipsyncGenerate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization.
text-to-audiofal-ai/kokoro/american-englishKokoro is a lightweight text-to-speech model that delivers comparable quality to larger models while being significantly faster and more cost-efficient.
image-to-imagefal-ai/flux-pro/kontext/multiExperimental version of FLUX.1 Kontext [pro] with multi image handling capabilities
text-to-speechfal-ai/qwen-3-tts/text-to-speech/1.7bBring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
text-to-audioelevenlabs/music/v2.5Generate high quality, realistic music with fine controls using Elevenlabs Music v2.5!
image-to-videofal-ai/kling-video/v1.6/pro/image-to-videoGenerate video clips from your images using Kling 1.6 (pro)
image-to-imagefal-ai/flux-lora/image-to-imageFLUX LoRA Image-to-Image is a high-performance endpoint that transforms existing images using FLUX models, leveraging LoRA adaptations to enable rapid and precise image style transfer, modifications, and artistic variations.
text-to-audiosonilo/v1.1/text-to-musicGenerates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.
text-to-imagefal-ai/gemini-3.1-flash-image-previewGemini 3.1 Flash Image (a.k.a Nano Banana 2) is Google's new state-of-the-art fast image generation and editing model
vision
image-to-videofal-ai/wan-i2vWan-2.1 is a image-to-video model that generates high-quality videos with high visual quality and motion diversity from images
video-to-videofal-ai/wan/v2.2-14b/animate/replaceWan-Animate Replace is a model that can integrate animated characters into reference videos, replacing the original character while preserving the scene’s lighting and color tone for seamless environmental integration.
text-to-audiosonilo/v1.1/text-to-sound-effectsGenerates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.
audio-to-audiofal-ai/ffmpeg-api/merge-audiosMerge audios into a single audio using FFmpeg API!
image-to-imagefal-ai/gpt-image-1/edit-imageOpenAI's latest image generation and editing model: gpt-1-image.
image-to-imagefal-ai/sam-3-1/imageSAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
text-to-speechfal-ai/minimax/speech-2.8-turboGenerate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
trainingfal-ai/krea-2-trainerTrain a custom LoRA on your own images to teach Krea 2 a new subject, character, or style. Provide a set of training images (and an optional trigger word), and the trainer outputs LoRA weights you can use for inference with the Krea 2 LoRA endpoint.
image-to-imagefal-ai/z-image/turbo/image-to-imageGenerate images from text and images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
text-to-audiofal-ai/elevenlabs/text-to-dialogue/eleven-v3Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.