text-to-imageGPT-Image 1.5
fal-ai/gpt-image-1.5GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/gpt-image-1.5GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.
image-to-imagefal-ai/recraft/upscale/crispEnhances a given raster image using 'crisp upscale' tool, boosting resolution with a focus on refining small details and faces.
text-to-imagefal-ai/flux-2/turboText-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities—all at turbo speed.
video-to-videofal-ai/ffmpeg-api/merge-videosUse ffmpeg capabilities to merge 2 or more videos.
image-to-imagefal-ai/ideogram/remove-backgroundRemove backgrounds from existing images with Ideogram's remove background feature. Isolate subjects cleanly for compositing and creative reuse.
image-to-videofal-ai/kling-video/ai-avatar/v2/standardKling AI Avatar v2 Standard: Endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters
text-to-videominimax/h3/text-to-videoMiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.
image-to-imagetopaz/upscale/image/precisionProfessional photo upscaling powered by Topaz Labs. Gigapixel precision models (Standard V2, High Fidelity, Low Resolution, CGI, Text Refine) enlarge images faithfully up to 4x. Best for photos that must stay true to the original.
image-to-imagepixelcut/background-removalPixelcut’s Background Remover enables fast, ultra high-quality removal of backgrounds from images. Perfect for e-commerce and image editing workflows. Powered by advanced AI for clean, perfect cutouts every time.
video-to-videofal-ai/ffmpeg-api/composeCompose videos from multiple media sources using FFmpeg API.
image-to-videofal-ai/kling-video/v3/turbo/pro/image-to-videoGenerate high quality 1080p videos from images using Kling's Turbo 3.0 model, with improved lipsync and multishot generation capabilities.
image-to-videogoogle/gemini-omni-flash/v1.1/reference-to-videoGemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video from combined multimodal references, images, videos and text together. Reasoning across all inputs to produce a single coherent result, with characters retaining their face, clothing, and voice throughout
image-to-videofal-ai/bytedance/omnihuman/v1.5Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.
text-to-imagemeta/muse-image/text-to-imageMeta's Muse Image model has faithful instruction-following and exceptional visual fidelity, with fine details like text, plots, and QR codes rendered accurately.
text-to-image
video-to-videofal-ai/kling-video/v3/pro/motion-controlTransfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.
image-to-imagefal-ai/ffmpeg-api/extract-frameffmpeg endpoint for first, middle and last frame extraction from videos
text-to-audiofal-ai/minimax-music/v2.6MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
image-to-videoalibaba/wan-3.0-prime/image-to-videoWan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.
text-to-videofal-ai/veo3.1Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!
image-to-imagefal-ai/qwen-image-edit-2511-multiple-anglesGenerates same scene from different angles (azimuth/elevation) with Qwen image Edit 2511 and the Lora Multiple Angles
text-to-speechfal-ai/minimax/voice-cloneClone a voice from a sample audio and generate speech from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality text-to-speech.
text-to-videofal-ai/kling-video/v2.5-turbo/pro/text-to-videoKling 2.5 Turbo Pro: Top-tier text-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
text-to-imagexai/grok-imagine-image/v2.0/text-to-imageGenerate images from text using xAi's Grok Imagine 2.0 model.
image-to-imagefal-ai/bria/expandBria Expand expands images beyond their borders in high quality. Trained exclusively on licensed data for safe and risk-free commercial use. Access the model's source code and weights: https://bria.ai/contact-us
video-to-videofal-ai/kling-video/v2.6/standard/motion-controlTransfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.
text-to-videofal-ai/kling-video/v3/standard/text-to-videoKling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.
video-to-videofal-ai/sync-lipsync/v2Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization with Sync Lipsync 2.0 model