image-to-imageKling O1 Image
fal-ai/kling-image/o1Perform precise image edits using strong reference control, transforming subjects, styles, and local details while preserving visual consistency.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/kling-image/o1Perform precise image edits using strong reference control, transforming subjects, styles, and local details while preserving visual consistency.
speech-to-speechresemble-ai/chatterboxhd/speech-to-speechTransform voices using Resemble AI's Chatterbox. Convert audio to new voices or your own samples, with expressive results and built-in perceptual watermarking.
image-to-videofal-ai/bytedance/omnihumanOmniHuman generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.
image-to-videoluma/agent/ray/v3.2/image-to-videoLuma Ray 3.2 animates a source image into cinematic motion guided by a text prompt, preserving the starting frame's look while controlling resolution, duration, and seamless looping.
text-to-videofal-ai/bytedance/seedance/v1/pro/text-to-videoSeedance 1.0 Pro, a high quality video generation model developed by Bytedance.
video-to-videofal-ai/birefnet/v2/videoVideo background removal version of bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
video-to-videoxai/grok-imagine-video/edit-videoEdit videos using xAI's Grok Imagine
text-to-imagefal-ai/ideogram/v2Generate high-quality images, posters, and logos with Ideogram V2. Features exceptional typography handling and realistic outputs optimized for commercial and creative use.
video-to-videofal-ai/kling-video/o3/standard/video-to-video/referenceKling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
text-to-imagefal-ai/recraft/v4.1/pro/text-to-imageRecraft V4.1 Pro pushes the V4.1 model into high-resolution territory — up to 2048×2048 and ultra-wide formats. Made for hero imagery, campaign work, and print, it preserves the same design taste at sizes ready for the final deliverable.
image-to-imagefal-ai/qwen-image-edit-2511/loraEndpoint for Qwen's Image Editing 2511 model with LoRa support.
image-to-image
text-to-imagefal-ai/wan/v2.7/text-to-imageGenerate high-quality images from text prompts using the WAN 2.7 model with advanced prompt understanding and detailed output.
text-to-imagefal-ai/flux-pro/kontext/max/text-to-imageFLUX.1 Kontext [max] text-to-image is a new premium model brings maximum performance across all aspects – greatly improved prompt adherence.
text-to-speechfal-ai/minimax/speech-2.6-hdGenerate speech from text prompts and different voices using the MiniMax Speech-2.6 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
image-to-imagefal-ai/kling-image/v3/image-to-imageKling Image V3: Latest kling image model
text-to-videofal-ai/wan/v2.2-a14b/text-to-videoWan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
text-to-videofal-ai/bytedance/seedance/v1/pro/fast/text-to-videoText to Video endpoint for Seedance 1.0 Pro Fast, a next-generation video model designed to deliver maximum performance at minimal cost
video-to-videofal-ai/wan/v2.2-14b/animate/moveWan-Animate is a video model that generates high-fidelity character videos by replicating the expressions and movements of characters from reference videos.
image-to-videofal-ai/kling-video/o3/4k/reference-to-videoKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
visionfal-ai/moondream3-preview/queryMoondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.
video-to-videoveed/video-background-removalRemove background from any video with people and objects. No green screen needed.
image-to-imagefal-ai/flux-2/klein/9b/base/editImage-to-image editing with Flux 2 [klein] 9B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
image-to-videoalibaba/happy-horse/reference-to-videoGenerate 1080p video with synchronized native audio from a text prompt and references. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.
text-to-speechfal-ai/qwen-3-tts/voice-design/1.7bCreate custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!
image-to-imagefal-ai/flux-2/klein/9b/edit/loraImage-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs and custom LoRA. Precise modifications using natural language descriptions and hex color control.
video-to-videofal-ai/wan/v2.7/edit-videoWan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
text-to-imagefal-ai/flux/kreaFLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.