image-to-imageWan
fal-ai/wan/v2.7/editTransform and edit existing images with text-guided instructions using the WAN 2.7 model for creative image manipulation.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/wan/v2.7/editTransform and edit existing images with text-guided instructions using the WAN 2.7 model for creative image manipulation.
video-to-videogoogle/gemini-omni-flash/editEdits generated video across multiple conversational turns while preserving scene coherence. Applies iterative changes through natural-language instructions without regenerating the full sequence from scratch.
text-to-videofal-ai/pixverse/v6/text-to-videoPixverse's latest v6 Model.
image-to-imagefal-ai/flux-kontext-loraFast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.
image-to-imagefal-ai/qwen-image-edit-plusEndpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.
text-to-speechfal-ai/inworld-ttsText to Speech Endpoint for Inworld's TTS-1.5 Max.
image-to-videoblackforestlabs/flux-3/first-last-frame-to-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.
video-to-videofal-ai/kling-video/o3/pro/video-to-video/referenceKling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
video-to-textopenrouter/router/videoRun any video-capable LLM with fal. Analyze, summarize, and understand video files using Gemini (Google) models. Supports mp4, mpeg, mov, webm, and YouTube links. Powered by OpenRouter.
text-to-videofal-ai/kling-video/v1.6/standard/text-to-videoGenerate video clips from your prompts using Kling 1.6 (std)
image-to-videofal-ai/kling-video/o1/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
image-to-videolightricks/ltx-2.5/image-to-video/fastLTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.
image-to-videofal-ai/minimax/hailuo-2.3-fast/standard/image-to-videoMiniMax Hailuo-2.3-Fast Image To Video API (Standard, 768p): Advanced fast image-to-video generation model with 768p resolution
video-to-videofal-ai/depth-anything-videoGenerates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.
text-to-videobytedance/seedance-2.0/mini/text-to-videoSeedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
image-to-videoalibaba/happy-horse/image-to-videoAlibaba's #1-ranked Happy Horse 1.0 — generate 1080p video with synchronized native audio and multilingual lip-sync from text prompts or images.
audio-to-audiofal-ai/sam-audio/separateAudio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
image-to-3dtripo3d/h3.1/multiview-to-3dGenerate 3D models from multiple view images using Tripo H3.1.
image-to-video
image-to-videofal-ai/minimax/hailuo-02/pro/image-to-videoMiniMax Hailuo-02 Image To Video API (Pro, 1080p): Advanced image-to-video generation model with 1080p resolution
video-to-videotopaz/upscale/video/generativeProfessional generative video upscaling powered by Topaz Labs. Starlight models rebuild detail that is not in the source, with Fast variants at half the price. Best for low-quality, compressed or archive footage.
video-to-videofal-ai/kling-video/o1/video-to-video/editEdit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.
text-to-speechxai/tts/v1Generate speech with expressive and realistic voices from xAI
image-to-3dmeshy/v7/multi-image-to-3deconstructs a high-fidelity textured 3D model from multiple angle views of one object, with game-ready topology and polygon control
text-to-videofal-ai/kling-video/o3/pro/text-to-videoGenerate realistic videos using Kling O3 from Kling Team!
text-to-imagefal-ai/z-image/turbo/loraText-to-Image endpoint with LoRA support for Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
image-to-imagefal-ai/qwen-image-layeredQwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers.
video-to-videobria/video/background-removal/v3Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.