EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 8 · 28 per page
image-to-image
Black Forest LabsREVIEW REQUIRED

PuLID Flux

fal-ai/flux-pulid

An endpoint for personalized image generation using Flux as per given description.

personalizationstyle transfer
text-to-image
AlibabaREVIEW REQUIRED

Qwen Image

fal-ai/qwen-image

Qwen-Image is an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing.

text-to-image
image-to-video
KlingREVIEW REQUIRED

Kling 1.6

fal-ai/kling-video/v1.6/standard/image-to-video

Generate video clips from your images using Kling 1.6 (std)

image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 Kontext [max]

fal-ai/flux-pro/kontext/max/multi

Experimental version of FLUX.1 Kontext [max] with multi image handling capabilities

image-to-image
clarityaiREVIEW REQUIRED

Crystal Upscaler

clarityai/crystal-upscaler

An advanced image enhancement tool designed specifically for facial details and portrait photography, utilizing Clarity AI's upscaling technology.

image-to-image
video-to-video
falREVIEW REQUIRED

Sync Lipsync

fal-ai/sync-lipsync/v2/pro

Generate high-quality realistic lipsync animations from audio while preserving unique details like natural teeth and unique facial features using the state-of-the-art Sync Lipsync 2 Pro model.

animationlip synchigh-quality
vision
falREVIEW REQUIRED

NSFW Filter

fal-ai/imageutils/nsfw

Predict the probability of an image being NSFW.

filtersafetyutility
video-to-video
falREVIEW REQUIRED

SeedVR2

fal-ai/seedvr/upscale/video

Upscale your videos using SeedVR2 with temporal consistency!

upscalevideo-to-video
text-to-video
GoogleREVIEW REQUIRED

Gemini Omni Flash 1.1 Text to Video

google/gemini-omni-flash/v1.1/text-to-video

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control expressed in natural language.

stylizedtransformlipsync
text-to-audio
falREVIEW REQUIRED

Lyria 3 Pro

fal-ai/lyria3/pro

Lyria 3 Pro is the latest music model from Google

audiosfx
video-to-video
Black Forest LabsREVIEW REQUIRED

Flux 3 FAST Edit Video

blackforestlabs/flux-3/edit-video

FLUX.3 Edit Video [FAST] is Black Forest Labs' frontier video model. This endpoint edits an existing video from natural-language instructions, applying targeted changes while preserving the rest of the scene.

stylizedtransformlipsync
image-to-image
ByteDanceREVIEW REQUIRED

Seedream 5.0 Pro Layerize

bytedance/seedream/v5/pro/layerize

Splits a finished image into independent, editable transparent-PNG layers — background plus separate elements, from a text description, returning 2 to 17 layers per call for non-destructive reuse in design tools.

utilityediting
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2 Turbo Edit

fal-ai/flux-2/turbo/edit

Image-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control—all at turbo speed.

text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 4B

fal-ai/flux-2/klein/4b

Text-to-image generation with FLUX.2 [klein] 4B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 4B

fal-ai/flux-2/klein/4b/edit

Image-to-image editing with FLUX.2 [klein] 4B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.

text-to-image
AlibabaREVIEW REQUIRED

Qwen Image 3 Text to Image

alibaba/qwen-image-3/text-to-image

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence

stylizedtransformtypography
image-to-video
GoogleREVIEW REQUIRED

Gemini Omni Flash

google/gemini-omni-flash/reference-to-video

Generates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.

stylizedtransformlipsync
image-to-image
falREVIEW REQUIRED

Segment Anything Model 2

fal-ai/sam2/image

SAM 2 is a model for segmenting images and videos in real-time.

segmentationmaskreal-time
video-to-video
falREVIEW REQUIRED

Workflow Utilities Trim Video

fal-ai/workflow-utilities/trim-video

FFMPEG Utility for Trim Video

video-to-video
video-to-video
KlingREVIEW REQUIRED

Kling Video v2.6 Motion Control [Pro]

fal-ai/kling-video/v2.6/pro/motion-control

Transfer movements from a reference video to any character image. Pro mode delivers higher quality output, ideal for complex dance moves and gestures.

text-to-video
Black Forest LabsREVIEW REQUIRED

Flux 3 Text to Video

blackforestlabs/flux-3/text-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.

stylizedtransformlipsync
text-to-video
GoogleREVIEW REQUIRED

Veo3.1 Lite Text to Video

fal-ai/veo3.1/lite

Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video

stylizedtransformlipsync
text-to-image
RecraftREVIEW REQUIRED

Recraft V4

fal-ai/recraft/v4/text-to-image

Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.

text-to-image
image-to-image
falREVIEW REQUIRED

FASHN Virtual Try-On V1.6

fal-ai/fashn/tryon/v1.6

FASHN v1.6 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 864x1296 resolution from both on-model and flat-lay photo references.

try-onfashionclothing
video-to-video
falREVIEW REQUIRED

MMAudio V2

fal-ai/mmaudio-v2

MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.

ai videofast
image-to-video
GoogleREVIEW REQUIRED

Veo 3.1 Fast

fal-ai/veo3.1/fast/first-last-frame-to-video

Generate videos from a first/last frame using Google's Veo 3.1 Fast

image-to-video
GoogleREVIEW REQUIRED

Veo 3.1

fal-ai/veo3.1/reference-to-video

Generate Videos from images using Google's Veo 3.1

text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Speech-02 Turbo

fal-ai/minimax/speech-02-turbo

Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech