EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 28 · 28 per page
vision
falREVIEW REQUIRED

Moondream3 Preview [Point]

fal-ai/moondream3-preview/point

Moondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.

Vision
image-to-image
briaREVIEW REQUIRED

Fibo Edit [Restore]

bria/fibo-edit/restore

Photo restoration model that automatically denoises, deblurs, and enhances old or damaged photos - removes imperfections while preserving original character.

image-restorationfibo-editbriajson
training
Black Forest LabsREVIEW REQUIRED

FLUX 2 [klein] 9b Base Trainer

fal-ai/flux-2-klein-9b-base-trainer

Fine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.

image-to-video
falREVIEW REQUIRED

Ltx 2.3 Quality

fal-ai/ltx-2.3-quality/image-to-video

Generate high-quality video with audio from images using LTX-2.3

image-to-video
text-to-image
RecraftREVIEW REQUIRED

Recraft V4.1 Text to Image Utility

fal-ai/recraft/v4.1/utility/text-to-image

Recraft V4.1 Utility is a faster, lighter variant of V4.1 made for high-volume creative workflows. Ideal for ideation, A/B exploration, and content pipelines, it keeps Recraft's design sensibility while optimizing for throughput and cost.

stylizedtransformtypography
image-to-json
falREVIEW REQUIRED

VGGT-1B

fal-ai/vggt-1b

Turn images or video into a detailed 3D scene with depth, camera poses, and a colored point cloud.

image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Max

fal-ai/qwen-image-max/edit

Image editing endpoint for Qwen-Image-Max. Qwen Image Max improves upon the Qwen Image Plus series by enhancing the realism and naturalness of images.

qwen-imagemax
image-to-image
falREVIEW REQUIRED

Z Image Turbo Controlnet

fal-ai/z-image/turbo/controlnet

Generate images from text and edge, depth or pose images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

text-to-image
falREVIEW REQUIRED

Fooocus Inpainting

fal-ai/fooocus/inpaint

Default parameters with automated optimizations and quality improvements.

stylizedediting
llm
falREVIEW REQUIRED

Video Prompt Generator

fal-ai/video-prompt-generator

Generate video prompts using a variety of techniques including camera direction, style, pacing, special effects and more.

motiontransformationchatclaude
audio-to-audio
falREVIEW REQUIRED

Stable Audio 2.5

fal-ai/stable-audio-25/audio-to-audio

Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI

audio
image-to-video
falREVIEW REQUIRED

LongCat Video Distilled

fal-ai/longcat-video/distilled/image-to-video/720p

Generate long videos in 720p/30fps from images using LongCat Video Distilled

image-to-image
pixelcutREVIEW REQUIRED

Pixelcut Product Photo

pixelcut/product-photo

Pixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery

utilityediting
video-to-video
soniloREVIEW REQUIRED

V1.1 Video to Video Sound Effects

sonilo/v1.1/video-to-video-sound-effects

Adds synchronized, royalty-free, commercial-use-safe sound effects to a video. Returns the finished video with the generated audio mixed in.

sfxaudioeffects
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] Depth with LoRAs

fal-ai/flux-lora-depth

Generate high-quality images from depth maps using Flux.1 [dev] depth estimation model. The model produces accurate depth representations for scene understanding and 3D visualization.

depthlorautilitycomposition
3d-to-3d
falREVIEW REQUIRED

Hunyuan 3D Smart Topology

fal-ai/hunyuan-3d/v3.1/smart-topology

Optimize 3D mesh topology with Hunyuan 3D Smart Topology.

3dhunyuantopology
video-to-video
decartREVIEW REQUIRED

Lucy 2.5

decart/lucy-2-5/realtime

Real-time, prompt-driven video editing over WebRTC. Restyle, swap backgrounds, and add or replace objects live on a webcam or streamed feed at interactive latency.

realtimevideo-to-videowebrtc
text-to-image
AlibabaREVIEW REQUIRED

Wan

fal-ai/wan/v2.2-a14b/text-to-image

Wan 2.2's 14B model generates high-resolution, photorealistic images with powerful prompt understanding and fine-grained visual detail

audio-to-audio
falREVIEW REQUIRED

Sam Audio

fal-ai/sam-audio/span-separate

Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.

audio-to-audiosam-audio
image-to-image
IdeogramREVIEW REQUIRED

Ideogram Upscale

fal-ai/ideogram/upscale

Ideogram Upscale enhances the resolution of the reference image by up to 2X and might enhance the reference image too. Optionally refine outputs with a prompt for guided improvements.

upscalinghigh-res
video-to-video
decartREVIEW REQUIRED

Lucy Restyle

decart/lucy-restyle

Restyle videos up to 30 min long - maintaining maximum detail quality.

video-edit
image-to-image
falREVIEW REQUIRED

try-on

fal-ai/cat-vton

Image based high quality Virtual Try-On

try-onfashionclothing
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Spanish)

fal-ai/kokoro/spanish

A natural-sounding Spanish text-to-speech model optimized for Latin American and European Spanish.

speech
image-to-image
briaREVIEW REQUIRED

Genfill

bria/genfill/v2

The GenFill Route enables the generation of objects by prompt in a specific region of an image. You can define the area for object generation by using a mask that outlines the region where the object will be created. Our model is optimized to work seamlessly with blob-shaped masks.

vision
falREVIEW REQUIRED

Florence 2 Large OCR

fal-ai/florence-2-large/ocr

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

ocrmultimodalvision
text-to-video
falREVIEW REQUIRED

Heygen v5 Digital Twin

fal-ai/heygen/avatar5/digital-twin

Create natural HeyGen Avatar V digital twin videos from text or audio, with lip-sync, optional backgrounds, captions, and MP4/WebM output.

avatardigital-twintalking-avatartext-to-video