EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 30 · 28 per page
text-to-audio
falREVIEW REQUIRED

Stable Audio 3

fal-ai/stable-audio-3/small/music/base/text-to-audio

Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.

musicon-devicelightweight
image-to-image
falREVIEW REQUIRED

Image Editing Retouch

fal-ai/image-editing/retouch

Retouch photos of faces. Remove blemishes and improve the skin.

video-to-video
falREVIEW REQUIRED

Segment Anything Model 2

fal-ai/sam2/video

SAM 2 is a model for segmenting images and videos in real-time.

segmentationmaskreal-time
audio-to-video
falREVIEW REQUIRED

Longcat Single Avatar

fal-ai/longcat-single-avatar/image-audio-to-video

LongCat-Video-Avatar is an audio-driven video generation model that can generates super-realistic, lip-synchronized long video generation with natural dynamics and consistent identity.

audio-to-videoimage-to-video
text-to-speech
MiniMaxREVIEW REQUIRED

Minimax

fal-ai/minimax/preview/speech-2.5-hd

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
text-to-image
falREVIEW REQUIRED

Hidream O1 Image

fal-ai/hidream-o1-image/dev

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.

image-to-video
PixVerseREVIEW REQUIRED

PixVerse V4.5 Transition

fal-ai/pixverse/v4.5/transition

Create seamless transition between images using PixVerse v4.5

stylizedtransform
image-to-video
falREVIEW REQUIRED

Vidu

fal-ai/vidu/q1/reference-to-video

Generate video clips from your multiple image references using Vidu Q1

stylizedtransform
text-to-image
falREVIEW REQUIRED

Nucleus Image

fal-ai/nucleus-image

Nucleus-Image is a text-to-image generation model built on a sparse mixture-of-experts (MoE) diffusion transformer architecture.

stylizedtransformtypography
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] Control LoRA Canny

fal-ai/flux-control-lora-canny/image-to-image

FLUX Control LoRA Canny is a high-performance endpoint that uses a control image using a Canny edge map to transfer structure to the generated image and another initial image to guide color.

lorastyle transfer
3d-to-3d
tripo3dREVIEW REQUIRED

Tripo3D Segment

tripo3d/tripo/segment

Automatically splits a 3D model into semantic parts for editing, texturing, and rigging.

stylizedtransform
video-to-video
Luma AIREVIEW REQUIRED

Luma Ray 2 Flash Reframe

fal-ai/luma-dream-machine/ray-2-flash/reframe

Adjust and enhance videos with Ray-2 Reframe. This advanced tool seamlessly reframes videos to your desired aspect ratio, intelligently inpainting missing regions to ensure realistic visuals and coherent motion, delivering exceptional quality and creative flexibility.

reframeoutpaintflash
text-to-speech
AlibabaREVIEW REQUIRED

Qwen 3 TTS - Text to Speech [0.6B]

fal-ai/qwen-3-tts/text-to-speech/0.6b

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

text-to-speech
image-to-image
falREVIEW REQUIRED

Hidream O1 Image

fal-ai/hidream-o1-image/dev/edit

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.

video-to-video
decartREVIEW REQUIRED

Lucy Edit [Pro]

decart/lucy-edit/pro

Edit outfits, objects, faces, or restyle your video - all with maximum detail retention.

video-edit
text-to-image
falREVIEW REQUIRED

GLM Image

fal-ai/glm-image

Create high-quality images with accurate text rendering and rich knowledge details—supports editing, style transfer, and maintaining consistent characters across multiple images.

text-to-image
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 SRPO [dev]

fal-ai/flux/srpo/image-to-image

FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

image-to-image
falREVIEW REQUIRED

Image Editing Background Change

fal-ai/image-editing/background-change

Replace your photo's background with any scene you desire, from beach sunsets to urban landscapes, with perfect lighting and shadows

stylizedtransform
text-to-3d
falREVIEW REQUIRED

Hyper3D - Rodin V2.5 - Text to 3D - Fast

fal-ai/hyper3d/rodin/v2.5/text-to-3d/fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

text-to-3d
video-to-video
falREVIEW REQUIRED

LTX Video 2.3 Pro

fal-ai/ltx-2.3/retake-video

LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.

stylizedtransformlipsync
video-to-video
falREVIEW REQUIRED

Wan 2.2 VACE Fun A14B

fal-ai/wan-22-vace-fun-a14b/depth

VACE Fun for Wan 2.2 A14B from Alibaba-PAI

text-to-image
imagineartREVIEW REQUIRED

Imagineart 1.5 Preview

imagineart/imagineart-1.5-preview/text-to-image

ImagineArt 1.5 text-to-image model generates high-fidelity professional-grade visuals with lifelike realism, strong aesthetics, and text that actually reads correctly.

visualsimagineartrealismtext
video-to-text
nvidiaREVIEW REQUIRED

Nemotron 3 Nano Omni

nvidia/nemotron-3-nano-omni/video

Video reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts video plus a prompt and returns text.

nemotronnvidiavideo-to-textvideo-understanding
video-to-video
veedREVIEW REQUIRED

Video Background Removal

veed/video-background-removal/green-screen

Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.

vision
falREVIEW REQUIRED

Florence 2 Large Caption

fal-ai/florence-2-large/caption

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

captioningmultimodalvision
video-to-video
AlibabaREVIEW REQUIRED

Wan

fal-ai/wan/v2.2-a14b/video-to-video

Wan-2.2 video-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and source videos.