EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 25 · 28 per page
training
Black Forest LabsREVIEW REQUIRED

FLUX 2 Trainer

fal-ai/flux-2-trainer

Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.

text-to-3d
falREVIEW REQUIRED

Hunyuan 3d

fal-ai/hunyuan-3d/v3.1/rapid/text-to-3d

Create detailed, fully-textured 3D models with text

3d
image-to-video
KlingREVIEW REQUIRED

Kling O1 Reference Image to Video [Standard]

fal-ai/kling-video/o1/standard/reference-to-video

Transform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.

text-to-image
falREVIEW REQUIRED

Phota Text to Image

fal-ai/phota

Phota's model empowers developers, photographers, and creators with personalized photograph generation and editing.

stylizedtransformtypographyphota
text-to-3d
MeshyREVIEW REQUIRED

Meshy 6

fal-ai/meshy/v6/text-to-3d

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

text-to-3d
image-to-image
falREVIEW REQUIRED

Z Image Turbo Inpaint

fal-ai/z-image/turbo/inpaint

Generate images from text, an image and a mask using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

inpainting
text-to-speech
falREVIEW REQUIRED

Maya1

fal-ai/maya

Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.

text-to-speechtts
video-to-audio
KlingREVIEW REQUIRED

Kling Video

fal-ai/kling-video/video-to-audio

Generate audio from input videos using Kling

audio-to-text
nvidiaREVIEW REQUIRED

Nemotron 3 Nano Omni

nvidia/nemotron-3-nano-omni/audio

Audio reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts audio plus a prompt and returns text.

nemotronnvidiaaudio-to-textaudio-understanding
audio-to-audio
falREVIEW REQUIRED

Stable Audio 3 Small Music Base Audio Outpainting

fal-ai/stable-audio-3/small/music/base/audio-outpainting

Stable Audio 3 Small Music Base audio outpainting is the foundational 459 million parameter checkpoint that extends music tracks via causal continuation guided by text prompts.

musicextensioncontinuation
image-to-video
falREVIEW REQUIRED

LTX-2.3 22B Distilled

fal-ai/ltx-2.3-22b/distilled/image-to-video

Generate video with audio from images using LTX-2.3 Distilled

text-to-audio
MiniMaxREVIEW REQUIRED

MiniMax (Hailuo AI) Music v1.5

fal-ai/minimax-music/v1.5

Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

music
image-to-video
falREVIEW REQUIRED

LTX-2 19B

fal-ai/ltx-2-19b/image-to-video

Generate video with audio from images using LTX-2

video-to-video
falREVIEW REQUIRED

Sam 3 1

fal-ai/sam-3-1/video-rle

SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.

segmentationmaskreal-time
image-to-image
falREVIEW REQUIRED

ben-v2-image

fal-ai/ben/v2/image

A fast and high quality model for image background removal.

background removal
audio-to-audio
falREVIEW REQUIRED

Stable Audio 3 Small Music Base Audio to Audio

fal-ai/stable-audio-3/small/music/base/audio-to-audio

Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.

musicstyle-transferremix
vision
falREVIEW REQUIRED

Moondream3 Preview [Caption]

fal-ai/moondream3-preview/caption

Moondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.

Vision
text-to-3d
MeshyREVIEW REQUIRED

Meshy 6 Preview

fal-ai/meshy/v6-preview/text-to-3d

Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.

text-to-3d
json
falREVIEW REQUIRED

FFmpeg API Waveform

fal-ai/ffmpeg-api/waveform

Get waveform data from audio files using FFmpeg API.

ffmpeg
text-to-image
Luma AIREVIEW REQUIRED

Luma Photon

fal-ai/luma-photon

Generate images from your prompts using Luma Photon. Photon is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.

image-to-image
IdeogramREVIEW REQUIRED

Ideogram Replace Background

fal-ai/ideogram/v3/replace-background

Replace backgrounds existing images with Ideogram V3's replace background feature. Create variations and adaptations while preserving core elements and adding new creative directions through prompt guidance.

text-to-video
PixVerseREVIEW REQUIRED

PixVerse C1 Text To Video

fal-ai/pixverse/c1/text-to-video

Generate film-grade videos from text prompts with native audio, up to 1080p and 15 seconds, using PixVerse C1.

video-generationtext-to-videopixversecinematic
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit Plus Lora

fal-ai/qwen-image-edit-plus-lora

LoRA endpoint for the Qwen Image Edit Plus model.

image-to-imageimage-editing
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] Control LoRA Depth

fal-ai/flux-control-lora-depth/image-to-image

FLUX Control LoRA Depth is a high-performance endpoint that uses a control image using a depth map to transfer structure to the generated image and another initial image to guide color.

lorastyle transfer
image-to-image
falREVIEW REQUIRED

Post Processing

fal-ai/post-processing

Post Processing is an endpoint that can enhance images using a variety of techniques including grain, blur, sharpen, and more.

stylizedutility
text-to-speech
falREVIEW REQUIRED

Orpheus TTS

fal-ai/orpheus-tts

Orpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation. This model has been finetuned to deliver human-level speech synthesis, achieving exceptional clarity, expressiveness, and real-time performances.

text to speechvoice synthesishigh-fidelity
image-to-image
topazREVIEW REQUIRED

Topaz Denoise Image

topaz/denoise/image

Professional photo denoising powered by Topaz Labs. Normal, Strong and Extreme presets clean noise at source resolution; Denoise Max adds generative detail recovery. Best for high-ISO and night photography.

restoreimage