EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 31 · 28 per page
video-to-video
falREVIEW REQUIRED

Auto-Captioner

fal-ai/auto-caption

Automatically generates text captions for your videos from the audio as per text colour/font specifications

captioningvideo
image-to-video
falREVIEW REQUIRED

Hunyuan Video Image-to-Video Inference

fal-ai/hunyuan-video-image-to-video

Image to Video for the high-quality Hunyuan Video I2V model.

motion
text-to-image
Black Forest LabsREVIEW REQUIRED

Juggernaut Flux Pro

rundiffusion-fal/juggernaut-flux/pro

Juggernaut Pro Flux by RunDiffusion is the flagship Juggernaut model rivaling some of the most advanced image models available, often surpassing them in realism. It combines Juggernaut Base with RunDiffusion Photo and features enhancements like reduced background blurriness.

image generation
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 SRPO [dev]

fal-ai/flux/srpo

FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

text-to-video
Luma AIREVIEW REQUIRED

Luma Ray 2 Flash

fal-ai/luma-dream-machine/ray-2-flash

Ray2 Flash is a fast video generative model capable of creating realistic visuals with natural, coherent motion.

motiontransformation
text-to-audio
falREVIEW REQUIRED

Stable Audio 3 Small SFX Base Text to Audio

fal-ai/stable-audio-3/small/sfx/base/text-to-audio

Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.

sfxsound-effectson-device
image-to-video
falREVIEW REQUIRED

Ovi

fal-ai/ovi/image-to-video

Ovi can generate videos with audio from image and text inputs.

image-to-audio-videoimage-to-video
speech-to-text
falREVIEW REQUIRED

Cohere Transcribe

fal-ai/cohere-transcribe

Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

speechtranscribestt
image-to-video
falREVIEW REQUIRED

Wan-2.1 Pro Image-to-Video

fal-ai/wan-pro/image-to-video

Wan-2.1 Pro is a premium image-to-video model that generates high-quality 1080p videos at 30fps with up to 6 seconds duration, delivering exceptional visual quality and motion diversity from images

image to videomotion
image-to-image
falREVIEW REQUIRED

Stable Diffusion V3

fal-ai/stable-diffusion-v3-medium/image-to-image

Stable Diffusion 3 Medium (Image to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.

diffusioneditingstyle
image-to-3d
falREVIEW REQUIRED

Hunyuan3D

fal-ai/hunyuan3d/v2/turbo

Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-image
Black Forest LabsREVIEW REQUIRED

Flux 2 Lora Gallery

fal-ai/flux-2-lora-gallery/face-to-full-portrait

Extends a face into a full body portrait

stylizedtransform
video-to-video
briaREVIEW REQUIRED

Video

bria/video/background-removal/green-screen-despill

Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.

video-to-video
falREVIEW REQUIRED

Bernini-R Edit Video

fal-ai/bernini-r/edit-video

Edit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping the rest of the scene intact.

edittransformstylized
vision
perceptronREVIEW REQUIRED

Isaac 0.1

perceptron/isaac-01

Isaac-01 is a multimodal vision-language model from Perceptron for various vision language tasks.

multimodalvision
audio-to-video
lightricksREVIEW REQUIRED

Ltx 2.5 Audio to Video Fast

lightricks/ltx-2.5/audio-to-video/fast

LTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates video timed to a supplied audio clip in a speed-optimized mode — useful for music-driven content, dialogue-led shorts, and ads keyed to a track.

stylizedtransformlip-sync
training
falREVIEW REQUIRED

Z Image Trainer

fal-ai/z-image-trainer

Train LoRAs on Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

turboz-imagefasttrainer
image-to-video
PixVerseREVIEW REQUIRED

PixVerse V4.5 Effects

fal-ai/pixverse/v4.5/effects

Generate high quality video clips with different effects using PixVerse v4.5

image-to-video
speech-to-text
nvidiaREVIEW REQUIRED

Nemotron Asr Multilingual

nvidia/nemotron-asr-multilingual/asr

Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.

utilitytranscribe
text-to-image
briaREVIEW REQUIRED

Fibo

bria/fibo/generate

SOTA open-source text-to-image model delivering high-fidelity outputs with accurate typography. JSON-structured prompts provide production-ready controllability for enterprise and agentic workflows. Trained exclusively on licensed data.

briafiboprompt-adherence
image-to-image
falREVIEW REQUIRED

Style Transfer

fal-ai/image-apps-v2/style-transfer

Apply artistic styles like impressionism, cubism, or surrealism to your images.

style-transfer
text-to-video
PixVerseREVIEW REQUIRED

PixVerse V5.5 Text To Video

fal-ai/pixverse/v5.5/text-to-video

Generate high quality video clips from text and image prompts using PixVerse v5.5

text-to-video
text-to-image
falREVIEW REQUIRED

Hunyuan Image

fal-ai/hunyuan-image/v2.1/text-to-image

Use the amazing capabilities of hunyuan image 2.1 to generate images that express the feelings of your text.

text-to-image
image-to-image
falREVIEW REQUIRED

ControlLight

fal-ai/control-light

ControlLight is a LoRA fine-tune of FLUX.2 [klein] 9B that enhances low-light images while preserving scene structure and fine details, with a single alpha parameter that gives continuous control over enhancement strength from subtle to full brightening.

stylizedtransform
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (French)

fal-ai/kokoro/french

An expressive and natural French text-to-speech model for both European and Canadian French.

speech
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] Canny with LoRAs

fal-ai/flux-lora-canny

Utilize Flux.1 [dev] Controlnet to generate high-quality images with precise control over composition, style, and structure through advanced edge detection and guidance mechanisms.

controlnetdetectionloraediting
image-to-image
falREVIEW REQUIRED

Product Photography

fal-ai/image-apps-v2/product-photography

Generate professional product photography with realistic lighting and backgrounds.

productmarketing