EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 9 · 28 per page
text-to-video
KlingREVIEW REQUIRED

Kling Video v2.6 Text to Video

fal-ai/kling-video/v2.6/pro/text-to-video

Kling 2.6 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation.

text-to-video
GoogleREVIEW REQUIRED

Veo3.1 Lite Text to Video

fal-ai/veo3.1/lite

Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video

stylizedtransformlipsync
video-to-video
falREVIEW REQUIRED

MMAudio V2

fal-ai/mmaudio-v2

MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.

ai videofast
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 Kontext [pro]

fal-ai/flux-pro/kontext/text-to-image

The FLUX.1 Kontext [pro] text-to-image delivers state-of-the-art image generation results with unprecedented prompt following, photorealistic rendering, and flawless typography.

image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit

fal-ai/qwen-image-edit

Endpoint for Qwen's Image Editing model. Has superior text editing capabilities.

image-editingimage-to-imagehigh-quality-text
text-to-video
ByteDanceREVIEW REQUIRED

Bytedance Seedance V1.5 Pro Text To Video

fal-ai/bytedance/seedance/v1.5/pro/text-to-video

Generate videos with audio with Seedance 1.5

bytedanceseedanceaudio
image-to-image
falREVIEW REQUIRED

Image Preprocessors

fal-ai/image-preprocessors/depth-anything/v2

Depth Anything v2 preprocessor.

depthpreprocessutilitycontrolnet
image-to-video
xAIREVIEW REQUIRED

Grok Imagine Reference to Video

xai/grok-imagine-video/reference-to-video

Generate videos using multiple reference images with xAI's Grok Imagine video model

video-editv2vgrokxai
image-to-video
KlingREVIEW REQUIRED

Kling Video

fal-ai/kling-video/v3/4k/image-to-video

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

stylizedtransformlipsync
text-to-image
xAIREVIEW REQUIRED

Grok Imagine Image

xai/grok-imagine-image/quality/text-to-image

Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.

stylizedtransformtypography
text-to-video
GoogleREVIEW REQUIRED

Gemini Omni Flash

google/gemini-omni-flash

Creates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.

stylizedtransformlipsync
speech-to-text
ElevenLabsREVIEW REQUIRED

Elevenlabs - Forced Alignment

fal-ai/elevenlabs/forced-alignment

Align the transcript and your audio recording using Elevenlab's forced alignment feature!

forced-alignmentspeech-to-text
video-to-video
veedREVIEW REQUIRED

Lipsync

veed/lipsync

Generate realistic lipsync from any audio using VEED's model.

lipsyncvideo-to-videoavatar
image-to-video
falREVIEW REQUIRED

LTX 2.3 Video Fast

fal-ai/ltx-2.3/image-to-video/fast

LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.

stylizedtransformlipsync
text-to-audio
falREVIEW REQUIRED

Lyria2

fal-ai/lyria2

Lyria 2 is Google's latest music generation model, you can generate any type of music with this model.

musicstylized
text-to-audio
falREVIEW REQUIRED

ACE Step

fal-ai/ace-step

Generate music with lyrics from text using ACE-Step

text-to-audiotext-to-music
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2 Pro Outpaint

fal-ai/flux-2-pro/outpaint

Outpainting generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.

image-to-imageoutpaintoutpainting
image-to-image
falREVIEW REQUIRED

EVF-SAM2 Segmentation

fal-ai/evf-sam

EVF-SAM2 combines natural language understanding with advanced segmentation capabilities, allowing you to precisely mask image regions using intuitive positive and negative text prompts.

segmentationmask
text-to-image
ByteDanceREVIEW REQUIRED

Seedream

bytedance/seedream/v5/lite/text-to-image

Text to Image endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent text-to-image generation.

text-to-imagebytedanceseedream-5.0-lite
image-to-3d
falREVIEW REQUIRED

Hyper3D - Rodin V2.5 - Image to 3D

fal-ai/hyper3d/rodin/v2.5

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

image-to-3d
text-to-video
AlibabaREVIEW REQUIRED

Wan 3.0 Prime

alibaba/wan-3.0-prime/text-to-video

Wan 3.0 Prime Text-to-Video transforms written prompts into polished videos with accelerated generation, fluid motion, strong scene fidelity, and coherent visual storytelling. Built for fast creative iteration, it brings complex ideas to life while preserving visual detail and cinematic consistency throughout each shot.

textvideo
video-to-video
KlingREVIEW REQUIRED

Kling O3 Edit Video [Standard]

fal-ai/kling-video/o3/standard/video-to-video/edit

Edit videos using Kling O3 from Kling Team!

video-to-video
text-to-image
GoogleREVIEW REQUIRED

Nano Banana Lite

google/nano-banana-lite

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

text-to-audio
falREVIEW REQUIRED

Stable Audio 3

fal-ai/stable-audio-3/medium/text-to-audio

Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.

musicaudiostereo
text-to-image
RecraftREVIEW REQUIRED

Recraft V4.1 Text to Image

fal-ai/recraft/v4.1/text-to-image

Recraft V4.1 builds on the design-first foundation of V4 with sharper prompt control and cleaner composition. Tuned for brand systems and editorial work, it delivers production-ready raster images that hold up next to a designer's hand.

stylizedtransformtypography
text-to-image
Black Forest LabsREVIEW REQUIRED

Flux 2 Flex

fal-ai/flux-2-flex

Text-to-image generation with FLUX.2 [flex] from Black Forest Labs. Features adjustable inference steps and guidance scale for fine-tuned control. Enhanced typography and text rendering capabilities.

stylizedtransform
video-to-video
veedREVIEW REQUIRED

Subtitles

veed/subtitles

VEED’s Subtitles API transforms raw footage into polished, publish-ready content with professional burned-in subtitles starting at a base rate of $0.10 per minute.

video-to-video
veedREVIEW REQUIRED

VEED Lipsync

veed/lipsync/v2

Generate production-quality lipsync from any audio using VEED's most advanced model yet.

veedlipsyncvideo-to-videoavatar