EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 3 · 28 per page
image-to-video
GoogleREVIEW REQUIRED

Veo 3.1 Fast

fal-ai/veo3.1/fast/image-to-video

Generate videos from your image prompts using Veo 3.1 fast.

text-to-audio
ElevenLabsREVIEW REQUIRED

ElevenLabs TTS Multilingual v2

fal-ai/elevenlabs/tts/multilingual-v2

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.

audio
image-to-image
falREVIEW REQUIRED

Segment Anything Model 3

fal-ai/sam-3/image

SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.

segmentationmaskreal-time
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Sound Effects V2

fal-ai/elevenlabs/sound-effects/v2

Generate sound effects using ElevenLabs advanced sound effects model.

sound
image-to-video
MiniMaxREVIEW REQUIRED

H3 Max Camera Controls

minimax/h3-max/camera-controls

H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space

stylizedtransformediting
text-to-speech
GoogleREVIEW REQUIRED

Gemini 3.1 Flash Tts

fal-ai/gemini-3.1-flash-tts

Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

lipsyncavatar
image-to-image
falREVIEW REQUIRED

Clarity Upscaler

fal-ai/clarity-upscaler

Clarity upscaler for upscaling images with high very fidelity.

upscaling
speech-to-text
ElevenLabsREVIEW REQUIRED

ElevenLabs Speech to Text - Scribe V2

fal-ai/elevenlabs/speech-to-text/scribe-v2

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

speech-to-text
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Music

fal-ai/elevenlabs/music

Generate high quality, realistic music with fine controls using Elevenlabs Music!

musictext-to-music
image-to-image
falREVIEW REQUIRED

Birefnet Background Removal

fal-ai/birefnet

bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)

background removalsegmentationhigh-resutility
image-to-image
GoogleREVIEW REQUIRED

Nano Banana Lite Edit

google/nano-banana-lite/edit

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

text-to-audio
MiniMaxREVIEW REQUIRED

MiniMax Music 3

minimax/music-3

MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long

sfxaudioeffects
image-to-video
ByteDanceREVIEW REQUIRED

Bytedance Seedance V1.5 Pro Image To Video

fal-ai/bytedance/seedance/v1.5/pro/image-to-video

Generate videos with audio with Seedance 1.5 (supports start & end frame)

bytedanceseedanceaudio
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2

fal-ai/flux-2

Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

image-to-video
GoogleREVIEW REQUIRED

Veo 3.1

fal-ai/veo3.1/image-to-video

Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind

image-to-video
xAIREVIEW REQUIRED

Grok Imagine Video

xai/grok-imagine-video/image-to-video

Generate videos from images with audio using xAI's Grok Imagine Video model.

grokxaiimage-to-videoi2v
image-to-image
GoogleREVIEW REQUIRED

Gemini 3 Pro Image Preview

fal-ai/gemini-3-pro-image-preview/edit

Gemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model

realismtypography
text-to-speech
ElevenLabsREVIEW REQUIRED

ElevenLabs TTS Turbo v2.5

fal-ai/elevenlabs/tts/turbo-v2.5

Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.

audio
text-to-image
xAIREVIEW REQUIRED

Grok Imagine Image

xai/grok-imagine-image

Generate highly aesthetic images with xAI's Grok Imagine Image generation model.

xaigroktext-to-image
image-to-image
GoogleREVIEW REQUIRED

Gemini 2.5 Flash Image

fal-ai/gemini-25-flash-image/edit

Google's famous original image generation and editing model, a.k.a Nano Banana

image-editing
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 9B

fal-ai/flux-2/klein/9b/edit

Image-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.

image-to-image
metaREVIEW REQUIRED

Meta Muse Image Edit

meta/muse-image/edit

Meta's Muse Image model does precise edits that change only what you ask, stay coherent across turns, and compose from multiple reference images.

realismtypographystylizedediting
image-to-image
xAIREVIEW REQUIRED

Grok Imagine Image 2.0

xai/grok-imagine-image/v2.0/edit

Edit images with xAi's Grok Imagine 2.0 model.

xaigrokimage-editing
image-to-video
xAIREVIEW REQUIRED

Grok Imagine Video 1.5

xai/grok-imagine-video/v1.5/image-to-video

Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.

stylizedtransformlipsync
image-to-video
GoogleREVIEW REQUIRED

Veo3.1 Lite Image to Video

fal-ai/veo3.1/lite/image-to-video

Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video

stylizedtransformlipsync
image-to-3d
falREVIEW REQUIRED

Hunyuan 3D Pro Image to 3D

fal-ai/hunyuan-3d/v3.1/pro/image-to-3d

Generate 3D models from images with Hunyuan 3D Pro

3dhunyuanimage-to-3d
image-to-image
OpenAIREVIEW REQUIRED

GPT-Image 1.5

fal-ai/gpt-image-1.5/edit

GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.

openaigpt-image
image-to-video
KlingREVIEW REQUIRED

Kling 2.1 (standard)

fal-ai/kling-video/v2.1/standard/image-to-video

Kling 2.1 Standard is a cost-efficient endpoint for the Kling 2.1 model, delivering high-quality image-to-video generation