EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 2 · 28 per page
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX 2 Pro Edit

fal-ai/flux-2-pro/edit

Text-to-image generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.

text-to-image
ByteDanceREVIEW REQUIRED

Seedream 5.0 Pro Text to Image

bytedance/seedream/v5/pro/text-to-image

ByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.

realismtypographystylized
image-to-image
falREVIEW REQUIRED

SeedVR2

fal-ai/seedvr/upscale/image

Use SeedVR2 to upscale your images

upscaleimage-to-image
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Tts Eleven V3

fal-ai/elevenlabs/tts/eleven-v3

Generate text-to-speech audio using Eleven-v3 from ElevenLabs.

audio
image-to-video
KlingREVIEW REQUIRED

Kling Video v3 Image to Video [Standard]

fal-ai/kling-video/v3/standard/image-to-video

Kling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

image-to-video
image-to-image
falREVIEW REQUIRED

Bria RMBG 2.0

fal-ai/bria/background/remove

Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04

background removalimage segmentationhigh resolutionutility
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX1.1 [pro] ultra

fal-ai/flux-pro/v1.1-ultra

FLUX1.1 [pro] ultra is the newest version of FLUX1.1 [pro], maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.

high-resrealism
text-to-video
MiniMaxREVIEW REQUIRED

H3 Max Turbo Text to Video

minimax/h3-max-turbo/text-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylizedtransformlipsync
image-to-video
MiniMaxREVIEW REQUIRED

MiniMax H3 Reference to Video

minimax/h3/reference-to-video

MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.

stylizedtransformlipsync
text-to-image
GoogleREVIEW REQUIRED

Nano Banana 2 Lite

google/nano-banana-2-lite

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

image-to-video
ByteDanceREVIEW REQUIRED

Seedance 2 Image to Video

bytedance/seedance-2.0/image-to-video

ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.

stylizedtransformlipsync
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 9B

fal-ai/flux-2/klein/9b

Text-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

image-to-video
ByteDanceREVIEW REQUIRED

Seedance 2 Reference to Video

bytedance/seedance-2.0/reference-to-video

ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.

stylizedtransformlipsync
vision
openrouterREVIEW REQUIRED

OpenRouter [Vision]

openrouter/router/vision

Run any Vision Language Model with fal. Analyze and understand images using Claude (Anthropic), GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), Llama (Meta), Qwen, Pixtral (Mistral), and more. Send one or multiple images for captioning, analysis, OCR, or visual Q&A. Powered by OpenRouter.

llm
OpenAIREVIEW REQUIRED

OpenRouter Chat Completions [OpenAI Compatible]

openrouter/router/openai/v1/chat/completions

OpenAI-compatible chat completions API. Drop-in replacement for the OpenAI API — use any OpenAI SDK or client to access Claude, Gemini, Grok, DeepSeek, Llama, Qwen, Mistral, and all OpenAI models (GPT-5, GPT-4o, o3) through fal. Powered by OpenRouter.

image-to-video
KlingREVIEW REQUIRED

Kling Video v2.6 Image to Video

fal-ai/kling-video/v2.6/pro/image-to-video

Kling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation.

image-to-image
xAIREVIEW REQUIRED

Grok Imagine Image

xai/grok-imagine-image/edit

Edit images precisely with xAI's Grok Imagine model

grokxaiimage-editing
text-to-image
falREVIEW REQUIRED

Z Image Turbo

fal-ai/z-image/turbo

Z-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

turboz-imagefast
text-to-video
ByteDanceREVIEW REQUIRED

Seedance 2.5 Text to Video

bytedance/seedance-2.5/text-to-video

Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.

stylizedtransformlipsync
image-to-image
ByteDanceREVIEW REQUIRED

Bytedance Seedream V4 Edit

fal-ai/bytedance/seedream/v4/edit

A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.

stylizedtransformediting
llm
openrouterREVIEW REQUIRED

OpenRouter

openrouter/router

Run any LLM with fal. Access Claude (Anthropic), ChatGPT / GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), DeepSeek, Llama (Meta), Qwen (Alibaba), Mistral, and 200+ more models through a single API. Supports reasoning, structured output, and streaming. Powered by OpenRouter.

text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] with LoRAs

fal-ai/flux-lora

Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.

lorapersonalization
image-to-video
MiniMaxREVIEW REQUIRED

MiniMax H3 Image to Video

minimax/h3/image-to-video

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

stylizedtransformlipsync
text-to-image
ByteDanceREVIEW REQUIRED

Bytedance Seedream V4.5 Text To Image

fal-ai/bytedance/seedream/v4.5/text-to-image

A new-generation image creation model ByteDance, Seedream 4.5 integrates image generation and image editing capabilities into a single, unified architecture.

stylizedtransform
text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Speech-02 HD

fal-ai/minimax/speech-02-hd

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
text-to-image
ByteDanceREVIEW REQUIRED

Bytedance Seedream V4 Text To Image

fal-ai/bytedance/seedream/v4/text-to-image

A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.

stylizedtransform
speech-to-text
falREVIEW REQUIRED

Wizper (Whisper v3 -- fal.ai edition)

fal-ai/wizper

[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!

transcriptionspeech