image-to-imageFLUX 2 Pro Edit
fal-ai/flux-2-pro/editText-to-image generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/flux-2-pro/editText-to-image generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.
text-to-imagebytedance/seedream/v5/pro/text-to-imageByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.
image-to-image
text-to-audiofal-ai/elevenlabs/tts/eleven-v3Generate text-to-speech audio using Eleven-v3 from ElevenLabs.
image-to-videofal-ai/kling-video/v3/standard/image-to-videoKling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
image-to-imagefal-ai/bria/background/removeBria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04
text-to-imagefal-ai/flux-pro/v1.1-ultraFLUX1.1 [pro] ultra is the newest version of FLUX1.1 [pro], maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.
text-to-videominimax/h3-max-turbo/text-to-videofal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
image-to-videominimax/h3/reference-to-videoMiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.
text-to-imagegoogle/nano-banana-2-liteNano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
image-to-videobytedance/seedance-2.0/image-to-videoByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.
text-to-imagefal-ai/flux-2/klein/9bText-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
image-to-videobytedance/seedance-2.0/reference-to-videoByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.
visionopenrouter/router/visionRun any Vision Language Model with fal. Analyze and understand images using Claude (Anthropic), GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), Llama (Meta), Qwen, Pixtral (Mistral), and more. Send one or multiple images for captioning, analysis, OCR, or visual Q&A. Powered by OpenRouter.
llmopenrouter/router/openai/v1/chat/completionsOpenAI-compatible chat completions API. Drop-in replacement for the OpenAI API — use any OpenAI SDK or client to access Claude, Gemini, Grok, DeepSeek, Llama, Qwen, Mistral, and all OpenAI models (GPT-5, GPT-4o, o3) through fal. Powered by OpenRouter.
image-to-videofal-ai/kling-video/v2.6/pro/image-to-videoKling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation.
image-to-imagexai/grok-imagine-image/editEdit images precisely with xAI's Grok Imagine model
text-to-imagefal-ai/z-image/turboZ-Image Turbo is a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
text-to-videobytedance/seedance-2.5/text-to-videoDreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.
image-to-imagefal-ai/bytedance/seedream/v4/editA new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.
llmopenrouter/routerRun any LLM with fal. Access Claude (Anthropic), ChatGPT / GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), DeepSeek, Llama (Meta), Qwen (Alibaba), Mistral, and 200+ more models through a single API. Supports reasoning, structured output, and streaming. Powered by OpenRouter.
text-to-imagefal-ai/flux-loraSuper fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.
image-to-videominimax/h3/image-to-videoMiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.
text-to-imagefal-ai/bytedance/seedream/v4.5/text-to-imageA new-generation image creation model ByteDance, Seedream 4.5 integrates image generation and image editing capabilities into a single, unified architecture.
image-to-image
text-to-speechfal-ai/minimax/speech-02-hdGenerate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
text-to-imagefal-ai/bytedance/seedream/v4/text-to-imageA new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.
speech-to-textfal-ai/wizper[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!