trainingFLUX 2 Trainer
fal-ai/flux-2-trainerFine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
trainingfal-ai/flux-2-trainerFine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.
text-to-3dfal-ai/hunyuan-3d/v3.1/rapid/text-to-3dCreate detailed, fully-textured 3D models with text
image-to-videofal-ai/kling-video/o1/standard/reference-to-videoTransform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.
text-to-imagefal-ai/photaPhota's model empowers developers, photographers, and creators with personalized photograph generation and editing.
text-to-3dfal-ai/meshy/v6/text-to-3dMeshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.
image-to-imagefal-ai/z-image/turbo/inpaintGenerate images from text, an image and a mask using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
text-to-speechfal-ai/mayaMaya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.
text-to-video
video-to-audiofal-ai/kling-video/video-to-audioGenerate audio from input videos using Kling
audio-to-textnvidia/nemotron-3-nano-omni/audioAudio reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts audio plus a prompt and returns text.
audio-to-audiofal-ai/stable-audio-3/small/music/base/audio-outpaintingStable Audio 3 Small Music Base audio outpainting is the foundational 459 million parameter checkpoint that extends music tracks via causal continuation guided by text prompts.
image-to-videofal-ai/ltx-2.3-22b/distilled/image-to-videoGenerate video with audio from images using LTX-2.3 Distilled
text-to-audiofal-ai/minimax-music/v1.5Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
image-to-videofal-ai/ltx-2-19b/image-to-videoGenerate video with audio from images using LTX-2
video-to-videofal-ai/sam-3-1/video-rleSAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
image-to-imagefal-ai/ben/v2/imageA fast and high quality model for image background removal.
audio-to-audiofal-ai/stable-audio-3/small/music/base/audio-to-audioStable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.
visionfal-ai/moondream3-preview/captionMoondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.
text-to-3dfal-ai/meshy/v6-preview/text-to-3dMeshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.
jsonfal-ai/ffmpeg-api/waveformGet waveform data from audio files using FFmpeg API.
text-to-imagefal-ai/luma-photonGenerate images from your prompts using Luma Photon. Photon is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.
image-to-imagefal-ai/ideogram/v3/replace-backgroundReplace backgrounds existing images with Ideogram V3's replace background feature. Create variations and adaptations while preserving core elements and adding new creative directions through prompt guidance.
text-to-videofal-ai/pixverse/c1/text-to-videoGenerate film-grade videos from text prompts with native audio, up to 1080p and 15 seconds, using PixVerse C1.
image-to-imagefal-ai/qwen-image-edit-plus-loraLoRA endpoint for the Qwen Image Edit Plus model.
image-to-imagefal-ai/flux-control-lora-depth/image-to-imageFLUX Control LoRA Depth is a high-performance endpoint that uses a control image using a depth map to transfer structure to the generated image and another initial image to guide color.
image-to-imagefal-ai/post-processingPost Processing is an endpoint that can enhance images using a variety of techniques including grain, blur, sharpen, and more.
text-to-speechfal-ai/orpheus-ttsOrpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation. This model has been finetuned to deliver human-level speech synthesis, achieving exceptional clarity, expressiveness, and real-time performances.
image-to-imagetopaz/denoise/imageProfessional photo denoising powered by Topaz Labs. Normal, Strong and Extreme presets clean noise at source resolution; Denoise Max adds generative detail recovery. Best for high-ISO and night photography.