Auto-Captioner
fal-ai/auto-captionAutomatically generates text captions for your videos from the audio as per text colour/font specifications
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
fal-ai/auto-captionAutomatically generates text captions for your videos from the audio as per text colour/font specifications
image-to-videofal-ai/hunyuan-video-image-to-videoImage to Video for the high-quality Hunyuan Video I2V model.
text-to-imagerundiffusion-fal/juggernaut-flux/proJuggernaut Pro Flux by RunDiffusion is the flagship Juggernaut model rivaling some of the most advanced image models available, often surpassing them in realism. It combines Juggernaut Base with RunDiffusion Photo and features enhancements like reduced background blurriness.
text-to-imagefal-ai/flux/srpoFLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
vision
text-to-videofal-ai/luma-dream-machine/ray-2-flashRay2 Flash is a fast video generative model capable of creating realistic visuals with natural, coherent motion.
text-to-audiofal-ai/stable-audio-3/small/sfx/base/text-to-audioStable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.
image-to-videofal-ai/ovi/image-to-videoOvi can generate videos with audio from image and text inputs.
speech-to-textfal-ai/cohere-transcribeCohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation
image-to-videofal-ai/wan-pro/image-to-videoWan-2.1 Pro is a premium image-to-video model that generates high-quality 1080p videos at 30fps with up to 6 seconds duration, delivering exceptional visual quality and motion diversity from images
image-to-imagefal-ai/stable-diffusion-v3-medium/image-to-imageStable Diffusion 3 Medium (Image to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.
image-to-3dfal-ai/hunyuan3d/v2/turboGenerate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.
image-to-imagefal-ai/flux-2-lora-gallery/face-to-full-portraitExtends a face into a full body portrait
video-to-videobria/video/background-removal/green-screen-despillRemove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.
video-to-videofal-ai/bernini-r/edit-videoEdit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping the rest of the scene intact.
visionperceptron/isaac-01Isaac-01 is a multimodal vision-language model from Perceptron for various vision language tasks.
audio-to-videolightricks/ltx-2.5/audio-to-video/fastLTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates video timed to a supplied audio clip in a speed-optimized mode — useful for music-driven content, dialogue-led shorts, and ads keyed to a track.
trainingfal-ai/z-image-trainerTrain LoRAs on Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
image-to-videofal-ai/pixverse/v4.5/effectsGenerate high quality video clips with different effects using PixVerse v4.5
speech-to-textnvidia/nemotron-asr-multilingual/asrNemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.
text-to-imagebria/fibo/generateSOTA open-source text-to-image model delivering high-fidelity outputs with accurate typography. JSON-structured prompts provide production-ready controllability for enterprise and agentic workflows. Trained exclusively on licensed data.
image-to-imagefal-ai/image-apps-v2/style-transferApply artistic styles like impressionism, cubism, or surrealism to your images.
text-to-videofal-ai/pixverse/v5.5/text-to-videoGenerate high quality video clips from text and image prompts using PixVerse v5.5
text-to-imagefal-ai/hunyuan-image/v2.1/text-to-imageUse the amazing capabilities of hunyuan image 2.1 to generate images that express the feelings of your text.
image-to-imagefal-ai/control-lightControlLight is a LoRA fine-tune of FLUX.2 [klein] 9B that enhances low-light images while preserving scene structure and fine details, with a single alpha parameter that gives continuous control over enhancement strength from subtle to full brightening.
text-to-audiofal-ai/kokoro/frenchAn expressive and natural French text-to-speech model for both European and Canadian French.
image-to-imagefal-ai/flux-lora-cannyUtilize Flux.1 [dev] Controlnet to generate high-quality images with precise control over composition, style, and structure through advanced edge detection and guidance mechanisms.
image-to-imagefal-ai/image-apps-v2/product-photographyGenerate professional product photography with realistic lighting and backgrounds.