image-to-videoVidu Image to Video
fal-ai/vidu/q1/image-to-videoVidu Q1 Image to Video generates high-quality 1080p videos with exceptional visual quality and motion diversity from a single image
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videofal-ai/vidu/q1/image-to-videoVidu Q1 Image to Video generates high-quality 1080p videos with exceptional visual quality and motion diversity from a single image
image-to-videofal-ai/hunyuan-video-v1.5/image-to-videoHunyuan Video 1.5 is Tencent's latest and best video model
video-to-videofal-ai/id-v2v/relightChange a video’s lighting using a relit reference frame while preserving the scene, subjects, and original performance. ID-V2V Relight propagates the new illumination across the video.
text-to-audiofal-ai/kokoro/italianA high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.
video-to-videofal-ai/pixverse/extendPixVerse Extend model is a video extending tool for your videos using with high-quality video extending techniques
image-to-imagehitem3d/hi3d/image-to-reliefGenerate a 3D relief depth map with Hi3D from a single image.
text-to-videoveed/avatars/text-to-videoGenerate high-quality videos with UGC-like avatars from text
video-to-videofal-ai/heygen/v3/lipsync/speedReplace or dub audio on an existing video with fast audio-only lip-sync.
text-to-imagerundiffusion-fal/juggernaut-flux-loraJuggernaut Base Flux LoRA by RunDiffusion is a drop-in replacement for Flux [Dev] that delivers sharper details, richer colors, and enhanced realism to all your LoRAs and LyCORIS with full compatibility.
3d-to-3dfal-ai/hunyuan-3d/v3.1/partSplit 3D models into parts with Hunyuan 3D
image-to-imagefal-ai/flux-1/schnell/reduxFLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
video-to-videofal-ai/scail-2SCAIL-2 is an end-to-end character animation model that drives a reference character from a source video without relying on intermediate pose representations like skeleton maps.
image-to-imagefal-ai/qwen-image-edit-plus-lora-gallery/integrate-productBlend products into backgrounds with automatic perspective and lighting correction
video-to-videofal-ai/krea-wan-14b/video-to-videoSuperfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.
text-to-imagefal-ai/ideogram/v2a/turboAccelerated image generation with Ideogram V2A Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.
image-to-imagefal-ai/image-editing/style-transferTransform your photos into artistic masterpieces inspired by famous styles like Van Gogh's Starry Night or any artistic style you choose.
text-to-imagefal-ai/wan/v2.2-a14b/text-to-image/loraWan 2.2's 14B model with LoRA support generates high-fidelity images with enhanced prompt alignment, style adaptability.
trainingrecraft/v4/create-styleCreates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.
image-to-imagefal-ai/sdxl-controlnet-union/image-to-imageAn efficent SDXL multi-controlnet image-to-image model.
image-to-imagefal-ai/playground-v25/image-to-imageState-of-the-art open-source model in aesthetic quality
image-to-imagefal-ai/image-apps-v2/makeup-applicationApply realistic makeup styles with adjustable intensity.
image-to-imagefal-ai/qwen-image-edit-2509-lora-gallery/multiple-anglesPrecise camera position and angle control (rotation, zoom, vertical movement)
text-to-audiofal-ai/csm-1bCSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.
text-to-videofal-ai/vidu/q2/text-to-videoUse the latest Vidu Q2 models which much more better quality and control on your videos.
image-to-3dfal-ai/meshy/v5/multi-image-to-3dMeshy-5 multi image generates realistic and production ready 3D models from multiple images.
audio-to-audiofal-ai/stable-audio-3/medium/base/audio-to-audioStable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.
video-to-videofal-ai/bernini-r/reference-edit-videoEdit a video guided by reference images with Bernini-R, bringing an object, material, background, style, or weather from a reference image into your video.
image-to-image