text-to-3dHyper3D - Rodin V2.5 - Text to 3D
fal-ai/hyper3d/rodin/v2.5/text-to-3dRodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-3dfal-ai/hyper3d/rodin/v2.5/text-to-3dRodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
image-to-imagetopaz/sharpen/imageProfessional photo sharpening powered by Topaz Labs. Models tuned per blur type (lens, motion, portrait, wildlife), plus Super Focus for generative recovery of severely blurred shots. Best for out-of-focus and motion-blurred photos.
image-to-videofal-ai/wan/v2.2-a14b/image-to-video/loraWan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and images. This endpoint supports LoRAs made for Wan 2.2
audio-to-videofal-ai/wan/v2.2-14b/speech-to-videoWan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applications
image-to-imagebria/extract-objectBria Extract Object uses text prompts to isolate a selected object from an image and return it as an RGBA PNG with a transparent background. Ideal for product, ecommerce, advertising, and creative editing workflows. Bria's Extract Object API leads in product shot extraction, outperforming SAM 3.1 where it counts most for commercial use.
text-to-imagefal-ai/wan-25-preview/text-to-imageWan 2.5 text-to-image model.
video-to-videofal-ai/luma-dream-machine/ray-2/modifyRay2 Modify is a video generative model capable of restyling or retexturing the entire shot, from turning live-action into CG or stylized animation, to changing wardrobe, props, or the overall aesthetic and swap environments or time periods, giving you control over background, location, or even weather.
video-to-videofal-ai/kling-video/o1/standard/video-to-video/editEdit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.
image-to-image
image-to-image
text-to-audiofal-ai/kokoro/british-englishA high-quality British English text-to-speech model offering natural and expressive voice synthesis.
video-to-videomirelo-ai/sfx1.6/video-to-videoGenerate synced sounds for any video, and return it with its new sound track (like MMAudio). Now up to 60 seconds!
text-to-imageimagineart/imagineart-1.5-pro-preview/text-to-imageImagineArt 1.5 Pro is an advanced text-to-image model that creates ultra-high-fidelity 4K visuals with lifelike realism, refined aesthetics, and powerful creative output suited for professional use.
text-to-imagefal-ai/stable-diffusion-v3-mediumStable Diffusion 3 Medium (Text to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.
text-to-imagefal-ai/flux-control-lora-cannyFLUX Control LoRA Canny is a high-performance endpoint that uses a control image to transfer structure to the generated image, using a Canny edge map.
audio-to-audiofal-ai/stable-audio-3/medium/audio-to-audioStable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.
text-to-videofal-ai/ltx-video-13b-distilledGenerate videos from prompts using LTX Video-0.9.7 13B Distilled and custom LoRA
text-to-imageluma/agent/uni-1/v1/maxLuma Uni-1 Max generates a single image at the model's highest fidelity, delivering richer detail and stronger prompt adherence than the base tier for hero-quality stills.
text-to-imagefal-ai/flux-2/klein/9b/base/loraText-to-image generation with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.
text-to-imagerundiffusion-fal/juggernaut-flux/lightningJuggernaut Lightning Flux by RunDiffusion provides blazing-fast, high-quality images rendered at five times the speed of Flux. Perfect for mood boards and mass ideation, this model excels in both realism and prompt adherence.
audio-to-videofal-ai/ltx-2.3/audio-to-videoLTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.
text-to-videofal-ai/luma-dream-machine/ray-2Ray2 is a large-scale video generative model capable of creating realistic visuals with natural, coherent motion.
image-to-imagefal-ai/flux-pro/v1.1-ultra/reduxFLUX1.1 [pro] ultra Redux is a high-performance endpoint for the FLUX1.1 [pro] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
image-to-imagefal-ai/image-editing/object-removalRemove unwanted objects or people from your photos while seamlessly blending the background.
video-to-videofal-ai/sync-lipsync/react-1Use React-1 from SyncLabs to refine human emotions and do realistic lip-sync without losing details!
llmfal-ai/bytedance/seed/v2/miniSeed 2.0 Mini is a high-performance multimodal model optimized for low latency and high concurrency. It supports text, image, and video input with 256K context and configurable thinking/reasoning modes.
text-to-videofal-ai/kling-video/o3/4k/text-to-videoKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
image-to-imagefal-ai/flux-2/klein/4b/edit/loraImage-to-image editing with FLUX.2 [klein] 4B from Black Forest Labs and custom LoRA. Precise modifications using natural language descriptions and hex color control.