image-to-imageQwen Image Layered
fal-ai/qwen-image-layered/loraQwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers. Use loras to get your custom outputs.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/qwen-image-layered/loraQwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers. Use loras to get your custom outputs.
video-to-videodecart/lucy2-vton/realtimeRealtime Try On experience with Decart Lucy 2.1 VTON
text-to-videofal-ai/infinitalk/single-textInfinitalk model generates a talking avatar video from a text and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.
image-to-imagefal-ai/lcm-sd15-i2iProduce high-quality images with minimal inference steps. Optimized for 512x512 input image size.
text-to-imagefal-ai/ideogram/v2/turboAccelerated image generation with Ideogram V2 Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.
text-to-videofal-ai/kandinsky5/text-to-video/distillKandinsky 5.0 Distilled is a lightweight diffusion model for fast, high-quality text-to-video generation.
image-to-imagefal-ai/image-editing/wojak-styleTransform your photos into wojak style while keeping the original characters likeness
text-to-videofal-ai/longcat-video/distilled/text-to-video/480pGenerate long videos from text using LongCat Video Distilled
text-to-imagefal-ai/bitdanceImage generation with BitDance. Fast, high-resolution photorealistic images using an autoregressive LLM— for efficient, high-quality results.
trainingfal-ai/phota/create-profileGenerate profiles using 30-50 images of a subject with Phota.
video-to-videofal-ai/ltx-video-13b-distilled/multiconditioningGenerate videos from prompts, images, and videos using LTX Video-0.9.7 13B Distilled and custom LoRA
text-to-imagefal-ai/omnigen-v1OmniGen is a unified image generation model that can generate a wide range of images from multi-modal prompts. It can be used for various tasks such as Image Editing, Personalized Image Generation, Virtual Try-On, Multi Person Generation and more!
audio-to-audiofal-ai/stable-audio-3/small/music/base/audio-inpaintingStable Audio 3 Small Music Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected music segments guided by text prompts.
image-to-3dfal-ai/hunyuan3d/v2/multi-view/turboGenerate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.
video-to-videoblackforestlabs/flux-3/extend-video/draftFLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that continue an existing clip, with a reusable draft cache for full-quality enhancement.
video-to-videofal-ai/lightx/relightUse tlightx capabilities to relight and recamera your videos.
image-to-imagefal-ai/telestyle-v2Restyle any image with TeleStyle v2 — provide an original image and a styling reference, and the model re-renders the original in the reference's visual style while preserving its content and composition.
trainingrecraft/v4/pro/create-styleCreates a reusable style from your reference images for use with Recraft V4 Styles Pro generation
trainingfal-ai/flux-2-klein-9b-base-trainer/editFine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.
text-to-imagefal-ai/ideogram/v2aGenerate high-quality images, posters, and logos with Ideogram V2A. Features exceptional typography handling and realistic outputs optimized for commercial and creative use.
text-to-imagefal-ai/z-image/turbo/tiling/loraGenerate seamlessly tiling photorealistic images from text using Z-Image Turbo and custom LoRA
image-to-videomoonvalley/marey/i2vGenerate a video starting from an image as the first frame with Marey, a generative video model trained exclusively on fully licensed data.
image-to-imagefal-ai/image-editing/professional-photoTurn your casual photos into stunning professional studio portraits with perfect lighting and high-end photography style.
visionfal-ai/sam-3/image/embedSAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.
visionfal-ai/sa2va/8b/imageSa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels
text-to-image
text-to-videofal-ai/longcat-video/text-to-video/720pGenerate long videos in 720p/30fps from text using LongCat Video