image-to-imagePuLID Flux
fal-ai/flux-pulidAn endpoint for personalized image generation using Flux as per given description.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/flux-pulidAn endpoint for personalized image generation using Flux as per given description.
text-to-imagefal-ai/qwen-imageQwen-Image is an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing.
image-to-videofal-ai/kling-video/v1.6/standard/image-to-videoGenerate video clips from your images using Kling 1.6 (std)
image-to-imagefal-ai/flux-pro/kontext/max/multiExperimental version of FLUX.1 Kontext [max] with multi image handling capabilities
image-to-imageclarityai/crystal-upscalerAn advanced image enhancement tool designed specifically for facial details and portrait photography, utilizing Clarity AI's upscaling technology.
video-to-videofal-ai/sync-lipsync/v2/proGenerate high-quality realistic lipsync animations from audio while preserving unique details like natural teeth and unique facial features using the state-of-the-art Sync Lipsync 2 Pro model.
visionfal-ai/imageutils/nsfwPredict the probability of an image being NSFW.
video-to-videofal-ai/seedvr/upscale/videoUpscale your videos using SeedVR2 with temporal consistency!
text-to-videogoogle/gemini-omni-flash/v1.1/text-to-videoGemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control expressed in natural language.
text-to-audiofal-ai/lyria3/proLyria 3 Pro is the latest music model from Google
video-to-videoblackforestlabs/flux-3/edit-videoFLUX.3 Edit Video [FAST] is Black Forest Labs' frontier video model. This endpoint edits an existing video from natural-language instructions, applying targeted changes while preserving the rest of the scene.
image-to-imagebytedance/seedream/v5/pro/layerizeSplits a finished image into independent, editable transparent-PNG layers — background plus separate elements, from a text description, returning 2 to 17 layers per call for non-destructive reuse in design tools.
image-to-imagefal-ai/flux-2/turbo/editImage-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control—all at turbo speed.
text-to-imagefal-ai/flux-2/klein/4bText-to-image generation with FLUX.2 [klein] 4B from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
image-to-imagefal-ai/flux-2/klein/4b/editImage-to-image editing with FLUX.2 [klein] 4B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
text-to-imagealibaba/qwen-image-3/text-to-imageGenerates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence
image-to-videogoogle/gemini-omni-flash/reference-to-videoGenerates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.
image-to-imagefal-ai/sam2/imageSAM 2 is a model for segmenting images and videos in real-time.
video-to-videofal-ai/workflow-utilities/trim-videoFFMPEG Utility for Trim Video
video-to-videofal-ai/kling-video/v2.6/pro/motion-controlTransfer movements from a reference video to any character image. Pro mode delivers higher quality output, ideal for complex dance moves and gestures.
text-to-videoblackforestlabs/flux-3/text-to-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.
text-to-videofal-ai/veo3.1/liteVeo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video
text-to-imagefal-ai/recraft/v4/text-to-imageRecraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.
image-to-imagefal-ai/fashn/tryon/v1.6FASHN v1.6 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 864x1296 resolution from both on-model and flat-lay photo references.
video-to-videofal-ai/mmaudio-v2MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.
image-to-videofal-ai/veo3.1/fast/first-last-frame-to-videoGenerate videos from a first/last frame using Google's Veo 3.1 Fast
image-to-videofal-ai/veo3.1/reference-to-videoGenerate Videos from images using Google's Veo 3.1
text-to-speechfal-ai/minimax/speech-02-turboGenerate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.