image-to-imageFILM
fal-ai/filmInterpolate images with FILM - Frame Interpolation for Large Motion
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/filmInterpolate images with FILM - Frame Interpolation for Large Motion
video-to-videoclarityai/crystal-video-upscalerDo high precision video upscaling that respects the original video perfectly using Crystal Upscaler's new video upscaling method!
video-to-videofal-ai/heygen/v2/translate/precisionHeygen Translate Model with Extreme Precision
trainingfal-ai/z-image-turbo-trainer-v2Fast LoRA trainer for Z-Image-Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
text-to-videominimax/h3/text-to-video/loraGenerate video with synchronized audio from a text prompt using MiniMax H3; load a trained LoRA at adjustable strength to lock in style, character, or motion.
image-to-imagefal-ai/image-apps-v2/relightingAdjust and enhance images with different lighting styles.
image-to-imagefal-ai/flux-2/klein/4b/base/edit/loraImage-to-image editing with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.
image-to-videofal-ai/live-portraitTransfer expression from a video to a portrait.
video-to-video
image-to-imagefal-ai/drct-super-resolutionUpscale your images with DRCT-Super-Resolution.
text-to-videofal-ai/hunyuan-videoHunyuan Video is an Open video generation model with high visual quality, motion diversity, text-video alignment, and generation stability. This endpoint generates videos from text descriptions.
audio-to-audiofal-ai/ace-step/audio-inpaintModify a portion of provided audio with lyrics and/or style using ACE-Step
video-to-videofal-ai/hunyuan-video-foleyUse the capabilities of the hunyuan foley model to bring life to your videos by adding sound effect to them.
video-to-videofal-ai/heygen/v2/translate/speedHeygen Translate Model with Extreme Speed
text-to-speechfal-ai/dia-ttsDia directly generates realistic dialogue from transcripts. Audio conditioning enables emotion control. Produces natural nonverbals like laughter and throat clearing.
audio-to-videolightricks/ltx-2.5/audio-to-video/proLTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates video timed to a supplied audio clip in a quality-optimized mode, for final visuals synchronized to music, dialogue, or a soundtrack.
video-to-videotopaz/interpolate/videoProfessional frame interpolation powered by Topaz Labs. Apollo, Chronos and Aion retime footage up to 120 fps, from smooth motion to extreme slow motion. Best for fluid 60fps output and slow-motion effects.
video-to-videofal-ai/wan-vace-14b/inpaintingVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
image-to-imagefal-ai/flux-vision-upscalerFlux Vision Upscaler for magnify/upscaling images with high fidelity and creativity.
image-to-imagefal-ai/hunyuan_worldHunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.
text-to-imagenvidia/cosmos-3-super/text-to-imageCosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.
text-to-imageimagineart/imagineart-2.0-preview/text-to-imageImagineArt 2.0 is ImagineArt's latest state-of-the-art visual reasoning text-to-image model, generating high-fidelity, professional-grade visuals with lifelike realism, cinematic effects, and strong aesthetic quality.
visionfal-ai/florence-2-large/detailed-captionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
video-to-videofal-ai/wan-vace-14bVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
text-to-imagefal-ai/z-image/turbo/tilingGenerate seamlessly tiling photorealistic images from text using Z-Image Turbo
image-to-videofal-ai/minimax/video-01-live/image-to-videoGenerate video clips from your images using MiniMax Video model
text-to-videofal-ai/minimax/hailuo-02/pro/text-to-videoMiniMax Hailuo-02 Text To Video API (Pro, 1080p): Advanced video generation model with 1080p resolution
text-to-audiofal-ai/diffrhythmDiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.