visionMarlin
fal-ai/marlinMarlin is a 2B video VLM tuned for the two questions developers actually want to ask of their videos: what is happening, and when?
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
visionfal-ai/marlinMarlin is a 2B video VLM tuned for the two questions developers actually want to ask of their videos: what is happening, and when?
text-to-videofal-ai/pixverse/v5.6/text-to-videoUse the latest pixverse v5.6 model to turn your texts into amazing videos.
text-to-videofal-ai/ltx-2.3-22b/distilled/text-to-videoGenerate video with audio from text using LTX-2.3 Distilled
speech-to-textfal-ai/speech-to-text/turboLeverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
image-to-videofal-ai/vidu/image-to-videoVidu Image to Video generates high-quality videos with exceptional visual quality and motion diversity from a single image
image-to-imagefal-ai/z-image/turbo/controlnet/loraGenerate images from text and edge, depth or pose images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
image-to-imagefal-ai/finegrain-eraserFinegrain Eraser removes objects—along with their shadows, reflections, and lighting artifacts—using only natural language, seamlessly filling the scene with contextually accurate content.
text-to-videofal-ai/pixverse/v4.5/text-to-video/fastGenerate high quality and fast video clips from text and image prompts using PixVerse v4.5 fast
text-to-imagefal-ai/wan/v2.2-a14b/text-to-image/loraWan 2.2's 14B model with LoRA support generates high-fidelity images with enhanced prompt alignment, style adaptability.
video-to-videocassetteai/video-sound-effects-generatorAdd sound effects to your videos
video-to-videofal-ai/wan-vace-14b/reframeVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
image-to-imagefal-ai/image-editing/expression-changeChange facial expressions in photos to any emotion you desire, from smiles to serious looks.
image-to-imagefal-ai/luma-photon/modifyEdit images from your prompts using Luma Photon. Photon is the most creative, personalizable, and intelligent visual models for creatives, bringing a step-function change in the cost of high-quality image generation.
image-to-videofal-ai/minimax/video-01-subject-referenceGenerate video clips maintaining consistent, realistic facial features and identity across dynamic video content
audio-to-videoveed/avatars/audio-to-videoGenerate high-quality videos with UGC-like avatars from audio
image-to-imagefal-ai/flux-1/krea/image-to-imageFLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
text-to-imagefal-ai/sana/v1.5/4.8bSana v1.5 4.8B is a powerful text-to-image model that generates ultra-high quality 4K images with remarkable detail.
image-to-imagefal-ai/image-preprocessors/mlsdM-LSD line segment detection preprocessor.
image-to-imagefal-ai/pasdPixel-Aware Diffusion Model for Realistic Image Super-Resolution and Personalized Stylization
video-to-videofal-ai/bernini-r/reference-edit-videoEdit a video guided by reference images with Bernini-R, bringing an object, material, background, style, or weather from a reference image into your video.
image-to-imagefal-ai/cartoonifyTransform images into 3D cartoon artwork using an AI model that applies cartoon stylization while preserving the original image's composition and details.
image-to-imagefal-ai/wan/v2.2-a14b/image-to-imageWan 2.2's 14B model edit high-resolution, photorealistic images with powerful prompt understanding and fine-grained visual detail
audio-to-audiofal-ai/personaplexPersonaPlex is a real-time, full-duplex speech-to-speech conversational model that enables persona control through text-based role prompts and audio-based voice conditioning.
image-to-videofal-ai/pika/v2.2/pikascenesPika Scenes v2.2 creates videos from a images with high quality output.
image-to-imagefal-ai/qwen-image-edit-plus-lora-gallery/add-backgroundAdd a realistic scene behind the object with white background
video-to-videofal-ai/ltx-2.3-quality/hdrGenerate HDR from reference video using LTX-2.3
image-to-videofal-ai/pixverse/v4/effectsGenerate high quality video clips with different effects using PixVerse v4
image-to-videofal-ai/ai-avatar/multiMultiTalk model generates a multi-person conversation video from an image and audio files. Creates a realistic scene where multiple people speak in sequence.