image-to-3dHi3D Multiview to 3D
hitem3d/hi3d/v3.0/multi-view-to-3dGenerate 3D models from multiple view images using Hi3D V3.0.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-3dhitem3d/hi3d/v3.0/multi-view-to-3dGenerate 3D models from multiple view images using Hi3D V3.0.
visionfal-ai/florence-2-large/more-detailed-captionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
image-to-imagefal-ai/florence-2-large/object-detectionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
text-to-imagefal-ai/stable-diffusion-v35-largeStable Diffusion 3.5 Large is a Multimodal Diffusion Transformer (MMDiT) text-to-image model that features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency.
image-to-imagefal-ai/feynobgFeyNobg is a state of the art AI model for background removal from feyninc
image-to-imagefal-ai/smart-resizeSmart image resize to arbitrary dimensions, powered by Nano Banana Pro with vision-LLM-guided prompting for composition-aware recomposition. Crop, cropping, resize ads.
image-to-videofal-ai/pika/v2.2/image-to-videoTurn photos into mind-blowing, dynamic videos in up to 1080p. Experience better image clarity and crisper, sharper visuals.
image-to-imagefal-ai/finegrain-eraser/maskFinegrain Eraser removes any object selected with a mask—along with its shadows, reflections, and lighting artifacts—seamlessly reconstructing the scene with contextually accurate content.
image-to-videoalibaba/happy-horse/v1.1/reference-to-videoHappy Horse 1.1 is Alibaba's #1-ranked video model. This reference-to-video endpoint turns up to 9 reference images into 1080p video with synchronized native audio and multilingual lip-sync for consistent characters.
text-to-speechfal-ai/bytedance/seed-speech/tts/v2Seed Speech developed by ByteDance, is a family of large-scale text-to-speech models capable of synthesizing speech that is virtually indistinguishable from human speech.
text-to-3dmeshy/v7/text-to-3dTurns text into a fully textured, PBR-ready 3D mesh with complete geometry, in game-ready Smart Topology at a target polygon count
image-to-imagefal-ai/ideogram/v3/editTransform existing images with Ideogram V3's editing capabilities. Modify, adjust, and refine images while maintaining high fidelity and realistic outputs with precise prompt control.
image-to-imagefal-ai/recraft/v3/image-to-imageRecraft V3 is a text-to-image model with the ability to generate long texts, vector art, images in brand style, and much more. As of today, it is SOTA in image generation, proven by Hugging Face's industry-leading Text-to-Image Benchmark by Artificial Analysis.
image-to-videofal-ai/pixverse/v6/transitionPixverse's latest v6 Model.
image-to-videofal-ai/pixverse/c1/image-to-videoAnimate images into cinematic videos with PixVerse C1, supporting 1080p resolution and native audio generation.
image-to-imagefal-ai/flux-general/image-to-imageFLUX General Image-to-Image is a versatile endpoint that transforms existing images with support for LoRA, ControlNet, and IP-Adapter extensions, enabling precise control over style transfer, modifications, and artistic variations through multiple guidance methods.
text-to-3dtripo3d/h3.1/text-to-3dGenerate 3D models from text descriptions using Tripo H3.1.
image-to-video
image-to-imagefal-ai/gpt-image-1-mini/editGPT Image 1 mini combines OpenAI's advanced language capabilities, powered by GPT-5, with GPT Image 1 Mini for efficient image generation.
visionfal-ai/moondream2/visual-queryMoondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.
image-to-videofal-ai/luma-dream-machine/ray-2/image-to-videoRay2 is a large-scale video generative model capable of creating realistic visuals with natural, coherent motion.
text-to-speechresemble-ai/chatterboxhd/text-to-speechGenerate expressive, natural speech with Resemble AI's Chatterbox. Features unique emotion control, instant voice cloning from short audio, and built-in watermarking.
text-to-audiofal-ai/minimax-musicGenerate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
image-to-imagefal-ai/image-apps-v2/virtual-try-onTry on clothes virtually by combining person and clothing images.
text-to-videofal-ai/ltx-2.3/text-to-videoLTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.
image-to-video
video-to-videoblackforestlabs/flux-3/extend-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint continues an existing clip beyond its final frame, generating additional footage that stays consistent with the original motion and scene.
text-to-imagefal-ai/z-image/baseZ-Image is the foundation model of the Z- Image family, engineered for good quality, robust generative diversity, broad stylistic coverage, and precise prompt adherence.