text-to-audioMinimax Music 2.5
fal-ai/minimax-music/v2.5MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-audiofal-ai/minimax-music/v2.5MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
text-to-audiomirelo-ai/sfx1.6/text-to-audioGenerate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.
text-to-audiofal-ai/minimax-music/v1.5Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
text-to-audiofal-ai/kokoro/british-englishA high-quality British English text-to-speech model offering natural and expressive voice synthesis.
text-to-audiofal-ai/diffrhythmDiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.
text-to-audiofal-ai/stable-audio-3/medium/base/text-to-audioStable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.
text-to-audiofal-ai/stable-audio-3/small/music/base/text-to-audioStable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.
text-to-audiofal-ai/stable-audio-3/small/sfx/base/text-to-audioStable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.
text-to-audiofal-ai/kokoro/spanishA natural-sounding Spanish text-to-speech model optimized for Latin American and European Spanish.
text-to-audiofal-ai/kokoro/frenchAn expressive and natural French text-to-speech model for both European and Canadian French.
text-to-audiofal-ai/kokoro/brazilian-portugueseA natural and expressive Brazilian Portuguese text-to-speech model optimized for clarity and fluency.
text-to-audiofal-ai/kokoro/italianA high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.
text-to-audiofal-ai/kokoro/japaneseA fast and natural-sounding Japanese text-to-speech model optimized for smooth pronunciation.
text-to-audiofal-ai/csm-1bCSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.
text-to-audiofal-ai/zonosClone voice of any person and speak anything in their voice using zonos' voice cloning.
text-to-audiofal-ai/kokoro/mandarin-chineseA highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.
text-to-audiofal-ai/kokoro/hindiA fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.
text-to-audiofal-ai/ltx-2.3-quality/text-to-audioText to Audio high-quality using LTX-2.3
text-to-audiofal-ai/ltx-2.3-quality/text-to-audio/loraText to Audio high-quality using LTX-2.3 with Lora
text-to-audioelevenlabs/music/v2.5Generate high quality, realistic music with fine controls using Elevenlabs Music v2.5!
text-to-audioelevenlabs/music/v2Generate high quality, realistic music with fine controls using Elevenlabs Music v2!