EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

text to audio

Page 2 · 28 per page
text-to-audio
MiniMaxREVIEW REQUIRED

Minimax Music 2.5

fal-ai/minimax-music/v2.5

MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

stylizedtransformlipsync
text-to-audio
mirelo-aiREVIEW REQUIRED

Mirelo SFX1.6

mirelo-ai/sfx1.6/text-to-audio

Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.

text-to-audiosfx
text-to-audio
MiniMaxREVIEW REQUIRED

MiniMax (Hailuo AI) Music v1.5

fal-ai/minimax-music/v1.5

Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

music
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (British English)

fal-ai/kokoro/british-english

A high-quality British English text-to-speech model offering natural and expressive voice synthesis.

speech
text-to-audio
falREVIEW REQUIRED

DiffRhythm: Lyrics to Song

fal-ai/diffrhythm

DiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.

music
text-to-audio
falREVIEW REQUIRED

Stable Audio 3 Medium Base Text to Audio

fal-ai/stable-audio-3/medium/base/text-to-audio

Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.

musicaudiostereo
text-to-audio
falREVIEW REQUIRED

Stable Audio 3

fal-ai/stable-audio-3/small/music/base/text-to-audio

Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.

musicon-devicelightweight
text-to-audio
falREVIEW REQUIRED

Stable Audio 3 Small SFX Base Text to Audio

fal-ai/stable-audio-3/small/sfx/base/text-to-audio

Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.

sfxsound-effectson-device
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Spanish)

fal-ai/kokoro/spanish

A natural-sounding Spanish text-to-speech model optimized for Latin American and European Spanish.

speech
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (French)

fal-ai/kokoro/french

An expressive and natural French text-to-speech model for both European and Canadian French.

speech
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Brazilian Portuguese)

fal-ai/kokoro/brazilian-portuguese

A natural and expressive Brazilian Portuguese text-to-speech model optimized for clarity and fluency.

speech
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Italian)

fal-ai/kokoro/italian

A high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.

speech
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Japanese)

fal-ai/kokoro/japanese

A fast and natural-sounding Japanese text-to-speech model optimized for smooth pronunciation.

speech
text-to-audio
falREVIEW REQUIRED

CSM-1B

fal-ai/csm-1b

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.

conversationaltext to speech
text-to-audio
falREVIEW REQUIRED

Zonos-Audio-Clone

fal-ai/zonos

Clone voice of any person and speak anything in their voice using zonos' voice cloning.

voice cloning
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Mandarin Chinese)

fal-ai/kokoro/mandarin-chinese

A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.

speech
text-to-audio
falREVIEW REQUIRED

Kokoro TTS (Hindi)

fal-ai/kokoro/hindi

A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.

speech
text-to-audio
falREVIEW REQUIRED

Ltx 2.3 Quality

fal-ai/ltx-2.3-quality/text-to-audio

Text to Audio high-quality using LTX-2.3

text-to-audio
text-to-audio
falREVIEW REQUIRED

Ltx 2.3 Quality

fal-ai/ltx-2.3-quality/text-to-audio/lora

Text to Audio high-quality using LTX-2.3 with Lora

text-to-audio
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Music v2.5

elevenlabs/music/v2.5

Generate high quality, realistic music with fine controls using Elevenlabs Music v2.5!

musictext-to-music
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Music v2

elevenlabs/music/v2

Generate high quality, realistic music with fine controls using Elevenlabs Music v2!

musictext-to-music