EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

text to audio

Page 1 · 28 per page
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Tts Eleven V3

fal-ai/elevenlabs/tts/eleven-v3

Generate text-to-speech audio using Eleven-v3 from ElevenLabs.

audio
text-to-audio
ElevenLabsREVIEW REQUIRED

ElevenLabs TTS Multilingual v2

fal-ai/elevenlabs/tts/multilingual-v2

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.

audio
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Sound Effects V2

fal-ai/elevenlabs/sound-effects/v2

Generate sound effects using ElevenLabs advanced sound effects model.

sound
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs Music

fal-ai/elevenlabs/music

Generate high quality, realistic music with fine controls using Elevenlabs Music!

musictext-to-music
text-to-audio
MiniMaxREVIEW REQUIRED

MiniMax Music 3

minimax/music-3

MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long

sfxaudioeffects
text-to-audio
falREVIEW REQUIRED

Stable Audio 2.5

fal-ai/stable-audio-25/text-to-audio

Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI

audio
text-to-audio
MiniMaxREVIEW REQUIRED

Minimax Music 2.6

fal-ai/minimax-music/v2.6

MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

stylizedtransformlipsync
text-to-audio
cassetteaiREVIEW REQUIRED

music generator

cassetteai/music-generator

CassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds. At 44.1 kHz stereo audio, expect a level of professional consistency with no breaks, no squeaks, and no random interruptions in your creations.

musiccassetteai
text-to-audio
ByteDanceREVIEW REQUIRED

Seed Audio 1.0

bytedance/seed-audio-1.0

Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.

text-to-audio
MiniMaxREVIEW REQUIRED

Minimax Music

fal-ai/minimax-music/v2

Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

musicaudio
text-to-audio
falREVIEW REQUIRED

Lyria 3 Pro

fal-ai/lyria3/pro

Lyria 3 Pro is the latest music model from Google

audiosfx
text-to-audio
falREVIEW REQUIRED

ACE Step

fal-ai/ace-step

Generate music with lyrics from text using ACE-Step

text-to-audiotext-to-music
text-to-audio
falREVIEW REQUIRED

Lyria2

fal-ai/lyria2

Lyria 2 is Google's latest music generation model, you can generate any type of music with this model.

musicstylized
text-to-audio
falREVIEW REQUIRED

Stable Audio 3

fal-ai/stable-audio-3/medium/text-to-audio

Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.

musicaudiostereo
text-to-audio
falREVIEW REQUIRED

Kokoro TTS

fal-ai/kokoro/american-english

Kokoro is a lightweight text-to-speech model that delivers comparable quality to larger models while being significantly faster and more cost-efficient.

speech
text-to-audio
soniloREVIEW REQUIRED

Sonilo V1.1 Text to Music

sonilo/v1.1/text-to-music

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.

stylizedtransformlipsync
text-to-audio
soniloREVIEW REQUIRED

V1.1 Text to Sound Effects

sonilo/v1.1/text-to-sound-effects

Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.

sfxaudioeffects
text-to-audio
ElevenLabsREVIEW REQUIRED

Elevenlabs

fal-ai/elevenlabs/text-to-dialogue/eleven-v3

Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.

audio
text-to-audio
cassetteaiREVIEW REQUIRED

Sound Effects Generator

cassetteai/sound-effects-generator

Create stunningly realistic sound effects in seconds - CassetteAI's Sound Effects Model generates high-quality SFX up to 30 seconds long in just 1 second of processing time

soundsfxsound-effectscassetteai
text-to-audio
falREVIEW REQUIRED

ACE Step Prompt To Audio

fal-ai/ace-step/prompt-to-audio

Generate music from a simple prompt using ACE-Step

text-to-audiotext-to-music
text-to-audio
GoogleREVIEW REQUIRED

Gemini TTS

fal-ai/gemini-tts

Use Gemini TTS Models to convert your prompts to real audio.

text-to-speechaudiogemini
text-to-audio
MiniMaxREVIEW REQUIRED

MiniMax (Hailuo AI) Music

fal-ai/minimax-music

Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

music
text-to-audio
falREVIEW REQUIRED

Stable Audio 3 Small Music Text to Audio

fal-ai/stable-audio-3/small/music/text-to-audio

Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.

musicon-devicelightweight
text-to-audio
falREVIEW REQUIRED

Stable Audio 3 Small SFX Text to Audio

fal-ai/stable-audio-3/small/sfx/text-to-audio

Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts, designed for on-device deployment on mobile phones and consumer laptops.

sfxsound-effectson-device
text-to-audio
falREVIEW REQUIRED

MMAudio V2 Text to Audio

fal-ai/mmaudio-v2/text-to-audio

MMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.

audiofast