EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

speech to text

Page 1 · 28 per page
speech-to-text
falREVIEW REQUIRED

Wizper (Whisper v3 -- fal.ai edition)

fal-ai/wizper

[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!

transcriptionspeech
speech-to-text
ElevenLabsREVIEW REQUIRED

ElevenLabs Speech to Text - Scribe V2

fal-ai/elevenlabs/speech-to-text/scribe-v2

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

speech-to-text
speech-to-text
ElevenLabsREVIEW REQUIRED

ElevenLabs Speech to Text

fal-ai/elevenlabs/speech-to-text

Generate text from speech using ElevenLabs advanced speech-to-text model.

speech
speech-to-text
ElevenLabsREVIEW REQUIRED

Elevenlabs - Forced Alignment

fal-ai/elevenlabs/forced-alignment

Align the transcript and your audio recording using Elevenlab's forced alignment feature!

forced-alignmentspeech-to-text
speech-to-text
falREVIEW REQUIRED

Speech-to-Text

fal-ai/speech-to-text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

speech-to-text
falREVIEW REQUIRED

Cohere Transcribe

fal-ai/cohere-transcribe

Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

speechtranscribestt
speech-to-text
nvidiaREVIEW REQUIRED

Nemotron Asr Multilingual

nvidia/nemotron-asr-multilingual/asr

Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.

utilitytranscribe
speech-to-text
falREVIEW REQUIRED

Speech-to-Text

fal-ai/speech-to-text/turbo

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

speech-to-text
falREVIEW REQUIRED

Pipecat's Smart Turn model

fal-ai/smart-turn

An open source, community-driven and native audio turn detection model by Pipecat AI.

speech-to-text
falREVIEW REQUIRED

Speech-To-text

fal-ai/speech-to-text/stream

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

streaming
speech-to-text
falREVIEW REQUIRED

Speech-to-Text

fal-ai/speech-to-text/turbo/stream

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

streaming