Models

Speech to Text

View as Markdown

The Speech to Text API transcribes audio into text. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-2.0. Use the REST endpoint for file-based batch transcription, or the streaming endpoint for real-time low-latency transcription.


At a glance

Details
ModalitiesAudio → Text
REST pricing/ hr
Streaming pricing/ hr
Regionus-east-1

Pricing

Details
REST (per hour)/ hr
Streaming (per hour)/ hr

Rate Limits

RESTStreaming
RPS (Requests per second)
Concurrent sessionsper team

Capabilities

  • REST and streaming transcription

  • Multiple audio formats (WAV, MP3, WebM, OGG, M4A)

  • Multiple languages

  • Real-time interim results (streaming)

  • Keyterm prompting for domain-specific vocabulary

  • Smart Turn end-of-turn detection (streaming) — ML-based prediction of whether the speaker has finished their thought


Availability

Details
Clusterus-east-1

Documentation


Last updated:September 18, 2026