Models
Speech to Text
The Speech to Text API transcribes audio into text. Use grok-voice-transcribe-1.0 or grok-voice-transcribe-2.0; the default is grok-voice-transcribe-2.0. Use the REST endpoint for file-based batch transcription, or the streaming endpoint for real-time low-latency transcription.
At a glance
| Details | |
|---|---|
| Modalities | Audio → Text |
| REST pricing | / hr |
| Streaming pricing | / hr |
| Region | us-east-1 |
Pricing
| Details | |
|---|---|
| REST (per hour) | / hr |
| Streaming (per hour) | / hr |
Rate Limits
| REST | Streaming | |
|---|---|---|
| RPS (Requests per second) | ||
| Concurrent sessions | — | per team |
Capabilities
REST and streaming transcription
Multiple audio formats (WAV, MP3, WebM, OGG, M4A)
Multiple languages
Real-time interim results (streaming)
Keyterm prompting for domain-specific vocabulary
Smart Turn end-of-turn detection (streaming) — ML-based prediction of whether the speaker has finished their thought
Availability
| Details | |
|---|---|
| Cluster | us-east-1 |
Documentation
Speech to Text Guide — Getting started with speech to text
Voice Overview — Overview of all voice capabilities
Pricing — Full pricing overview
Last updated:September 18, 2026