The new speech-to-text model is built on the audio model behind Grok Voice, detects the language itself and follows a switch mid-recording. The price stays at $0.10 per hour of audio for batch and $0.20 for streaming.
- Per the company, the model ranks first for accuracy among 32 streaming transcription models on the Artificial Analysis leaderboard.
- On short voice commands in their internal test the word error rate falls from 20.6 to 6.8 per cent.
- Bulgarian is not in their language chart.
An account number dictated over the phone with noise in the background. An email address said in a hurry. That is exactly where transcription is hardest, and exactly where a mistake costs the most.
SpaceXAI measures right at that spot, and says so.
Price is the strong part. Streaming transcription is about half of what they show for ElevenLabs and Deepgram in their own chart. For a call centre with thousands of hours a month, that is already a budget line, not a rounding error.
The ranking, though, is their citation of someone else's leaderboard, and the language chart cites no outside source. And one thing that matters to us: Bulgarian is not in the language chart. That does not mean it cannot transcribe it. It means they do not measure it in front of us.
For conversations in Bulgarian, the only honest test is your own recording with your own noise. Run one hour and count the mistakes in names and numbers.