Speech-to-speech model with parallel reasoning reduces latency by thinking while speaking; tool calls fire before agent finishes first sentence.
Summary
Eliminates reasoning latency in voice agents and handles real-world audio conditions (background noise, telephony compression) without transcription degradation. Fewer reasoning tokens means faster tool invocation in production voice flows.
Why it matters
Eliminates reasoning latency in voice agents and handles real-world audio conditions (background noise, telephony compression) without transcription degradation. Fewer reasoning tokens means faster tool invocation in production voice flows.
Implementation verdict
Replaces previous Grok Voice for real-time voice applications. Requires short-lived token server (provided pattern) and AI SDK's realtime API integration. Ready now—deploy via `xai/grok-voice-think-fast-2.0` model identifier through Vercel AI Gateway.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.