Save time, make money and get customers with FREE AI! CLICK HERE →

Grok Voice Transcribe 2.0: Twice the Accuracy, Same Price

In this Grok Voice Transcribe 2.0 breakdown, here’s what we’re going to cover: what xAI actually shipped on 18 September 2026 (per the official announcement on x.ai), what it costs, the accuracy claims worth double-checking, and how I’d plug it into an AI content workflow this week. Transcription is one of those boring layers that quietly decides how good your repurposing pipeline is, so a price-flat accuracy jump matters.

Short answer

  • Grok Voice Transcribe 2.0 is xAI’s new speech-to-text model, released 18 September 2026 per the official x.ai announcement.
  • xAI claims twice the accuracy of 1.0 at identical pricing: $0.10/hour batch, $0.20/hour streaming.
  • Diarisation, word-level timestamps and key-term biasing are included, not add-ons.
  • Automatic language detection across dozens of languages, including mid-recording language switching.
  • 1.0 stays available by pinning grok-voice-transcribe-1.0; 2.0 becomes the default, with 1.0 deprecated in the coming weeks.


What’s new in Grok Voice Transcribe 2.0

The headline, straight from the 18 September 2026 announcement: same price, roughly double the accuracy. xAI’s internal evaluations report improvements across telephony, conversational, credentials and multilingual audio, with the most dramatic jump on multilingual short phrases — word error rate falling from 20.6% to 6.8% versus 1.0.

⚠️ Vendor benchmark caveat: the “twice as accurate” figure, the 20.6% → 6.8% word-error-rate numbers and the claimed first-place accuracy ranking among 32 streaming models on the Artificial Analysis leaderboard all come from xAI’s own announcement and internal evaluations. Treat them as the vendor’s claims until you’ve tested your own audio — accents, audio quality and domain jargon move these numbers a lot.

Feature-wise, the announcement lists a genuinely complete package for the base price:

  • Batch and streaming transcription with word-level timestamps and confidence scores
  • Speaker diarisation at no extra charge
  • Multichannel transcription up to 8 channels
  • Key-term biasing: up to 100 domain terms per request, so it stops mangling your product names
  • Text formatting for numbers, dates, currencies, phone numbers and emails
  • Filler-word removal and smart turn detection for voice agents
  • Automatic language detection across dozens of languages, with mid-recording language switching

Grok Voice Transcribe 2.0 pricing and how it compares to 1.0

Pricing is unchanged, which is the whole pitch:

Grok Voice Transcribe 1.0 Grok Voice Transcribe 2.0
Batch $0.10 / hour of audio $0.10 / hour of audio
Streaming $0.20 / hour of audio $0.20 / hour of audio
Accuracy Baseline ~2× better (xAI’s claim)
Diarisation / timestamps / key terms Included Included
Status Deprecated in the coming weeks; pinnable as grok-voice-transcribe-1.0 Becoming the default Speech-to-Text model

Migration is a model-id swap: request grok-voice-transcribe-2.0 now, or wait for it to become the default. Per xAI’s developer release notes, the default is still 1.0 at the time of writing, so pin explicitly if you want 2.0 behaviour today — and pin 1.0 if you need old behaviour to survive the deprecation window.

🔥 Want this set up without the guesswork? Inside the AI Profit Boardroom we build transcription-to-content pipelines exactly like the one below — 3,700+ members, four live calls a week, daily tutorials, done-for-you templates and a 30-day roadmap. Or if you’d rather map your content engine 1-on-1 first, book a free SEO strategy session.

Where it fits in an AI SEO workflow

Here’s why a transcription model launch belongs on an SEO site: at $0.10/hour, transcribing a 100-video YouTube back-catalogue costs a few pounds, and every accurate transcript is raw material for articles, newsletters, FAQs and schema. The workflow I’d run:

  1. Batch-transcribe your video and podcast archive with diarisation on, so speaker turns survive into the transcript.
  2. Load your brand and product names as key terms (up to 100 per request) — this is the difference between a transcript you can publish from and one you babysit.
  3. Feed transcripts to your writing agent to draft posts, pull quotes and build FAQ sections from what you actually said on camera.
  4. For live use, the streaming tier’s smart turn detection is aimed at voice agents — relevant if you’re building call assistants or live-show tooling.

Accuracy compounding is the quiet win: every percentage point of word error rate you remove is one less hallucinated quote downstream. That’s why I’d re-run a sample of old, badly-transcribed audio through 2.0 before anything else — and it’s the kind of leverage-first automation we prioritise in the Boardroom roadmap.

The bottom line on Grok Voice Transcribe 2.0

Free accuracy is the best kind of upgrade: same $0.10/hour batch price, a claimed 2× accuracy jump, and diarisation, timestamps and key-term biasing bundled in. Verify xAI’s numbers on your own audio, pin the model id you want, and if the multilingual gains hold up in your tests, this becomes the default transcription layer for content repurposing at this price point. Want me to look at where transcription-led content fits your site’s strategy? Book a free SEO strategy session — it costs nothing and you’ll leave with a plan.

FAQ: Grok Voice Transcribe 2.0

When was Grok Voice Transcribe 2.0 released?

xAI released it on 18 September 2026, per the official announcement “Introducing Grok Voice Transcribe 2.0” on x.ai.

How much does Grok Voice Transcribe 2.0 cost?

Pricing is unchanged from 1.0: $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming, with diarisation, timestamps and key terms included at no extra cost, per the announcement.

Is Grok Voice Transcribe 2.0 more accurate than 1.0?

xAI claims it is twice as accurate as 1.0 at the same price, and cites a first-place accuracy ranking among 32 streaming models on the Artificial Analysis leaderboard. Those are vendor-reported numbers, so run your own audio through both before migrating anything critical.

What languages does it support?

The announcement describes automatic language detection across dozens of languages, with support for switching languages mid-recording. xAI’s internal evaluations report the multilingual short-phrase word error rate dropping from 20.6% to 6.8% versus 1.0.

How do I use it in the API?

Request the model id grok-voice-transcribe-2.0 through xAI’s Speech-to-Text API. Per xAI’s release notes the current default is still grok-voice-transcribe-1.0, with 2.0 set to become the default; 1.0 will be deprecated in the coming weeks but stays accessible by pinning its id.

Does it do speaker diarisation and timestamps?

Yes — speaker diarisation, word-level timestamps with confidence scores, multichannel transcription up to 8 channels, and key-term biasing of up to 100 domain terms per request are all included in the base price, per the announcement.

Related reading

Next step: turn your existing videos and calls into ranking content — join 3,700+ members inside the AI Profit Boardroom for the pipelines, templates and 30-day roadmap, or book a free SEO strategy session and we’ll build your repurposing engine around your niche.

About the author

Julian Goldie is an SEO agency owner with 10+ years in SEO, 394K+ YouTube subscribers, a 100% Upwork job-success score, 75K+ community members across his groups, and a best-selling SEO book to his name. He tests new AI and SEO tools the week they ship and publishes what actually works.

Watch the latest experiments on YouTube, learn AI SEO alongside 3,700+ members inside the AI Profit Boardroom, or book a free SEO strategy session and get a plan built for your site.

Last updated September 2026. This is the living guide to grok voice transcribe 2.0 — it gets updated as the tools change.