Voice Model Deep Dives•6 min read•

Best Speech-to-Text API 2026: Benchmarks & Prices

Choosing a speech-to-text API in late 2026 means choosing between eight serious options — up from two last year. Prices, flagships, and benchmarks all moved this year, so here is the whole field in one table, with picks by use case.

Figures verified September 2026 against vendor pricing pages and published benchmarks. Vendor-published WER numbers are labeled as such — reproduce on your audio before committing.

The Full Comparison Table

API / Model Async price Streaming price Published accuracy signal BAA Best for
AssemblyAI Universal-3.5 Pro $0.21/hr flat $0.45/hr (RT); $4.50/hr agent API 7.69% avg WER; 6.99% realtime (vendor) Standard, self-serve Transcription products, medical, entities
Deepgram Nova-3 / Flux $0.26–0.29/hr (promo lower) $0.35/hr streaming; Flux EN $0.39/hr Flux 15.58% realtime WER (vendor turf) Enterprise tier Voice bots, batch at volume
ElevenLabs Scribe v2 $0.22/hr + add-ons $0.39/hr Realtime 5.87% EN WER; 5.5% independent (Coval) Enterprise-gated ElevenLabs TTS users, media pipelines
Gemini 3.5 Transcribe Token-based (check Google) Live variant available No independent numbers yet Cloud enterprise path Google Cloud shops, polished multilingual output
Microsoft MAI-Transcribe-1.5 Foundry pricing Via Azure Speech paths 43 langs, entity biasing; no public benchmarks Azure inherited Foundry/Azure-standardized teams
Soniox v5 $0.12/hr Same platform card Structured-output positioning; verify independently Check vendor Entity-dense audio on a budget
OpenAI GPT-4o Transcribe $0.36/hr ($0.006/min) Via Realtime-Whisper (usage) Reference baseline, not leader Enterprise/DPA path OpenAI-native stacks
Whisper Large v3 (open) Free (self-hosted) Community streaming builds The baseline everything cites You are the BAA Prototypes, offline, cost-zero

Picks by Use Case

  • Voice agent where interruptions matter: Deepgram Flux for turn-taking, but benchmark Universal-3.5 Pro Realtime on entities first — 15.58% vs 6.99% realtime WER is too big to ignore.
  • Transcription product (Otter-style): Universal-3.5 Pro at $0.21/hr flat. Cheapest flagship async with the best published accuracy.
  • Medical / ambient scribe: AssemblyAI with Medical Mode ($0.36/hr async, 3.2% published missed-entity rate). See our medical head-to-head.
  • Media, podcasts, dubbing: Scribe v2 (timestamps + 99 languages + Dubbing in one house) or Gemini 3.5 Transcribe for pre-polished output.
  • Tightest budget: Soniox v5 at $0.12/hr — if v5 benchmarks verify on your audio type.
  • Google Cloud shop: Gemini 3.5 Transcribe. Azure shop: MAI-Transcribe-1.5 or classic Speech for 140+ languages.
  • OpenAI-native agent: GPT-Realtime-Whisper keeps STT→reasoning in one API.

How to Run Your Own Bake-Off

  1. Collect 500 representative clips (your accents, your noise, your jargon).
  2. Score entity accuracy (names, numbers, drug/part codes) separately from overall WER — entities decide production quality.
  3. Measure first-token latency on live traffic, not finals — Scribe v2 Realtime's 123ms finals hide a 2.1s first token.
  4. Price the stacked cost (diarization, redaction, Medical Mode), not the headline rate.

FAQ

What is the cheapest speech-to-text API in 2026? Soniox v5 lists $0.12/hr; among flagships, Universal-3.5 Pro at $0.21/hr flat undercuts Nova-3 pay-as-you-go. Deepgram Growth commits and promos can go lower at volume.

What is the most accurate STT API? On published numbers, Universal-3.5 Pro (7.69% avg WER, 6.99% realtime). Independently, Scribe v2 Realtime tests at 5.5% WER on Coval's suite. Different test sets — run your own bake-off.

Which STT API is best for voice agents? No single winner: Flux for turn-taking, Universal-3.5 Pro RT for entity accuracy, GPT-Realtime-Whisper for OpenAI-native stacks. See the per-model reviews below.

In-Depth Reviews

Keep reading