Best Speech-to-Text API 2026: Benchmarks & Prices
Choosing a speech-to-text API in late 2026 means choosing between eight serious options — up from two last year. Prices, flagships, and benchmarks all moved this year, so here is the whole field in one table, with picks by use case.
Figures verified September 2026 against vendor pricing pages and published benchmarks. Vendor-published WER numbers are labeled as such — reproduce on your audio before committing.
The Full Comparison Table
| API / Model | Async price | Streaming price | Published accuracy signal | BAA | Best for |
|---|---|---|---|---|---|
| AssemblyAI Universal-3.5 Pro | $0.21/hr flat | $0.45/hr (RT); $4.50/hr agent API | 7.69% avg WER; 6.99% realtime (vendor) | Standard, self-serve | Transcription products, medical, entities |
| Deepgram Nova-3 / Flux | $0.26–0.29/hr (promo lower) | $0.35/hr streaming; Flux EN $0.39/hr | Flux 15.58% realtime WER (vendor turf) | Enterprise tier | Voice bots, batch at volume |
| ElevenLabs Scribe v2 | $0.22/hr + add-ons | $0.39/hr Realtime | 5.87% EN WER; 5.5% independent (Coval) | Enterprise-gated | ElevenLabs TTS users, media pipelines |
| Gemini 3.5 Transcribe | Token-based (check Google) | Live variant available | No independent numbers yet | Cloud enterprise path | Google Cloud shops, polished multilingual output |
| Microsoft MAI-Transcribe-1.5 | Foundry pricing | Via Azure Speech paths | 43 langs, entity biasing; no public benchmarks | Azure inherited | Foundry/Azure-standardized teams |
| Soniox v5 | $0.12/hr | Same platform card | Structured-output positioning; verify independently | Check vendor | Entity-dense audio on a budget |
| OpenAI GPT-4o Transcribe | $0.36/hr ($0.006/min) | Via Realtime-Whisper (usage) | Reference baseline, not leader | Enterprise/DPA path | OpenAI-native stacks |
| Whisper Large v3 (open) | Free (self-hosted) | Community streaming builds | The baseline everything cites | You are the BAA | Prototypes, offline, cost-zero |
Picks by Use Case
- Voice agent where interruptions matter: Deepgram Flux for turn-taking, but benchmark Universal-3.5 Pro Realtime on entities first — 15.58% vs 6.99% realtime WER is too big to ignore.
- Transcription product (Otter-style): Universal-3.5 Pro at $0.21/hr flat. Cheapest flagship async with the best published accuracy.
- Medical / ambient scribe: AssemblyAI with Medical Mode ($0.36/hr async, 3.2% published missed-entity rate). See our medical head-to-head.
- Media, podcasts, dubbing: Scribe v2 (timestamps + 99 languages + Dubbing in one house) or Gemini 3.5 Transcribe for pre-polished output.
- Tightest budget: Soniox v5 at $0.12/hr — if v5 benchmarks verify on your audio type.
- Google Cloud shop: Gemini 3.5 Transcribe. Azure shop: MAI-Transcribe-1.5 or classic Speech for 140+ languages.
- OpenAI-native agent: GPT-Realtime-Whisper keeps STT→reasoning in one API.
How to Run Your Own Bake-Off
- Collect 500 representative clips (your accents, your noise, your jargon).
- Score entity accuracy (names, numbers, drug/part codes) separately from overall WER — entities decide production quality.
- Measure first-token latency on live traffic, not finals — Scribe v2 Realtime's 123ms finals hide a 2.1s first token.
- Price the stacked cost (diarization, redaction, Medical Mode), not the headline rate.
FAQ
What is the cheapest speech-to-text API in 2026? Soniox v5 lists $0.12/hr; among flagships, Universal-3.5 Pro at $0.21/hr flat undercuts Nova-3 pay-as-you-go. Deepgram Growth commits and promos can go lower at volume.
What is the most accurate STT API? On published numbers, Universal-3.5 Pro (7.69% avg WER, 6.99% realtime). Independently, Scribe v2 Realtime tests at 5.5% WER on Coval's suite. Different test sets — run your own bake-off.
Which STT API is best for voice agents? No single winner: Flux for turn-taking, Universal-3.5 Pro RT for entity accuracy, GPT-Realtime-Whisper for OpenAI-native stacks. See the per-model reviews below.
In-Depth Reviews
- Deepgram Nova-3 Review 2026: Speed, Accuracy & Pricing
- Deepgram vs. AssemblyAI (2026): Which STT API Should You Choose?
- ElevenLabs Scribe v2 Review 2026: Realtime & Pricing
- Gemini 3.5 Transcribe Review: Smart STT Tested (2026)
- Microsoft MAI-Transcribe-1.5 Review: In-House STT
- OpenAI GPT-Realtime-Whisper: Streaming STT Guide
- Soniox Review 2026: Accuracy, Pricing & Alternatives
- AssemblyAI Pricing 2026: $0.21/hr Plans Explained
Keep reading
Voice Model Deep Dives
ElevenLabs Scribe v2 Review 2026: Realtime & Pricing
Scribe v2 at $0.22/hr plus a sub-150ms Realtime model. Benchmarks, pricing vs Deepgram and AssemblyAI, and the diarization catch.
Voice Model Deep Dives
Fastest Speech-to-Text 2026: Latency Benchmarks
Who is actually fastest? First-token, final-segment and endpointing numbers for Flux, Nova-3, 3.5 Pro RT, Scribe v2 and Parakeet.
Voice Model Deep Dives
Gemini 3.5 Transcribe Review: Smart STT Tested (2026)
Google’s Aug 2026 model polishes ramblings into formatted text. 85+ languages, smart transcription, and how it compares to Nova-3.
