ElevenLabs Scribe v2 Review 2026: Realtime & Pricing
ElevenLabs built its name on text-to-speech. With Scribe v2 — and the newer Scribe v2 Realtime — it now wants both sides of the conversation: the voice that speaks and the ear that listens.
If you already live in ElevenLabs for TTS, dubbing, or voice cloning, Scribe is the obvious STT to evaluate. But how does it hold up against Deepgram Nova-3 and AssemblyAI Universal-3.5 Pro on accuracy, latency, and price?
The Lineup: Scribe v2 vs Scribe v2 Realtime
| Scribe v2 (async) | Scribe v2 Realtime | |
|---|---|---|
| Designed for | Batch transcription, media workflows | Live agents, meetings, streaming |
| Claimed latency | N/A (file-based) | Sub-150ms |
| Languages | 99 | 90+ |
| Word timestamps | Yes | Yes |
| Speaker diarization | Yes | No streaming diarization |
| Price | $0.22/hr (~$0.0037/min) + add-ons | $0.39/hr |
Add-ons on async: entity detection +$0.07/hr, keyterm prompting +$0.05/hr (rates via third-party price trackers, May–September 2026 — confirm on elevenlabs.io before budgeting).
Independent Benchmarks (Coval, September 2026)
Daily third-party testing on fixed public audio (clean, accented, noisy, far-field, phone codecs) places Scribe v2 Realtime mid-pack:
- Word error rate: 5.5% — rank #14 of 26. Respectable, behind the leaders.
- Time to final segment: 123ms — rank #9 of 24. Genuinely fast finalization.
- Time to first token: 2,162ms — rank #20 of 22. The weak spot: first words arrive slowly even though finals land quickly.
That profile — slow to start, fast to finish — suits agents that can tolerate a beat before responding but need crisp completed turns.
Head-to-Head: Scribe v2 vs the Incumbents
Using AssemblyAI's published comparison (vendor turf — reproduce on your audio):
| ElevenLabs Scribe v2 | AssemblyAI Universal-3.5 Pro | |
|---|---|---|
| English WER (real-world) | 5.87% | 4.35% |
| Code-switching WER | 10.07% | 7.60% |
| Async price | $0.22/hr | $0.21/hr |
| Realtime price | $0.39/hr | From $0.15/hr (budget) |
| Streaming diarization | No | Yes |
| BAA | Enterprise-gated | Standard |
Against Deepgram: Nova-3 remains cheaper for batch ($0.0043–0.0048/min) with a bigger developer ecosystem, while Scribe's edge is the single-vendor voice loop — transcribe with Scribe, reason, and speak with ElevenLabs TTS voices without juggling providers.
Strengths
- The TTS pairing. If your product already uses ElevenLabs voices or Dubbing, Scribe keeps audio in one platform with one bill.
- 99-language async coverage with word timestamps — strong for media, podcast, and captioning pipelines.
- 10,000 free credits, no card — easy to trial.
Weaknesses
- No streaming diarization on the Realtime model. Multi-speaker live audio (meetings, calls) needs another solution or post-processing.
- Slow time-to-first-token (2.1s in independent tests) undermines the sub-150ms marketing for snappy back-and-forth.
- Enterprise-gated BAA — healthcare and finance buyers face a sales process AssemblyAI and Deepgram don't require.
- Add-on metering for entities and keyterms complicates the headline $0.22/hr.
Verdict: Who Should Use Scribe v2?
- ElevenLabs TTS users — the integration dividend is real. One vendor, one API key, matched voices.
- Media and dubbing pipelines — timestamps + 99 languages + Dubbing in one house.
- Skip it if you need live speaker separation, the fastest first-token latency, or a frictionless BAA. Benchmark Universal-3.5 Pro and Nova-3 on your audio first.
FAQ
Is Scribe v2 better than Whisper? For structured output (timestamps, diarization, entities) yes. For raw multilingual robustness, Whisper Large v3 remains a strong open baseline — see our Whisper architecture breakdown.
Scribe v2 or Scribe v2 Realtime? Files → v2. Live microphone or phone audio → Realtime, and verify the 2-second first-token behavior against your UX bar.
Related comparisons & reviews
Keep reading
Voice Model Deep Dives
Best Speech-to-Text API 2026: Benchmarks & Prices
All 8 leading STT APIs compared: Nova-3, Universal-3.5 Pro, Scribe v2, Gemini 3.5, MAI, Soniox, GPT-4o. Prices, WER and picks.
Voice Model Deep Dives
Fastest Speech-to-Text 2026: Latency Benchmarks
Who is actually fastest? First-token, final-segment and endpointing numbers for Flux, Nova-3, 3.5 Pro RT, Scribe v2 and Parakeet.
Voice Model Deep Dives
Gemini 3.5 Transcribe Review: Smart STT Tested (2026)
Google’s Aug 2026 model polishes ramblings into formatted text. 85+ languages, smart transcription, and how it compares to Nova-3.
