Best AI Speech-to-Text APIs for Developers, Ranked by Accuracy, Latency, and Cost
We benchmarked the leading speech-to-text APIs on real-world English audio, streaming latency, multilingual accuracy, feature depth, and per-minute cost, with 2026 pricing verified against each vendor's pricing page.
Deepgram takes the top slot on the combination of real-time latency, per-second billing, and a batch rate that undercuts every hosted competitor except Whisper. AssemblyAI is the pick when you need transcript intelligence bundled in and the tightest streaming word error rate. ElevenLabs Scribe v2 Realtime wins multilingual real-time on both latency and FLEURS accuracy. Speechmatics is the answer for regulated, on-prem, or accent-heavy deployments. OpenAI Whisper and GPT-4o Transcribe are the cheapest managed batch option and the open-source baseline; skip them if you need streaming.
Five speech-to-text APIs, one fixed evaluation frame. We picked the platforms developers actually shortlist when they're wiring transcription into a voice agent, a meeting product, a call-center pipeline, or a media workflow, and held the evaluation axes constant so the differences on the table trace back to the APIs rather than the audio.
Every provider was scored on the same five axes: English word accuracy on real-world audio, streaming latency for voice-agent workloads, multilingual coverage and accuracy, feature and workflow depth (diarization, keyterm prompting, entity/PII detection, formatting, deployment modes), and effective cost per audio hour at the vendor's published 2026 rates. Independent benchmark numbers are cited where available; vendor-published figures are labeled as such.
Scores reflect API behavior as of the 2026-08-01 test date at each vendor's default settings on a paid tier with no custom vocabulary or fine-tuning. Where independent third-party benchmarks are available (Coval, Hamming.ai, Hugging Face Open ASR Leaderboard, Artificial Analysis AA-WER v2), we anchored the accuracy and latency scores to them rather than to vendor-published numbers. Pricing was verified against each vendor's pricing page during the week of testing. Cost is reported alongside quality but isn't folded into the quality score.
Scored against public independent benchmarks that use real conversational and code-switched audio rather than clean read speech. Anchors included the Hugging Face Open ASR Leaderboard, Artificial Analysis AA-WER v2 (AgentTalk subset), Coval's 2,400-run STT benchmark across five providers, and Hamming.ai's benchmark across 4M+ production calls. 100 corresponds to 0% WER on the reference set; providers that only publish vendor benchmarks on their own test sets were penalized for transparency. Weighted 30%.
Time-to-first-partial and end-of-turn detection for real-time streaming, measured against published third-party numbers where available. Anchors: Hamming.ai P50 latency across 4M+ production calls, Coval time-to-first-token across 2,400 runs, and vendor-documented end-of-speech detection figures. For voice-agent workloads, sub-300ms is treated as the passing bar because a voice-to-voice round-trip budget under 800ms leaves 150-300ms for STT. Weighted 25%.
Number of supported languages, real-time multilingual streaming availability, and independent multilingual accuracy where measured (FLEURS across 30 languages was the primary anchor). Providers with English-first models were scored lower even when their English number was strong. Weighted 15%.
Presence and quality of production features developers actually ship with: streaming speaker diarization, keyterm/vocabulary prompting, PII/PHI/PCI entity detection, word-level timestamps, formatting and punctuation, HIPAA/SOC 2 posture, on-prem or self-hosted deployment, and SDK ergonomics (REST + WebSocket, official SDKs). Each capability was scored present-and-good, present-but-weak, or absent. Weighted 20%.
Effective dollar cost per audio hour at each vendor's published pay-as-you-go rate for the primary streaming model, verified against the pricing page during the week of testing. Add-on charges (diarization, redaction, keyterm prompting, audio intelligence tokens) are noted where they materially change the effective rate. Normalized so a lower cost-per-hour scores higher. Reported alongside quality but never folded into the quality score. Weighted 10%.
Deepgram ships two flagship models: Nova-3 for general-purpose transcription and Flux for voice-agent pipelines with integrated end-of-turn detection. Flux Multilingual, released April 29, 2026, posts median end-of-turn detection under 300ms and saves 200-600ms on agent response time versus a separate STT-plus-VAD pipeline. Pay-as-you-go pricing is $0.0043 per minute batch and $0.0077 per minute streaming for Nova-3 monolingual, with per-second billing rather than 15-second rounding. The weaknesses sit on the accuracy side: on independent AA-WER v2 real-world sets, Nova-3 records higher WER than the AssemblyAI, ElevenLabs, and Speechmatics flagships, and speaker diarization, redaction, and keyterm prompting are each priced separately and stack onto the base rate.
Source: Deepgram ↗Strengths
- Lowest published end-of-turn detection latency in the field via Flux
- Per-second billing rather than per-minute or 15-second rounding
- $200 in free credit with no card required, covering roughly 45,000 minutes of Nova-3 batch
- Voice Agent API bundles STT, TTS, and LLM into one endpoint
Weaknesses
- Higher WER than AssemblyAI, ElevenLabs, and Speechmatics flagships on independent real-world benchmarks
- Diarization, redaction, and keyterm prompting are add-ons that stack onto the base per-minute rate
- Growth plan requires a $4,000/year prepaid commitment
How it scored, by metric
AssemblyAI ships Universal-3 Pro (batch, released February 3, 2026), Universal-3 Pro Streaming (March 2026), and Slam-1, a speech-language model that adds natural-language keyterm prompting up to 1,500 words. Universal-3 Pro posts 5.6% mean WER on Artificial Analysis AA-WER v2 and ranked third on the AgentTalk subset, and Hamming.ai's benchmark across 4M+ production calls measured Universal-3 Pro Streaming at 307ms P50 latency and 8.14% WER against Deepgram Nova-3 at 516ms P50 and 9.8% WER. Beyond raw transcription, the platform bundles sentiment, topic detection, entity recognition, PII detection, and an LLM Gateway priced separately. The weaknesses are language coverage for real-time (live multilingual streaming works in six languages) and add-on stacking that inflates the effective per-hour rate.
Source: AssemblyAI ↗Strengths
- Lowest streaming WER of the hosted flagships on independent third-party benchmarks
- Natural-language keyterm prompting up to 1,500 words via Slam-1
- Bundled sentiment, topic, entity, and PII detection in one API
- Dated public changelog for model improvements
Weaknesses
- Live multilingual streaming currently supports only six languages
- Audio intelligence features are priced separately from base transcription
- Streaming latency trails Deepgram Flux and ElevenLabs Scribe v2 Realtime
How it scored, by metric
Scribe v2 Realtime, released November 11, 2025, targets approximately 150ms first-partial latency across 90+ languages and posts 93.5% accuracy on the FLEURS benchmark across 30 languages, ahead of Gemini Flash 2.5, GPT-4o Mini Transcribe, and Deepgram Nova-3 on the same set. Scribe v2 batch supports 99 languages, speaker diarization up to 48 speakers, entity detection across 65 categories, keyterm prompting up to 1,000 terms, and dynamic audio tagging for non-speech events. The weaknesses are pricing model and real-time diarization: STT bills per credit rather than a flat per-minute rate (with a 45% price cut on May 7, 2026), and the realtime model doesn't offer speaker diarization, pushing multi-speaker meeting workloads to the batch tier.
Source: ElevenLabs ↗Strengths
- Highest independently measured multilingual real-time accuracy in the field
- Sub-150ms latency across 90+ languages via WebSocket
- Batch model supports speaker diarization up to 48 speakers, 99 languages
- SOC 2, ISO 27001, HIPAA, and GDPR posture out of the box
Weaknesses
- Real-time model doesn't currently offer speaker diarization
- Credit-based pricing is harder to model than flat per-minute rates
- Keyterm prompting adds a 30% premium over base transcription cost on third-party hosts
How it scored, by metric
Speechmatics Ursa 2 posts an 18% WER reduction across 55 languages versus Ursa 1 and code-switching accuracy 35% better than the nearest competitor. The platform supports 56+ languages with bilingual packs, claims ISO/IEC 27001:2022, SOC 2 Type II, GDPR, and HIPAA alignment through its trust center, and ships three documented deployment modes: cloud, on-premises, and on-device. Pro-tier pricing is usage-based from around $0.24 per hour for batch and $0.0067 per minute for real-time, with a Flow voice-agent endpoint at $0.0537 per minute and an opt-in data-sharing program that takes 33% off the rate. The weaknesses are the intelligence stack and the developer-first polish: transcript features are less bundled than AssemblyAI's, and language availability at enterprise tier is contract-dependent.
Source: Speechmatics ↗Strengths
- Strongest accent and code-switching accuracy on independent testing
- Cloud, on-premises, and on-device deployment fully documented
- 56+ languages with bilingual packs and speaker diarization included as a core feature
- $100 in credit and 480 free minutes per month to evaluate without a card
Weaknesses
- Transcript-intelligence stack is less bundled than AssemblyAI
- Enterprise tier language availability can be contract-dependent
- Developer-first polish trails Deepgram and AssemblyAI on ergonomics
How it scored, by metric
OpenAI ships three transcription products: open-source Whisper (self-hostable, 99+ languages, batch-only), gpt-4o-transcribe and gpt-4o-mini-transcribe (managed batch endpoints at $0.006/min and $0.003/min respectively), and GPT-Realtime-Whisper, released May 7, 2026 at $0.017/min for dedicated low-latency streaming. Whisper Large V3 Turbo, released October 2024, delivers 5.4x speed improvements by reducing decoder layers from 32 to 4. The strengths are ecosystem, price, and self-hostability; the weaknesses are that base Whisper is batch-only with no speaker diarization or streaming, and third-party benchmarks have caught GPT-4o Transcribe collapsing to as high as 43.8% WER on long-form financial earnings-call audio in one July 2026 test.
Source: OpenAI ↗Strengths
- gpt-4o-mini-transcribe at $0.003/min is the cheapest managed API in the field
- Whisper is fully open-source and self-hostable with 99+ language support
- GPT-Realtime-Whisper adds a dedicated $0.017/min streaming option
- Large ecosystem of third-party hosts (Groq, Fireworks, Replicate) for cheap Whisper inference
Weaknesses
- Base Whisper is batch-only with no streaming, no diarization, no speaker labels
- GPT-4o Transcribe collapsed to 43.8% WER on long-form earnings audio in one independent July 2026 benchmark
- Fewer production features (redaction, keyterm prompting, entity detection) than the specialists
- Streaming endpoint is newer and less tested in production than Deepgram or ElevenLabs
How it scored, by metric
The ranking above reflects each API’s default settings on a paid tier during the week of testing, with independent third-party benchmarks used to anchor the accuracy and latency scores where available. The single largest separator at the top of the table isn’t raw English WER (the top four are within roughly two points on clean English audio) but the interaction between streaming latency, feature bundling, and how much of the effective per-hour cost lives in add-ons the vendor doesn’t put on the headline pricing page.
Why the top three are close
Deepgram, AssemblyAI, and ElevenLabs are within four points of each other because each one wins a different axis and loses another. Deepgram wins on latency and per-second billing but trails on independent real-world WER. In benchmarks from Hamming.ai across 4M+ production calls, AssemblyAI’s Universal-3 Pro Streaming posted 307ms P50 latency and 8.14% WER, versus Deepgram Nova-3’s 516ms P50 . AssemblyAI wins on transcript intelligence and streaming WER, but its live multilingual coverage is narrower. ElevenLabs wins on multilingual real-time, Scribe v2 Realtime delivers 150 ms latency with 93.5% accuracy across 30 languages, outperforming Gemini Flash 2.5, GPT-4o Mini Transcribe, and Deepgram Nova 3 on the FLEURS benchmark , but its real-time model doesn’t currently offer speaker diarization.
What the WER numbers really mean
Every provider in this ranking publishes at least one vendor-run WER figure that flatters the model. Published WER benchmarks often use clean, well-recorded audio with clear speech. Real-world audio includes background noise, overlapping speakers, accents, domain jargon, and poor recording quality. A provider showing 5% WER on benchmarks might deliver 15-20% WER on challenging production audio . We anchored the accuracy scores to independent benchmarks (Coval, Hamming.ai, Artificial Analysis AA-WER v2, Hugging Face Open ASR Leaderboard) rather than to vendor decks, and penalized providers whose only public numbers came from their own test sets.
The batch-versus-streaming split
The pricing gap between batch and streaming isn’t marginal. Deepgram’s pricing has a critical decision point that drastically affects your costs: real-time streaming vs batch processing. Nova-3 costs $0.0077/min on Pay-As-You-Go or $0.0065/min on Growth plans , roughly 1.8x the batch rate. AssemblyAI’s streaming premium runs about 25%. The right call is to route by workload: We don’t use a single STT API, we route based on content type. Clean podcast audio goes to self-hosted Whisper large-v3 (lowest cost, excellent accuracy on studio-quality audio). Multi-speaker meeting recordings go to AssemblyAI (diarization quality justifies the cost premium). Real-time transcription for our live features runs on Deepgram’s streaming API . That routing pattern is the production norm at scale.
Add-ons are where the effective per-hour rate lives
Base per-minute pricing isn’t total cost. On Deepgram, speaker diarization is available as an add-on supporting up to 20 channels; at production volume, diarization, keyterm prompting, and redaction each carry separate per-minute charges that stack onto the base rate. A pipeline processing millions of minutes annually will see that cost compound quickly . AssemblyAI’s intelligence features (summaries, sentiment, topic detection, speaker labels) are similarly priced separately, and ElevenLabs charges a 30% premium for keyterm prompting on third-party hosts. Model the workload with add-ons enabled before you treat the headline rate as the budget.
Where OpenAI sits in 2026
Whisper is no longer the accuracy leader on any independent benchmark we could find, but it’s still the cheapest self-hostable option and the default entry point for teams already inside the OpenAI ecosystem. The open-source Whisper Large V3 Turbo (October 2024) delivers 5.4x speed improvements through architectural optimization, reducing decoder layers from 32 to 4. In March 2025, OpenAI released gpt-4o-transcribe and gpt-4o-mini-transcribe models with lower error rates than Whisper. OpenAI now recommends gpt-4o-mini-transcribe over gpt-4o-transcribe for best results . The streaming gap closed in May 2026 when OpenAI shipped GPT-Realtime-Whisper on May 7, 2026 at $0.017/min streaming, alongside refreshed Realtime-2 voice models. First time OpenAI separated streaming-optimized STT from the batch Whisper line , but it’s newer and less tested in production than the incumbents.
- https://deepgram.com/
- https://www.assemblyai.com/
- https://elevenlabs.io/speech-to-text
- https://www.speechmatics.com/
- https://platform.openai.com/docs/guides/speech-to-text
- https://deepgram.com/pricing
- https://www.speechmatics.com/pricing
- https://elevenlabs.io/realtime-speech-to-text
- https://www.coval.ai/blog/best-speech-to-text-providers-in-2026-independent-benchmarks-and-how-to-choose/
- https://futureagi.com/blog/speech-to-text-apis-in-2026-benchmarks-pricing-developer-s-decision-guide/
Q.Which speech-to-text API has the lowest word error rate in 2026?
On independent, real-world benchmarks the picture splits by mode. For clean English batch, NVIDIA Canary Qwen (5.63% WER) and Deepgram Nova-3 (5.26% WER in batch mode) currently lead independent ASR leaderboards , though those Deepgram figures are measured on Deepgram's own test set. For streaming, Hamming.ai's benchmark across 4M+ production calls measured AssemblyAI's Universal-3 Pro Streaming at 307ms P50 latency and 8.14% WER, versus Deepgram Nova-3's 516ms P50 . For multilingual real-time, ElevenLabs Scribe v2 Realtime leads FLEURS. Test on your own audio before you commit.
Q.What is the cheapest speech-to-text API for developers?
On base per-minute rates, OpenAI currently lists gpt-4o-transcribe at $0.006/minute and gpt-4o-mini-transcribe at $0.003/minute , which is the floor among managed APIs. Deepgram Nova-3 is $0.0043/min for pre-recorded audio and $0.0077/min for streaming with per-second billing and $200 in free credit. AssemblyAI batch starts at roughly $0.15/hr. The catch is that streaming, diarization, redaction, and keyterm prompting are typically priced separately, so the effective per-hour cost at production can be 2-3x the sticker rate.
Q.Which STT API is best for building a voice agent?
Deepgram Flux is purpose-built for that workload. Deepgram Flux Multilingual (April 29, 2026) is the first conversational STT with integrated end-of-turn detection. No external VAD needed. Median EOT under 300ms, saves 200-600ms on agent response time vs. STT + VAD pipelines . ElevenLabs Scribe v2 Realtime is the closest alternative when the workload's multilingual, at approximately 150ms first-partial latency across 90+ languages.
Q.When does it make sense to self-host Whisper instead of using a managed STT API?
Self-hosting Whisper is the right call when data residency, air-gapped deployment, or high-volume batch cost is the binding constraint. Open-source Whisper is batch-only with no speaker diarization or streaming, so real-time voice-agent workloads should go to Deepgram, ElevenLabs, or AssemblyAI regardless of cost. For cheap Whisper batch inference at scale, Groq Whisper-v3 at $0.04/hr, Lemonfox at $0.17/hr, fal.ai Wizper, Replicate, 10-30× cheaper than OpenAI's first-party Whisper endpoint for cost-sensitive workloads .
Q.Which STT API is best for regulated industries and on-prem deployments?
Speechmatics is the strongest pick. It offers three documented deployment modes (cloud, on-premises, and on-device), supports 56+ languages with bilingual packs, claims ISO/IEC 27001:2022, SOC 2 Type II, GDPR, and HIPAA alignment through a trust center, and offers a free plan with 2,400 minutes per month . Deepgram and AssemblyAI both offer self-hosted deployments as well, but Speechmatics is the only entry in this ranking with on-device as a documented option.
Devon Mizrahi measures what a model costs to run and how fast it answers. He maintains the price-per-token tables and the latency rigs, and he is the reason the Tracker reports tokens-per-second next to every quality score.
Other leaderboards
- Tooling Best AI Figma-to-Code Tools for Product Teams, Ranked by Fidelity, Component Reuse, and Workflow
- AI sales tools Best AI Cold Email Platforms for Sales Teams, Ranked by Personalization, Deliverability, and Cost
- Multimodal Best AI Video Dubbing Platforms for Global Video Teams, Ranked by Lip-Sync, Voice, and Cost