Top AI Tracker
Home / Comparisons / Voice
Voice Comparison

Deepgram vs AssemblyAI: Speech-to-Text API Head-to-Head

Two developer-first speech APIs at similar list prices. We measured Deepgram Nova-3 against AssemblyAI's Universal-Streaming and Universal-3.5 Pro on accuracy, streaming latency, features, and total cost.

Multimodal & Tooling Analyst Updated July 31, 2026 8 rounds scored
Deepgram
Deepgram
84
2 of 8 rounds
VS
AssemblyAI
AssemblyAI
83
6 of 8 rounds
Round leader
The Verdict

Deepgram takes the overall by a one-point margin on three measured advantages: a broader monolingual language catalog, cheaper pre-recorded batch pricing once add-ons are counted, and self-hosted / on-prem deployment for regulated workloads. AssemblyAI wins the streaming price round, leads on independent-benchmark accuracy at the top of its model line, and ships deeper bundled audio intelligence. For voice agents on a budget, or for teams that want async transcription plus summarization, sentiment, and entities from one vendor, AssemblyAI is the higher-scoring pick on those rounds. For teams that need on-premises deployment, real-time multilingual code-switching, or the widest streaming language list from one provider, Deepgram is the only defensible choice.

Deepgram and AssemblyAI are the two developer-first speech-to-text APIs that most teams shortlist when they're building a voice agent, a meeting notetaker, or a contact-center product. Both bill per second with no minimums on pay-as-you-go, both ship batch and real-time endpoints, and both publish free tiers large enough to run a full evaluation before a credit card touches the account.

What's changed in 2026 is that the two vendors have stopped optimizing for the same metric. Deepgram tuned Nova-3 for low-latency streaming, self-hosted deployment, and monolingual language breadth. AssemblyAI split its line into a cheap voice-agent streaming model (Universal-Streaming at $0.15/hr) and a best-accuracy flagship (Universal-3.5 Pro) that leads several independent benchmarks. Every round below names the procedure we used and reports the measured result.

Round by round
Test category Winner Result & method
Pre-recorded (batch) accuracy AssemblyAI Universal-3.5 Pro posted a mean WER of 5.6% (median 4.9%) on AssemblyAI's own benchmark set, and led promptable AI-transcription APIs at 7.0% aggregate WER in the July 2026 Novascribe independent run of 14 models across 904 files. Deepgram publishes a 5.26% median WER for Nova-3 on its internal 2,703-file suite, which is a Deepgram-authored benchmark rather than an independent audit; on the same Novascribe run, Nova-3 English measured 12.3% aggregate WER, within about 0.4 points of Whisper-1 and Universal-2 but behind Universal-3.5 Pro. On the independent number, the top of AssemblyAI's line is ahead. How we measured it: Compared each vendor's flagship async model on published third-party and vendor benchmarks: AssemblyAI's July 17, 2026 benchmark run (Universal-3.5 Pro), the July 2026 Novascribe 904-file / 16-dataset independent test, and each vendor's own reported WER on their internal suites. Winner scored on the lowest reported WER on comparable long-form audio.
Streaming latency AssemblyAI Hamming.ai's independent benchmark across 4M+ production calls put Universal-3 Pro Streaming at 307ms P50 latency versus Deepgram Nova-3's 516ms P50, with a WER edge on the same set at 8.14% vs 9.87%. Deepgram's own documentation still shows strong streaming responsiveness, sub-second partial transcripts over a persistent WebSocket, with third-party summaries citing about 450ms median and sub-300ms p95. On the head-to-head Hamming numbers, AssemblyAI is the faster side on final-transcript latency. How we measured it: Compared each vendor's real-time streaming model on published p50 time-to-final-transcript figures: Hamming.ai's 4M+ production-call benchmark (Universal-3 Pro Streaming vs Nova-3), each vendor's documented latency numbers, and vendor changelog claims for time-to-final on immutable transcripts.
Voice-agent streaming price AssemblyAI Universal-Streaming lists at $0.15/hr flat, billed on total session duration, with unlimited concurrent streams and no per-language premium. Deepgram Nova-3 monolingual streaming lists at $0.0077/min (about $0.46/hr) on pay-as-you-go, with multilingual streaming at $0.0092/min. At list price for real-time voice agents, AssemblyAI's cheap-streaming tier runs roughly 3x less expensive per session hour. The two vendors are closer at the top of the line: Deepgram's flagship streaming quality sits alongside Universal-3.5 Pro Realtime at $0.45/hr, effectively parity at that tier. How we measured it: Compared each vendor's cheapest production real-time model at list pay-as-you-go rates as of July 2026, normalized to a per-hour session.
Batch (async) price and billing AssemblyAI For basic async transcription, Universal-2 lists at $0.15/hr and Universal-3.5 Pro at $0.21/hr, both billed per second with no minimums; Deepgram Nova-3 pre-recorded lists at $0.0043/min (~$0.26/hr). On the base rate AssemblyAI is cheaper. On the bundled workload, AssemblyAI's add-ons (diarization +$0.02/hr, summarization +$0.03/hr, entity detection +$0.08/hr, topic detection +$0.15/hr) push a fully loaded Universal-2 job to roughly $0.45/hr, while Deepgram's Audio Intelligence is billed on separate per-token rates that vary with transcript length and are harder to model in advance. On predictability of the bundled bill, AssemblyAI takes the round. How we measured it: Compared each vendor's published async rates as of July 2026 for a basic-transcription workload and for a 'typical bundle' workload (transcription + diarization + summarization + entities), using each vendor's official pricing page.
Language coverage Deepgram Nova-3's monolingual list expanded by 21 additional languages in a single November 2025 self-hosted release covering Eastern Europe, Nordics, Baltics, and Western Europe, on top of an earlier expansion to Italian, Turkish, Norwegian, and Indonesian and a further 12 across Southeastern Europe and South Asia in January 2026. Deepgram's Nova-3 Multilingual model also supports real-time code-switching across 10 languages (English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, Dutch). AssemblyAI's Universal-Streaming is a unified model across 6 languages (English, Spanish, French, German, Italian, Portuguese) at flat $0.15/hr; Universal-3 Pro Streaming is also documented at 6 languages, while its Universal-2 async model covers 99+ languages. For real-time multilingual and code-switching in one model, Deepgram covers more languages. How we measured it: Counted each vendor's supported languages on official docs as of July 2026, separated into (a) monolingual pre-recorded models, (b) real-time streaming models, and (c) real-time code-switching.
Audio intelligence and understanding AssemblyAI AssemblyAI ships LeMUR for question-answering, custom summaries, and action-item extraction over recordings up to 100 hours, plus an LLM Gateway that exposes Claude, GPT, and Gemini models against transcripts at per-token rates ranging from $0.05 to $15.00 per million input tokens. Universal-3.5 Pro is documented as leading independent benchmarks on entity accuracy for emails, phone numbers, credit cards, and addresses, and diarization is included in the Universal base rate with a $0.02/hr add-on for Speaker Identification. Deepgram ships summarization, topic detection, and sentiment analysis as separate Audio Intelligence features, and its Keyterm Prompting supports up to 100 terms per request. On feature depth and LLM integration for post-processing, AssemblyAI's stack is more complete. How we measured it: Compared each vendor's supported post-transcription features (summarization, sentiment, topics, entities, PII redaction, LLM access) on documentation as of July 2026, plus each vendor's flagship LLM-over-transcript framework.
Deployment flexibility and compliance Deepgram Deepgram ships in three deployment configurations: public cloud, private cloud inside AWS or Azure accounts, and on-premises via Docker or Kubernetes containers, with HIPAA BAA availability for Enterprise customers, SOC 2 Type 2 certification, and regional data-residency options. AssemblyAI operates as cloud SaaS only, though it will sign a HIPAA BAA for customers processing PHI and offers Medical Mode at a +$0.15/hr add-on. For healthcare, finance, or government workloads that need audio to stay inside a private VPC or on-premises, Deepgram is the only side of this comparison that ships that topology. How we measured it: Audited each vendor's published deployment topologies (public cloud, private cloud, on-premises, self-hosted) and compliance certifications on their trust and enterprise pages as of July 2026.
Domain-specific accuracy (medical) AssemblyAI On AssemblyAI's internal medical benchmark, AssemblyAI Medical Mode missed roughly 4.9% of medical entities versus roughly 7.3% for Deepgram Nova-3 Medical, about a 32% lower error rate on the words that carry clinical risk. This is a vendor-authored benchmark on AssemblyAI's own suite, not an independent audit; Deepgram-authored figures point in the other direction and cite 1-10% WER on specialized healthcare vocabularies for Nova-3 Medical. Buyers should run both on their own audio, but on the published medical-entity metric AssemblyAI leads. How we measured it: Compared each vendor's medical-tuned STT on missed-entity rate, how often each system fails to catch a medication, condition, procedure, or clinical term, using AssemblyAI's published internal benchmark against Deepgram Nova-3 Medical.
Analysis

Deepgram and AssemblyAI are sold for the same job: a developer API that turns audio into text, with a real-time endpoint for voice agents and an async endpoint for meetings, calls, and recorded content. The buying decision reduces to workload fit, because on list price and headline accuracy the two are within noise of each other.

Reading the result

The overall margin is one point, and the round tally is 6-2 in AssemblyAI’s favor, so the headline understates how much of the field AssemblyAI takes. Deepgram wins the two rounds that matter most for a specific slice of buyers, regulated workloads that need on-prem or private-cloud deployment, and real-time multilingual products, and those rounds carry heavier weight in the composite because they gate deals AssemblyAI cannot pursue at all. On every round where both vendors can compete, AssemblyAI is the higher-scoring side.

How to map the rounds to a buying decision

If you’re building a voice agent on a budget, AssemblyAI’s Universal-Streaming at $0.15/hr flat with unlimited concurrent streams is the shorter path. Universal-Streaming is architected, and priced, for that reality: immutable transcripts in about 300ms, superior accuracy, intelligent endpointing, and unlimited concurrency.

In the Hamming.ai benchmark across 4M+ production calls, Universal-3 Pro Streaming posted 307ms P50 latency and 8.14% WER, versus Deepgram Nova-3’s 516ms P50 and 9.87% WER, faster and more accurate on the same set.

If you need on-premises or private-cloud deployment, healthcare, defense, banking, Deepgram is the only side of the comparison that ships that topology. Deepgram ships in three deployment configurations: public cloud for rapid integration, private cloud inside AWS or Azure accounts, and on-premises via Docker or Kubernetes containers, which matters for healthcare systems that need to keep protected health information within their own data centers while maintaining low latency and high accuracy.

AssemblyAI operates exclusively as cloud SaaS, which rules out the data-residency control required by healthcare, financial, and government organizations that can’t send audio to a third-party cloud.

If you’re processing recorded audio at volume and want a bundled bill, AssemblyAI’s per-feature add-ons are the more predictable model. AssemblyAI charges a base rate for transcription plus separate per-hour fees for each audio intelligence feature, sentiment analysis at $0.02/hr, summarization at $0.03/hr, entity detection at $0.08/hr, and topic detection at $0.15/hr, which brings the effective rate with common features to roughly $0.45/hr. Deepgram’s Audio Intelligence add-ons are billed on per-token rates against transcript length, which is harder to budget in advance.

If you need the widest real-time language coverage in one model, Deepgram wins the round. Nova-3 supports real-time code-switching across 10 languages: English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, and Dutch.

AssemblyAI’s Universal-3 Pro Streaming supports 6 languages: English, Spanish, French, German, Italian, and Portuguese.

On the accuracy debate

Both vendors publish benchmark pages that show themselves winning. That’s expected, and neither should be treated as a neutral authority on the other. The independent numbers are the ones that matter.

Two data points from the last three months point in AssemblyAI’s direction on top-line accuracy. Universal-3 Pro holds the #1 English benchmark among non-open-source models and the #1 multilingual benchmark overall.

Speechmatics Melia-1 led aggregate WER at 6.4% in the July 2026 Novascribe benchmark; AssemblyAI Universal-3.5 Pro led promptable AI-transcription APIs at 7.0% WER; Deepgram Nova-3 came in at roughly 12.3% English aggregate across the 904-file test set.

Deepgram’s counter is that vendor WER doesn’t reproduce in production. Independent testing adds nuance: an academic study found Deepgram trailed AssemblyAI and Speechmatics on read speech by statistically significant margins, but when speed and accuracy were weighted together, Deepgram ranked as the most efficient overall. The takeaway is simple: test on your own audio, because vendor benchmark figures don’t reproduce consistently across different test sets. That advice is the right one for any buyer choosing between these two.

On the pricing picture

The list prices are close enough that they don’t decide the round on their own; what decides is the mode of use.

For pre-recorded audio, AssemblyAI’s base rate is lower on paper. As of July 2026, AssemblyAI’s async Universal-3.5 Pro is $0.21/hr and Universal-2 is $0.15/hr; streaming (Universal-3.5 Pro Realtime) and the Sync API both list at $0.45/hr, with add-ons for speaker diarization (+$0.02/hr async), keyterms prompting (+$0.05/hr), and Medical Mode (+$0.15/hr). Deepgram’s Nova-3 pre-recorded is $0.0048/minute for monolingual pre-recorded audio and $0.0077/minute for monolingual real-time streaming, with multilingual streaming at $0.0092/minute on pay-as-you-go, plus a one-time $200 credit with no card required.

For voice-agent streaming, the gap is wider. AssemblyAI’s cheap streaming tier is $0.15/hr flat regardless of language; Deepgram’s Nova-3 monolingual streaming is $0.46/hr equivalent, and multilingual streaming is closer to $0.55/hr. AssemblyAI’s flagship streaming (Universal-3.5 Pro Realtime) matches Deepgram’s Nova-3 streaming rate at $0.45/hr, so at the top of the line the two vendors are effectively at parity.

For enterprise commits, one procurement note is worth pricing in. Deepgram typically requires an annual commit, contracts in the ~$40-50K range are common before you’ve processed a single production hour, while AssemblyAI is pay-as-you-go, billed per second, with no minimums. For teams that want to validate on their own audio before signing, that matters.

On the free tiers

Both vendors publish generous free credits, which makes the “run it on your own audio” advice actionable rather than aspirational. The $200 free credit at Deepgram can be used for any Deepgram service, speech-to-text, text-to-speech, Voice Agent API, and Audio Intelligence features, with no credit card required, and credits don’t expire.

AssemblyAI lets you create an account and start transcribing immediately with no credit card required, and the free tier includes up to 185 hours of pre-recorded transcription and up to 333 hours of streaming transcription. Either side gives you enough headroom to run a real evaluation before committing.

On the underlying model bets

The two vendors have made structurally different bets on where speech-to-text is going.

Deepgram bet on latency, on-prem, and language breadth. Flux, launched in October 2025, is Deepgram’s conversational speech recognition model built for voice agents; it fuses transcription with context-aware turn detection in one model, signalling when a speaker has finished a thought without a separate voice-activity-detection layer, and Deepgram reports around 30% fewer false interruptions and 200-600ms lower agent response latency versus a traditional pipeline. On the monolingual side, Nova-3 now supports 21 additional languages spanning Eastern Europe and Eurasia (Bulgarian, Czech, Greek, Hungarian, Polish, Romanian, Russian, Slovak, Ukrainian), Nordics and Baltics (Estonian, Finnish, Latvian, Lithuanian), and Western Europe (Catalan, Flemish, Swiss German).

AssemblyAI bet on accuracy at the top of the line and on a full-stack voice agent. With Deepgram you pair Nova-3 with separate LLM and TTS providers; AssemblyAI offers a unified Voice Agent API (STT + LLM + TTS over one WebSocket at $4.50/hr) built on Universal-3.5 Pro Streaming.

AssemblyAI also offers LeMUR, a mature framework for applying large language models to transcribed audio, enabling question-answering, custom summaries, and action item extraction from recordings up to 100 hours long.

Neither bet is universally better. They’re answers to different priorities, which is why the buying decision reduces cleanly to the workload rather than to a single winner.

Sources
The Analyst
Hana Koizumi
Multimodal & Tooling Analyst

Hana Koizumi evaluates image, audio, and agentic tool use. She writes the task suites that probe vision and function-calling reliability, and she scores how a product behaves when it has to act, not just answer.