Speech-to-text API priced by the hour of audio
AssemblyAI is an API and nothing else: audio goes in, and text returns with word timings, speaker labels, summaries and content-safety classifications attached. Charging is per hour of audio rather than per seat, at $0.15/hr on Universal-2 and $0.21/hr on Universal-3.5 Pro, with diarisation, PII redaction and medical vocabulary billed as surcharges that stack on top. A separate Voice Agent API bundles recognition, model reasoning and speech output into one $0.075 per minute rate.
Pay-as-you-go per hour of audio with no monthly plan: $0.15/hr Universal-2, $0.21/hr Universal-3.5 Pro, $0.15/hr streaming. Add-ons stack additively — diarisation +$0.02/hr async or +$0.12/hr streaming, PII redaction +$0.08/hr, medical mode +$0.15/hr. Signup carries $50 of credits with no card, and no volume tiers are published.
Use tool ↗The transcription API to reach for when accuracy and per-hour pricing matter more than a user interface, but the headline rate is a base rate and every label you actually want adds to it.
Watch out: The advertised $0.15/hr buys plain transcription only; diarisation, redaction and medical mode each add their own per-hour charge on top of it.
Side-by-side fact sheets against the competitors our research names — pricing, platforms, API and both verdicts on one page.
No reviews yet — be the first to review AssemblyAI.
Reviews are tied to your EffectHub account — one review per tool, so the rating for AssemblyAI reflects real users.
No questions yet — ask the first one about AssemblyAI.
Questions about AssemblyAI are tied to your EffectHub account, so answers can reach you and the section stays free of spam.