For narrated marketing and e-learning with a timeline editor, use Murf or Speechify Studio, both metered by generation time. For the widest choice of models in one API — 70+ languages on Eleven v3, roughly 75 ms on Flash v2.5 — use ElevenLabs, billed per character. If raw language count is the constraint, Azure publishes the longest list at 100+ languages and locales. For regulated, high-volume or SSML-heavy production, use Azure AI Speech or Google Cloud Text-to-Speech. Choose by billing unit first: seconds of audio or characters of text.
Some links below are affiliate links. We may earn a commission if you subscribe, at no extra cost to you. This does not affect which tools we recommend or how we rank them.
Key takeaways
- ElevenLabs bills 1 credit per character on its V2 multilingual models, with plans from $6 per month (Starter, 30,000 credits) to $990 per month (Business, 6,000,000 credits).
- Speechify Studio bills output duration at one credit per second, so the $19 per month Starter plan's 7,200 credits equal two hours of finished audio.
- Murf's API is pay-as-you-go at $0.03 per 1,000 characters, and its Falcon 2 streaming model at 1 cent per 1,000 characters with time-to-first-audio under 100 ms.
- No major text-to-speech vendor offers unrestricted self-serve voice cloning: ElevenLabs allows only your own verified voice for Professional clones, Azure requires a recorded consent statement from the voice talent, Google's Instant Custom Voice is allow-list only, and Murf gates cloning behind an enterprise sales conversation.
- EU AI Act Article 50 applies from 2 August 2026, requiring machine-readable marking of synthetic audio and clear deepfake disclosure, with fines to 3 percent of worldwide turnover.
- PlayHT is no longer a buying option: Meta announced the PlayAI acquisition on 13 July 2025 and framed the deal around AI Characters rather than the product, and no public PlayHT plan or signup page remains.
The six tools compared
Six products are worth shortlisting, split between editor-first tools, API-first tools and hyperscaler speech services. Every figure below comes from the vendor's own pricing or documentation pages.
| Tool | Best for | Pricing model | Voice cloning | Main limitation | API surface |
|---|---|---|---|---|---|
| Murf | Marketing and e-learning voiceover in a timeline editor | Studio subscription metered in voice generation time; API pay-as-you-go at $0.03 per 1,000 characters | Enterprise only, not self-serve | Cloning is gated behind a sales conversation | REST plus streaming; Falcon 2 and Gen2 models, regional endpoints |
| ElevenLabs | Widest language range and expressive narration | Credits, 1 credit per character on V2 multilingual; $6 to $990 per month | Instant cloning from Starter, Professional from Creator | Professional clones are restricted to your own voice | Full REST and streaming API; separate API pricing tiers |
| Speechify Studio | Predictable per-minute budgeting for video voiceover and dubbing | Credits by output duration: 1 per second voiceover, 3 dubbing, 30 avatar | Included from the $19 per month Starter plan | Free tier carries no commercial rights | Separate prepaid API with per-million-character rates |
| Azure AI Speech | Regulated, high-volume production needing fine prosody control | Per billable character with monthly commitment tiers; custom voice billed in compute hours plus hosting | Limited Access; application and recorded consent required | Custom voice needs approval before you can start | Speech SDK, REST, batch synthesis for audio over 10 minutes |
| Google Cloud Text-to-Speech | Multilingual apps already running on Google Cloud | Per character, with a monthly free allowance that differs by voice class | Chirp 3: Instant Custom Voice, allow-list only | Chirp 3: HD rejects SSML on streaming requests | REST and gRPC, streaming synthesis, long-form endpoint |
| OpenAI audio models | Teams already billing through the OpenAI API | tts-1 at $15.00 and tts-1-hd at $30.00 per 1M characters; gpt-4o-mini-tts at $12.00 per 1M audio output tokens | Not offered | No cloning and a small fixed voice set | Same API key and SDK as the text models |
Per-character versus per-second pricing
The most consequential difference between these tools is what they count. Character-billed services charge for input text regardless of how long the audio runs; duration-billed services charge for output seconds regardless of script density. A slow read with long pauses costs the same as a fast one on a character-billed API, but noticeably more in a duration-billed editor.
Azure documents its counting rules most precisely: every character in a processed request is billable, including numbers, spaces and punctuation, and SSML markup inside the text field counts except the <speak> and <voice> tags. Each Chinese character counts as two. Speechify Studio is the opposite: credits are defined purely by output, and re-exporting unchanged content costs nothing.
Murf: editor plus a low-cost streaming API
Murf fits if you want a timeline editor and a cheap streaming API from one vendor, and do not need self-serve cloning.
Murf
Best for: voiceover where a non-technical editor produces the audio and a developer later automates it against the same voice library.
Trade-off: cloning is not self-serve at any tier, and Studio and API do not share a billing unit.
Murf's API documentation describes 150+ voices across 35 languages and 20+ speaking styles, split between Gen2 for studio-quality synthesis and Falcon 2 for streaming. Falcon 2 is documented at time-to-first-audio under 100 ms and 1 cent per 1,000 characters, the lowest published rate of the six. The help centre lists a free trial of 100,000 characters, and pay-as-you-go at $0.03 per 1,000 characters with a $2 minimum, concurrency 15 and 10,000 requests per minute.
Delivery control is proprietary rather than SSML: rate and pitch as integers from -50 to 50, a variation setting from 0 to 5, a [pause] tag accepting 0.1 to 5 seconds, and an audio-duration parameter for matching a fixed video slot. Enough for narration, not a drop-in for an existing SSML pipeline.
All paid plans grant commercial rights over Studio voiceovers; the free plan does not. Cloning "is not a self-serve solution" and is Enterprise-only. Benchmark quality on the free API tier at Murf.
ElevenLabs: three models behind one API key
ElevenLabs is the default when you need many languages, expressive delivery, or both a fast and a high-quality model behind one API key.
Tiers published on the ElevenLabs pricing page in August 2026: Free $0 with 10,000 credits, Starter $6 with 30,000, Creator $22 with 121,000, Pro $99 with 600,000, Scale $299 with 1,800,000 and Business $990 with 6,000,000. On V2 multilingual models one character equals one credit; other models cost 0.5 to 1 credit per character via the API.
The model line-up matters more than the tier. Eleven v3 covers 70+ languages with a 5,000-character request limit for nuanced narration; Multilingual v2 covers 29 languages at 10,000 characters; Flash v2.5 covers 32 languages at roughly 75 ms and 40,000 characters. The documentation is candid that Flash offers limited emotional range and does not normalise numbers by default.
The commercial licence starts at Starter; free accounts are non-commercial and must attribute ElevenLabs in the title. The terms state you retain all rights in your output while granting a broad licence for service improvement.
Speechify: duration credits and a prepaid API
Speechify is the easiest of the six to budget for: Studio credits map onto seconds of output, not script length.
The Speechify Studio pricing page lists a free plan with 600 credits, no cloning and no commercial rights; Studio Starter at $19 per month with 7,200 credits, cloning and commercial rights; and Studio Creator at $49 with 28,800 credits. At one credit per second that is 10 minutes free, two hours on Starter and eight on Creator.
The API is a separate prepaid product: Free caps at 50,000 characters per month at three concurrent calls, Starter $10 includes 1M then $10 per million at six concurrent, Pro $99 includes 3M then $8 per million, and Scale $499 includes 10M then $6 per million at thirty.
Watch the language gap: the consumer reading app advertises 60+ languages, but the Studio pricing page enumerates roughly twenty.
Azure AI Speech: the deepest prosody control
Azure gives more documented control over delivery than any other service here, which keeps it the default for IVR, accessibility and broadcast. Microsoft's text-to-speech overview describes standard neural voices in 100+ languages and locales, with a batch synthesis API for files longer than 10 minutes.
The SSML surface is the differentiator. Azure supports mstts:express-as with a named speaking style and a styledegree from 0.01 to 2, where 1 is the predefined intensity and 2 doubles it. A role attribute lets a voice imitate another age or gender persona. Prosody rate accepts 0.5 to 2 times the original, and an audio-duration element can force a passage to a target length up to 300 seconds. The caveat: HD, personal and embedded voices do not support every tag, so verify per voice.
Pricing is published as a structure rather than a number: rates per 1M characters, a free tier of 0.5 million characters per month, and commitment tiers at 80M, 400M and 2,000M characters, with regional amounts only in the calculator. Custom voice adds training in compute hours plus hosting per hour — typically 20 to 40 compute hours for a single-style voice and around 90 for multi-style, capped at 96.
Google Cloud Text-to-Speech: Chirp 3 and the SSML gap
Google is the right pick if your stack already runs on Google Cloud, but its newest voices carry a markup limitation. The service offers five voice classes — Standard, WaveNet, Neural2, Studio and Chirp 3: HD — the last documented with 28 named voices across 50+ language and locale combinations.
The catch is markup. Standard, WaveNet and Neural2 support SSML, and Studio voices support it except <mark>, <emphasis> and <prosody pitch>. Chirp 3: HD accepts SSML for non-streaming requests but, in Google's words, "SSML tags are not currently supported for streaming requests". On a streaming agent your control narrows to the pace parameter (0.25x to 2x), bracketed [pause short] and [pause long] markers, and pronunciations as IPA or X-SAMPA.
Billing is per character with a monthly free allowance that differs by voice class, and generative tiers cost materially more than Standard. Confirm current per-class rates before modelling costs: the class moves the figure by roughly an order of magnitude.
OpenAI audio models: lowest friction if you are already there
OpenAI is not a voice platform, but if your app already authenticates against the OpenAI API, its speech models are the shortest path to audio. Per the OpenAI API pricing page checked in August 2026, tts-1 is $15.00 and tts-1-hd $30.00 per 1M characters, while gpt-4o-mini-tts bills $12.00 per 1M audio output tokens and gpt-realtime-2.1 $64.00.
What you give up is everything around the voice: no cloning, no voice library, no editor. Right trade for notification audio; wrong one for a brand voice you intend to keep. Our comparison of the major AI models in 2026 covers the same vendors from the language-model side.
What happened to PlayHT
PlayHT, later branded PlayAI, still appears in most 2026 roundups but is no longer purchasable. Meta announced the acquisition on 13 July 2025; an internal memo said the "entire PlayAI team" would join the following week, and Meta framed the deal around AI Characters, Meta AI and wearables rather than continuing the product.
What remains is documentation without a product behind it. The play.ht and play.ai apex domains no longer serve a marketing or signup page, while the API reference at docs.play.ht is still published — Quickstart, PlayDialog, the Node and Python SDKs, the v2.3 endpoint list — last dated September 2025. Treat that reference as an archive, not an offer: there is no public plan to buy and no vendor commitment behind the endpoints it documents. Anyone still holding integration code should plan a migration rather than a version bump.
Migrate by capability, not price. Instant-clone workflows map to ElevenLabs Instant Voice Cloning and conversational agents to Flash v2.5 or Murf Falcon 2, while multi-speaker PlayDialog scripts must be rebuilt as per-speaker requests. Same class of forced migration as the 2026 platform shifts and August deadlines.
Voice cloning: what each vendor demands first
Every vendor gates cloning behind a consent or identity step, and the mechanism decides which you can use. ElevenLabs offers Instant Voice Cloning from Starter and Professional Voice Cloning from Creator. Professional cloning recommends 30 minutes of audio as a bare minimum and 2 to 3 hours for best results, requires a verification recording made on similar equipment to the samples, and takes roughly 3 to 6 hours to fine-tune. Its policy is the strictest here: you can only clone your own voice, and "even with their consent, you cannot clone someone else's". Slots run from one on Creator to ten on Business.
Azure treats custom voice as Limited Access requiring an application. You must upload a recording of the voice talent reading a predefined statement in the training language — in English, "I [state your first and last name] am aware that recordings of my voice will be used by [state the name of the company] to create and use a synthetic version of my voice." Microsoft reserves the right to run speaker-recognition biometrics on it and match it against the training audio.
Google's Chirp 3: Instant Custom Voice is allow-list only and needs two recordings of up to 10 seconds made in the same environment: a clean reference sample and a consent recording using Google's script. Speechify includes cloning from the $19 Starter tier, Murf only on Enterprise, and OpenAI not at all.
Emotional and prosody control in practice
Three approaches to controlling delivery, and they are not interchangeable.
Full SSML
Azure and Google's non-generative classes accept standard SSML with vendor extensions. This is the only approach that survives a vendor change with modest edits, because the core tags are a W3C spec.
Proprietary numeric parameters
Murf exposes rate, pitch and variation as integers plus a bracketed pause tag; Chirp 3: HD exposes a pace multiplier and pause markers. Simpler to surface in a UI, but they do not port.
Model-inferred prosody
Eleven v3 infers emotion and pacing from context rather than markup, which is why the documentation positions it for character dialogue and Flash for speed. If you need identical output across regenerations, avoid it; if you need it to sound unrehearsed, avoid heavy SSML.
Language coverage and real-time latency
Languages and accents
Published coverage in August 2026: Azure at 100+ languages and locales; ElevenLabs at 70+ on Eleven v3, 32 on Flash v2.5 and 29 on Multilingual v2; Google's Chirp 3: HD at 50+ combinations across 28 voices; Murf at 35 languages with 150+ voices; Speechify Studio at roughly twenty.
Two caveats outrank the counts. Coverage is per model, not per vendor, so a voice on Eleven v3 may not exist on Flash. And locale count is not accent count: a service listing "German" may ship only standard High German with no Austrian or Swiss variant, a real constraint for DACH advertising.
Latency for real-time use
Only two publish figures suitable for conversational agents: ElevenLabs at roughly 75 ms for Flash v2.5, and Murf at under 100 ms time-to-first-audio for Falcon 2. Both exclude network round-trip, and for an EU-hosted agent the endpoint region usually dominates. Concurrency matters as much: Murf allows 15 concurrent calls on pay-as-you-go and Speechify's API tiers run 3 to 30. Our guide to AI agents that take actions covers the orchestration layer, and our comparison of the best AI video generators covers the platforms that bundle synthesis into video instead.
EU AI Act transparency and personality rights
Two regimes apply if you publish synthetic voice to an EU audience: the AI Act's transparency rules, in force today, and national personality rights, which bite the moment a cloned voice is recognisable.
Article 50 transparency, from 2 August 2026
The European Commission's Article 50 FAQ splits the duty in two. Providers of generative systems must ensure outputs, explicitly including synthetic audio, are marked in a machine-readable format and detectable as artificially generated. Deployers publishing deepfake content must disclose it in a clear and distinguishable manner at first exposure, and the Commission states an embedded machine-readable mark alone does not satisfy that duty.
One narrow grace period exists: for systems placed on the market before 2 August 2026, the Article 50(2) marking obligation applies only from 2 December 2026, and earlier content needs no retroactive labelling. Penalties reach 15 million euro or 3 percent of worldwide annual turnover. The Commission's Code of Practice on Transparency of AI-generated Content, published 10 June 2026 and signed by roughly 190 organisations by late July, is the voluntary route to demonstrating compliance.
Personality rights when the voice is a real person's
Complying with Article 50 does not give you the right to use someone's voice. In Germany that right is protected under the general right of personality derived from Articles 1(1) and 2(1) of the Basic Law, with civil claims through section 823(1) BGB. The Federal Court of Justice recognised it in the Marlene Dietrich decision (I ZR 49/97), and the Hamburg Higher Regional Court held in 1989 (3 W 45/89) that commercial imitation of a voice can be inadmissible.
A voice recording that uniquely identifies a person is also biometric data under Article 9 GDPR, so cloning needs explicit, purpose-specific, revocable consent. A photography model release will not cover it.
Commercial usage rights and what you own
Commercial rights are a plan feature, not a default, and three of the six withhold them on the free tier most people evaluate with. ElevenLabs restricts free accounts to non-commercial use and requires attribution to elevenlabs.io or 11.ai in the title, with the licence attaching from Starter. Speechify Studio's free plan excludes commercial rights, which begin on Starter. Murf grants them on all paid plans and not on free.
Two caveats apply regardless of vendor. The vendor licence does not override a distribution platform's own rules on synthetic voice. And owning the audio is not owning the voice: on a cloned voice your rights are bounded by the consent you obtained from the speaker, not by the tier you bought.
How to choose in under ten minutes
Four questions:
- Who produces the audio? A marketer or instructional designer needs an editor: Murf or Speechify Studio. A developer goes straight to an API.
- What shape is the workload? Long scripts with sparse delivery favour character billing. Fixed-length video slots favour duration billing.
- Do you need a cloned voice? If it is your own, the self-serve routes are ElevenLabs (Instant Voice Cloning from Starter, Professional from Creator with a documented 3-to-6-hour fine-tune) and Speechify Studio from its $19 Starter tier. If it is someone else's, only Azure and Google offer a documented consent workflow that survives legal review.
- Voice track or video? If you need a presenter on screen, a platform like Synthesia bundles voice with avatar. Our roundups of AI tools for graphic designers and AI image generators cover the visual half.
Before publishing to an EU audience, confirm the vendor marks output machine-readably, add your own disclosure where the audio imitates a real person, and keep the consent recording with the project files.
Frequently Asked Questions
How much does an AI voice generator cost per hour of finished audio?
It depends on the billing unit. Speechify Studio charges one credit per second of voiceover, so an hour costs 3,600 credits and the $19 per month Starter plan's 7,200 credits cover two hours. Character-billed APIs need a script estimate first: an hour of speech runs roughly 8,000 to 10,000 words.
Can I use AI voice-over commercially on YouTube?
Only on a paid plan. ElevenLabs limits free accounts to non-commercial use and requires attribution in the title. Speechify Studio's free tier excludes commercial rights, which begin on the $19 per month Starter plan. Murf grants them on all paid plans but not on free.
Do I need consent to clone someone else's voice?
Yes, and vendors enforce it technically. ElevenLabs states you can only create a Professional Voice Clone of your own voice, and will not accept someone else's even with their permission. Azure requires a recorded consent statement from the voice talent and may biometrically match it against the training audio.
What happened to PlayHT?
Meta announced the acquisition of PlayAI, the company behind PlayHT, on 13 July 2025, with an internal memo saying the entire PlayAI team would join the following week, and framed the deal around AI Characters, Meta AI and wearables rather than continuing the product. No public PlayHT plan or signup page remains. The API reference at docs.play.ht is still published, last dated September 2025, so treat it as an archive rather than an offer and plan a migration rather than a version bump.
Which AI voice generator has the lowest latency for real-time apps?
On published figures, ElevenLabs Flash v2.5 and Murf Falcon 2. ElevenLabs documents roughly 75 ms for Flash v2.5 across 32 languages, and Murf under 100 ms time-to-first-audio for Falcon 2. Both exclude network round-trip, so measure from your own region before committing.
Do I have to label AI-generated audio in the EU?
From 2 August 2026, yes. AI Act Article 50 requires providers to mark synthetic audio in a machine-readable, detectable format, and deployers publishing deepfake audio to disclose it clearly at first exposure. The Commission states machine-readable marking alone does not satisfy the deployer duty.
Does every AI voice tool support SSML?
No, and the gaps are model-specific rather than vendor-specific. Azure has the deepest SSML surface, including mstts:express-as with a styledegree from 0.01 to 2, but its HD and personal voices do not support every tag. Google's Chirp 3: HD accepts SSML except on streaming requests.