How to pick an AI voice for a project
Three decisions that narrow the field before the listening starts, and the one that cannot be undone.
How we checked what a voice has to clear
A voice is chosen in the reverse order to most software. The feature list is not the hard part and it is not what the decision turns on: what turns it is whether the audio that comes out suits the room it plays in, and no plan table makes that judgement for you. What a plan table does settle is everything around it, and there is more around a voice than there first looks.
Three questions narrow the field before the listening starts, ordered by how much it costs to get them wrong. The first is what the voice is doing, because narration, dubbing and a live conversation pull the same product in different directions. The second is which language, and which voice inside it, since the highest-quality model and the widest language list are not always the same one. The third is licensing, and it is the only one of the three that cannot be fixed once the audio exists.
Nothing here was recorded or listened to. The model names, character limits, language counts, clone requirements and credit allowances below are read from what the platforms publish, checked on the date at the end of the page.
Three decisions, in order
Decision 1: Does the voice narrate, dub or answer?
The same platform sells these as separate models, and they are not interchangeable. A long-form model is built to hold up across an audiobook and trades speed for it. A conversational model is built to answer inside a pause short enough that the silence does not give the machine away, and it trades some warmth for that. An expressive model sits above both on direction, with tags written into the script to place emphasis and emotion, and a character ceiling per request that a full chapter will not fit inside.
Pick the model from the job rather than the other way round. A product page read aloud needs the cheap fast one, a narrated course needs the long-form one, and a voice answering inside an app needs the low-latency pair. Choosing by price alone lands a long audiobook on a model that renders it in fragments.
Decision 2: Which language, and which voice inside it?
A headline language count is the weakest number on the page, because it is usually the best case across a whole catalogue rather than a property of the model you are about to use. On the platform we reviewed in this group, the long-form model covers 29 languages, the fast pair covers 32 and the expressive model is listed at more than 70, while the pricing page advertises the widest of those as a single figure.
Read the count off the model you intend to use, then ask the second half of the question, which is depth. Two vendors can both list a language and differ in how many distinct voices sit inside it, and that difference decides whether a cast can be voiced from the catalogue or has to be cloned. No pricing page settles it, and the free tier is where to test it.
Decision 3: Do you own the voice, and may you use it?
Licensing is the decision that cannot be revisited later, and it splits into two parts that are easy to run together. The first is the licence on the output: on the platform we reviewed, commercial use begins at the first paid tier and the free plan carries none, so a voiceover that earns anything at all needs a subscription before it needs a voice.
The second is the clone. Instant cloning needs only a minute or two of audio and returns a usable model in seconds, enough to hold one narrator across a batch. A professional clone is a different purchase: it wants half an hour of clean audio to work at all and closer to three hours to be at its best, it needs verifiable permission from the person whose voice it is, and it is limited to a handful of slots starting on a mid-upper tier. Voicing a cast is therefore a count of clone slots rather than a count of voices in a library.
Where the decisions stop
The three decisions above settle which platform and which tier. Three other things decide the outcome, and none of them appears on a plan table.
The script, and the meter behind it
Voice platforms in this group bill by the character rather than by the seat, so the script is the invoice. A long script is billed as a long script, and on the platform we reviewed the same pool of credits also pays for music, sound effects, voice changing and dubbing, which means the voice budget and the everything-else budget are one number. Where a project is audio-led that is simply the cost of the medium; where the voice is one ingredient, it is worth knowing the allowance is shared before it runs out mid-project.
The room the audio plays in
A voice that reads well through headphones can still be wrong for the place it is heard. Narration for a training module, a voice answering on a phone line and a track sitting under music in a video are three different mixes, and none of them is settled by choosing a model. What matters at the cheap end is whether the platform gives you the format and bitrate your editing chain expects, because below a certain tier one of the platforms here holds output at a lower bitrate and reserves the higher one for the plans above it.
Whether the voice is the whole job
This is the line that decides which kind of product to buy. A dedicated voice platform stops at the audio track, so the picture stays somebody else’s job, and a video that needs a presenter as well is a job for a video tool rather than a voice one. Those tools carry a voice of their own, and cloning often arrives earlier on them than on a voice platform: the free plan of the video tool we reviewed in the neighbouring group includes a single voice clone, where cloning on the platform reviewed here begins at its first paid tier and the free plan carries no commercial licence at all. Where the deliverable is the audio itself that separation is a feature rather than a gap; where the voice is one track inside a video, it may be a subscription that does not need to exist.
Where the three decisions leave you
Two voice platforms are reviewed on this site and they answer different questions. One is built around a model catalogue and bills by the character; the other starts as a transcription tool and treats a synthetic voice as one output among several. Which is right depends on whether the audio is the deliverable or a step on the way to one, and that is the question Decision 1 is there to settle.
Where the voice has to end up inside a picture, the tools that carry it there are collected in our guide to the AI video generators we have reviewed, and the brief those tools need covers the parts of the job this page leaves alone.
Sources
ElevenLabs (elevenlabs.io), checked 16 September 2026: the pricing page for tier rates, credit allowances, the tier the commercial licence starts on and the professional clone slots; the text-to-speech page for the model list with each model's language coverage and character limit; and the voice-cloning page for the audio a clone needs and the consent requirement.
HeyGen (heygen.com), checked 16 September 2026: the plan comparison and help centre for the voice-clone allowance included with each tier, used here for the point where a video tool's own voice work replaces a separate subscription.
Reviews and comparisons
Last updated Sep 16, 2026 · Sources for every figure are stated on this page · Editorial Policy
