XavierFok
← all posts

Text to speech in 2026: what makes a synthetic voice convincing

2026-08-15 · by Xavier Fok

# Text to speech in 2026: what makes a synthetic voice convincing

Every voice on my channel is synthetic. No second of narration I have published came out of a human throat. That fact has forced me to think harder about fake voices than most people ever need to, because my entire production depends on a listener sitting through ten minutes of machine speech without being jarred out of the content. I have tested plenty of options, rejected most of them, and sat through hours of robotic output wondering whether I would have to settle. Eventually I found something that holds, and most viewers never think to ask. This post is what I learned along the way. It is deliberately free of product names and rankings, because the products change every few months while the lessons underneath them stay put.

Learning what to listen for

My first obstacle was vocabulary. Early on I could tell a voice sounded wrong without being able to name why, and a fault you cannot name is a fault you cannot fix. Over time the vague discomfort resolved into specific things I now listen for separately, and once you hear them separately you cannot unhear them.

The first is prosody, the rise and fall of pitch across a sentence that signals meaning and emphasis. Human speech does this constantly without effort. A voice that delivers every syllable at the same pitch and speed reads as a machine even when every individual sound is rendered perfectly. Prosody is the music of speech, and without it a voice sounds like something reading coordinates off a screen.

The second is breath and pause. Real speech is full of tiny gaps, small hesitations, the occasional audible breath, a slight slowdown before a word that matters. These signal that a mind sits behind the voice. When they are missing, or inserted in the wrong places, the listener's brain flags it immediately, even if the listener cannot say what bothered them.

The third is consistency across a long stretch, and this one matters most for what I make. The question I care about is whether the voice can carry ten minutes without losing the plot.

The cheap robotic tier

The options I have tried fall into rough categories, and the categories stay stable even as the products inside them churn. The bottom tier is the cheap robotic one: free tiers, browser built ins, the voices you meet on phone menus and government websites. They have improved over the years, and they still hit a ceiling you can hear. Individual words come out right while the prosody around them stays missing or misplaced. Emphasis lands on the wrong word. Clauses that should slow down get rushed. There is no breath anywhere, just a flat march through the text.

For a few seconds of utility audio, none of that matters. Phone prompts, brief interface messages, internal tooling. In those settings this tier is fine at its actual job. Ask it to carry a ten minute narration and listeners feel the effort of listening within a minute or two, then leave.

The good neutral tier

The middle tier covers the better voices from the big cloud providers and the dedicated synthesis platforms, and this is where I lost the most time, because this tier contains a trap. Many of these voices sound genuinely impressive for one paragraph. They have learned real prosody. They breathe. They vary their pace. Paste in a single well punctuated sentence and you will be convinced.

Then you run a full narration and the seams open. The model has learned patterns that hold for short bursts, and over a long piece the inconsistencies compound. One paragraph lands beautifully and the next goes strangely flat. A sentence that needed weight gets rushed. A breath lands mid thought. No single fault disqualifies the voice, but they accumulate, and by minute eight the listener has been quietly jarred enough times that something feels off even if they could never articulate what.

For short content this tier is honestly useful. Explainer clips, product demos, anything under two minutes sits comfortably inside its abilities. Long form narration is where it gets unreliable.

Voice clones

The third category is clones, where you upload recordings of a real voice, your own or one you have licensed, and the platform trains a model to reproduce it. Results range from uncanny to convincing, and the biggest variable is the source audio. Clean, consistent, well paced recordings produce a far better clone than a mixed bag of conditions.

Clones fit the cases where identity is the point: a recognizable voice tied to a person or a brand, a narrator the audience comes to know over time. Two honest caveats come with them. Even a good clone inherits a ceiling from its source, because it is constrained to sound like that specific speaker, who may never have been a dynamic performer in the first place. And most clone platforms bill by usage, so the cost scales with the volume you produce.

One sentence proves nothing

The most expensive mistake I made was testing wrong, and I want to spell it out because I suspect most people shopping for a voice are repeating it. I used to paste in a clean showcase sentence, listen, and judge. Plenty of voices passed that test. Then I would produce a full ten minute video and discover halfway through that the voice was falling apart, rushing paragraphs it had handled smoothly before, drifting away from the tone I needed.

A single sentence proves nothing. The only test that predicts production is the load of production itself. My method now is to paste in a full thousand word block of real narration, written the way I actually write, and listen to the whole output in one sitting. I count the stumbles: broken prosody, a punched word that should have been soft, a swallowed word that should have been clear, pacing that runs away. More than two or three problems in a thousand words and I move on. And if using a voice means babysitting its output line by line with punctuation tricks and phonetic overrides, that time is a real cost that wipes out whatever the service saved me.

The test material matters as much as the length. Showcase sentences are engineered to sound good. Real scripts have density, awkward constructions, and sentence variety, and the voice has to survive all of it at production settings. I also pay particular attention to the ends of long paragraphs, because that is where a voice usually starts to drift first.

What I use, and the honest reason

I use a metered, paid voice synthesis service. The cost per video is small, and at my volume it is a real but manageable line item. What the money buys is consistency. The voice holds across a full ten minute narration without supervision. Its prosody is imperfect, as every synthetic voice's is, but paragraph eight sounds like paragraph one, and across a whole video that steadiness matters more to me than any single beautiful sentence.

I picked it with the exact method above. Several candidates, the same thousand word block of real narration, a full listen to each, and the ones that lost pace got cut. This one held, and that is the entire story. I am not naming it because the field shifts constantly, because I have no financial relationship with any voice provider, and because a voice that holds for my scripts in my setup may behave differently on yours.

The honest thing about cost

A lot of noise surrounds free TTS tools, so let me say the quiet part plainly. The free tiers exist and some are usable for short content. The voices that hold up over long narration, with real prosody and real consistency, almost all sit behind a paywall or a meter. There is no scandal in that. Training a good voice model is expensive and the companies charge accordingly.

The framing that helps is to treat the voice as a production line item and weigh it against the alternative. Recording, editing, and re recording human narration for every video costs hours. A good synthesis service removes that entire block of work for a small fee per video. Evaluated as a budget line, the paid voice usually wins comfortably. Evaluated against an imagined free lunch, everything looks overpriced.

How to actually pick one

For someone starting out, my advice compresses to a short sequence. Work out how long your narrations will run, because a voice good enough for a two minute clip may collapse at ten, and your test has to match your real length. Test on your actual scripts rather than anything resembling a showcase. Listen to the whole output in one sitting and count the faults instead of savoring the good moments. Put an honest price on your own time, since a paid voice that saves an hour of audio wrangling per video justifies itself quickly. And decide whether identity matters to you: a recognizable branded voice points at a clone, while clean neutral narration is well served by the middle tier.

The bar I hold a voice to is steadiness rather than perfection. It needs credible prosody, breath in sensible places, and above all the ability to stay itself for the full length of a piece without being supervised. A voice that clears that bar over a real thousand word test is good enough to build on, and voices that clear it do exist right now. If this kind of plain accounting of how a faceless channel actually gets made is useful to you, there is more of it on [xavierfok.com](/).

Get new guides and videos first — join the Telegram channel.