XavierFok
← all posts

Local text-to-speech vs paid voice APIs: how I decide

2026-08-14 · by Xavier Fok

# Local text-to-speech vs paid voice APIs: how I decide

Voiceover is a line item in my life. Every video I publish needs one, so the local versus paid question is a bill I either pay or avoid each month, and a quality decision that shows up in the finished work. I have shipped real videos with both. Modern text to speech is genuinely good on either path, and the two paths are unequal in specific, predictable ways. Here is the grounded comparison.

Where local voices stand now

If your mental image of free local TTS is a flat robotic voice, it is out of date. Voice models running on a home GPU now produce clear, natural speech with sensible pacing and correct pronunciation on nearly everything ordinary. Narration, audiobooks, internal tools, accessibility, drafts: all comfortably within reach. A casual listener would rarely flag the better local models as synthetic on a straight read.

Being specific matters here, because vague praise is useless. The better local models get sentence rhythm right, land emphasis correctly on most words, treat commas and periods as real pauses, and lift the pitch slightly on questions. Where they slip is the unusual stuff. An odd proper noun comes out mangled. A strangely built sentence lands with the stress in the wrong place. Big emotional swings come out flatter than they should. For straight reading of clear text, the output is genuinely pleasant to listen to.

The one thing paid still does better

Emotion. When a line needs a laugh, a sigh, real excitement, subtle sarcasm, the top paid voices deliver it more convincingly than local models usually manage. Local excels at steady, clear narration and remains weaker at acting. So a project that lives or dies on a performing voice still points at the paid services. A project that needs pleasant, accurate narration of information can go local, because that gap has closed almost entirely.

Speed on a home GPU

Generating speech is light work compared with images or a big language model. On a decent card, a good local TTS model produces audio several times faster than real time, so a ten minute narration comes back in roughly two or three minutes. Generating a full voiceover happens while the coffee brews.

The speed matters for a second reason that took me a while to appreciate: free regeneration changes how you work. When every attempt costs money, you ration attempts, and the rationing quietly makes the work worse. When regeneration is free, you iterate on the script and the pacing until it is actually right. That freedom to redo cheaply has improved my output more than any single quality number.

The cost crossover

Paid voice services charge by the amount of audio generated, per character or per minute. At low volume that is cheap, and honestly the right choice. The bill scales with output though, and for someone producing hours of narration every week it becomes a real recurring cost. Local TTS carries no per minute charge. You paid for the GPU, you pay for electricity, and the audio after that is effectively free at any volume.

So there is a crossover point. Below some volume, the paid route is cheaper and simpler. Above it, local wins, and the gap widens with every hour you produce. My own crossover arrived quickly, because a channel publishing on a steady cadence generates narration by the hour, and hours are exactly what per minute pricing punishes. The question is never which is better in the abstract. How much audio do you make, and how much does the voice have to carry? High volume plus steady narration points hard at local. Low volume plus high emotional demand points at paid.

Control and consistency

Running the model on your own machine buys you things no paid tier sells. Unlimited generation with no quota hanging over the pipeline. Offline operation. Tight integration with your own scripts. And a voice that cannot be retired out from under you. Paid services change and remove voices over time, which can quietly break the identity of a long running series. The voice on my disk is the voice my channel keeps, for as long as I want it, and for a body of work built over years that ownership is worth a lot.

Integration deserves its own mention. Because the model runs on my machine, my scripts call it directly, so a written script turns into finished audio and flows straight to the next stage of the pipeline without me touching anything. At volume, that automation is the whole game. A paid service can be automated too, but every run costs money and a quota hangs over the loop.

Fixing pronunciation

Mispronunciation is the most common local TTS frustration, and it has a real fix. When a model mangles a name or a technical term, spell the word phonetically in the script, written the way it sounds. The model then says it correctly. For words you use often, keep a small substitution list and apply it automatically before generating. It is slightly manual at first, and each fix is permanent, so after a few weeks the problem words are all handled. That bit of script preparation separates a professional sounding voiceover from one with a jarring stumble in the middle.

The setup tax is real

A paid service is an API call: send text, receive audio, no maintenance. Local TTS asks more of you. Install the model and its dependencies, confirm it actually uses the GPU, audition voices, tune settings until the sound is right. The better local options can be fussy on first contact. Budget an afternoon to dial everything in, after which it mostly just runs. Whether that afternoon is worth it comes down to volume again. For someone needing an occasional clip, paid with zero setup is the smarter use of time. For anyone producing seriously, the afternoon pays for itself within weeks.

A hybrid worth stealing

Nothing forces one tool for everything. Route the bulk of the narration, the steady informational reading, through the local model at zero marginal cost. For the rare lines that genuinely need to perform, the emotional hook or the dramatic beat, generate just those few lines with a top paid voice and stitch them in. You pay for a handful of seconds instead of the whole runtime and keep most of the paid quality exactly where it counts. It takes a little more orchestration, and for serious content work it is the best of both worlds.

The line on cloning

Voice models can clone a specific person's voice from a short sample, and that capability deserves care. Cloning your own voice for your own work is one thing. Using these tools on someone else's voice without clear consent is a line I will not cross and neither should you. Running the model privately on your own machine does not soften any of that. It only means the responsibility sits with you instead of with a platform's review process.

Practical output details

Local TTS hands you a standard audio file that drops into any editor. You can generate one long file per script, or generate in pieces and stitch, which lets you regenerate a single flawed line without redoing the whole read. That granularity is a quiet advantage of running the model yourself. You control how the audio is chunked and assembled rather than accepting one block from a service. I keep the pieces small enough that a regenerated line costs seconds rather than a full read.

How I actually route it

My channel narration is steady, informational, high volume. It runs local, because the quality clears the bar for clear narration and the volume would make a paid bill real. The rare project where the voice must perform goes to a top paid voice without hesitation, since that is precisely the gap local has yet to close. The same pattern shows up everywhere in my stack: high volume steady work stays home, and money goes only to the specific cases where the difference is audible in the result.

Two questions decide it. How much audio do you make, and how much does the voice have to carry? Answer both honestly and the choice makes itself.

More honest comparisons like this one, from systems I actually run, live on [the home page](/).

Get new guides and videos first — join the Telegram channel.