XavierFok
← all posts

Building a word list for what AI mishears

2026-08-10 · by Xavier Fok

The errors are boring, and that's the point

Nobody talks about transcription corrections because they're not dramatic. A model doesn't hallucinate a paragraph out of nowhere very often. What actually happens, day after day, is smaller and duller: it swaps a brand name for a common word, drops an acronym, or turns a technical term into something that sounds similar but means nothing. If you run any pipeline that converts speech to text at volume, this is the failure mode you live with, not the sci-fi one.

I run a video pipeline where scripts get read aloud, recorded, and in some cases transcribed back for captions or for feeding into other automation steps. The transcription layer is either a cloud API or a local model, depending on the job and what I'm willing to send off-machine. Either way, the same handful of words break every single time, and they break the same way. That repetition is the whole reason a word list is worth building. Random errors aren't worth chasing. Systematic ones are.

Why the same words keep breaking

Speech-to-text models are trained on huge amounts of general audio. They're good at common words because common words are common in the training data. The words that trip them up are the ones that are rare in general speech but frequent in your specific content: brand names, product names, acronyms specific to your industry, and proper nouns that sound like something more ordinary.

In my case, the offenders are predictable. Brand names built from ordinary words get "corrected" back into the ordinary words. Acronyms tied to AI infrastructure get expanded into something that sounds plausible but is wrong. Anything with an unusual capitalization pattern or a compound word gets split apart because the model has no way to know it's supposed to be one token. None of this is a defect in the model exactly. It's just doing what it was trained to do: predict the most statistically likely sequence of words for that sound. If your content lives outside the statistically likely zone, you eat the cost of that.

What a word list actually does

A word list here isn't a dictionary and it isn't a style guide. It's a mapping of "what the model tends to output" to "what it should have said," built from real transcripts, not guesswork about what might go wrong. The distinction matters. If you sit down and try to predict every term the model might mangle, you'll write a list full of edge cases that never happen and miss the two or three that happen constantly.

The only reliable way to build it is to look at actual output. I keep a folder of transcripts from real runs and scan them for the recurring misses. When the same substitution shows up three or four times across different recordings, it goes on the list. When it shows up once, I ignore it. One-off errors are noise. Recurring ones are signal, and the list is where you store that signal so you don't have to rediscover it every time.

Where the corrections cluster

After doing this for a while, the errors fall into a few clear buckets.

The first is brand and product names, especially ones built from real words. If a name reads like an ordinary phrase, the model has almost no way to know it should be treated as a proper noun instead of two normal words strung together.

The second is acronyms. Anything in the AI and automation space is acronym-heavy, and short acronyms are exactly the kind of thing a general-purpose model will "helpfully" expand into a wrong but plausible full phrase, or mishear as a similar-sounding short word.

The third is homophones tied to technical vocabulary. Two terms that sound close but mean different things in a technical context will get swapped constantly, because the model has no domain signal telling it which one fits.

The fourth, and the one people underestimate, is numbers and version-style strings read aloud. Spoken numbers, especially ones with decimals or mixed letters and digits, get transcribed inconsistently depending on how they're phrased in the sentence.

Knowing these four buckets in advance doesn't eliminate the work of building the list. It just tells you where to look first when you're reviewing a new batch of transcripts.

Wiring the list into the pipeline, not just a notes file

A word list that lives in a document you occasionally glance at is close to useless. The value shows up when it's applied automatically, as a post-processing step right after the transcript comes back and before anything downstream touches it. In my setup, that's a straightforward find-and-replace pass keyed off the list, run on every transcript before it's used for captions or fed into the next stage of the pipeline. It's not clever. It doesn't need to be. The list is small, the substitutions are exact-match, and running it costs almost nothing compared to the transcription step itself.

The part that actually takes discipline is maintenance. Every so often I go back through recent transcripts and check for new recurring errors, especially after adding a new brand or a new recurring term to the content. If I skip that step for a few weeks, the list quietly goes stale and the same corrections start happening by hand again, which is exactly the manual work the list was supposed to remove.

What this fixes and what it doesn't

A word list catches known, repeating errors. It does not catch novel ones, and it does nothing for a transcript that's wrong in a way you haven't seen before. It's a targeted fix for a targeted problem, not a general accuracy improvement. If the underlying transcription is consistently rough for a whole recording, no amount of find-and-replace after the fact will save it. In that case the fix is upstream: cleaner audio, a different model, or in some cases just re-recording the line.

It's also worth being honest that this is manual curation work, not something a model does for you unsupervised. I'm not going to claim a tool automatically builds a perfect correction list, because that's not how it works in practice. Someone has to read the transcripts, notice the pattern, and decide it's worth adding. That's a small but real ongoing cost, and it's one of the quieter parts of running a content pipeline that people don't budget time for.

Running speech-to-text locally versus through a cloud API changes latency, cost, and how much you're comfortable sending off your own machine, but it doesn't change this particular problem. Local models trained on general audio have the same blind spot for your specific vocabulary as cloud ones do. A word list is a fix you apply after the transcription step regardless of where that step happens, because the failure mode is about vocabulary coverage, not about where the compute runs.

The actual payoff

The reason this is worth doing isn't that it makes the model smarter. It doesn't. It's that it turns a recurring manual annoyance into a one-time entry in a list that then runs automatically forever. Every brand name you don't have to manually fix in a caption file again is a few seconds saved, and a few seconds saved on every single piece of content adds up fast once you're publishing at any real volume. It's a small, unglamorous piece of infrastructure, and those are usually the ones that pay for themselves the fastest.

If you're building out your own automation and want to see how pieces like this fit into a bigger content and AI setup, [head back to the homepage](/) for more on what running this stuff in production actually looks like.

Get new guides and videos first — join the Telegram channel.