XavierFok
← all posts

Transcribing audio locally with Whisper: the whole workflow

2026-08-13 · by Xavier Fok

# Transcribing audio locally with Whisper: the whole workflow

I transcribe hours of audio every week and I have never paid a transcription service to do it. Everything runs through Whisper, an open speech to text model, on the GPU already sitting in my machine. The audio never leaves my computer, there is no per minute charge, and the quality holds up against the paid options for most real work. This is the workflow I actually use, from install to a clean transcript, including the places where it fails.

Why bother when paid services exist

Two reasons, and the first is bigger than money. Transcription touches sensitive material. Interviews, meetings, voice notes, personal recordings. Uploading those to a service means trusting a company's data policy with some of the most private audio you own. Local Whisper keeps every second of it on your own disk.

The second reason is volume. Paid services charge per minute of audio, so the bill scales in a straight line with how much you transcribe. Local Whisper has no per minute price at all. A hundred hours costs the electricity to keep a GPU busy for a while. If you transcribe occasionally, per minute pricing is fine and simple. If transcription is a regular part of your week, the local math wins by more every single month.

The one decision that matters: model size

Whisper comes in several sizes, from tiny to very large, and choosing between them is most of the skill. Bigger models are more accurate, especially on messy audio, and they cost you speed and memory. Smaller models are quick and light, and they stumble more on hard speech.

Here is how I choose in practice. Clean audio, one clear speaker, decent microphone: a medium model is plenty, and it is fast. Hard audio, several speakers, accents, background noise, phone call quality: step up to a larger model, because the accuracy gain is worth the wait. Throwaway transcripts where I only need the gist: a small model, nearly instant. The beginner mistake is reaching for the largest model every time. On clean audio you often cannot hear the difference in the result, and you paid for it in speed anyway.

Install and first run

The cleanest path is one of the optimized Whisper builds made to run efficiently on a GPU. It installs as a Python package with a single command, and the main thing to verify is that your graphics drivers are current so the acceleration actually engages. If you never want to see a terminal, graphical apps exist that wrap the whole thing, and they are a fine starting point. I run the command line version because it slots into an automated pipeline, where audio arrives and transcripts appear without me clicking anything.

Running it is genuinely simple. Point it at an MP3 or WAV, name the model size, and wait. The first use of any size downloads that model once, after which it lives on your disk and runs fully offline. Ask for timed output if you want subtitles, and each chunk of text comes stamped with when it was spoken.

Speed lives and dies on the GPU

On a decent graphics card, Whisper runs far faster than real time. An hour of audio finishes in a handful of minutes with a mid sized model. On CPU alone, the same job can take longer than the recording itself. So if transcription feels slow, check whether it quietly fell back to the CPU before blaming anything else. That is the usual culprit. When the model is properly on the card, waiting stops being part of the workflow.

Where accuracy breaks

No transcription is perfect, local or paid, and Whisper struggles where humans struggle. Heavy crosstalk with people talking over each other. Thick accents over a bad connection. Specialist jargon and unusual proper nouns it has rarely seen. On clean speech from a good microphone, a strong model lands well above ninety percent accuracy, often high enough that proofreading takes minutes. On a noisy multi person meeting, expect real correction work. There is no model, at any price, that makes bad audio perfectly accurate.

One quirk deserves its own warning: invented phrases. During long silences, stretches of music, or very noisy passages with no clear speech, Whisper sometimes writes out a plausible sentence nobody said. It is filling the gap with its best guess. The defenses are simple. Trim silent and musical sections before transcribing, and for anything critical, check the transcript against the audio wherever the timestamps show long gaps, because that is where inventions hide.

Knowing the quirk exists is most of the protection.

Cleaning the raw output

The raw transcript is a wall of text, or a stream of timed lines, and it usually needs work before a human wants to read it. I chain a second local model for this step. The raw Whisper output goes to a local language model with one instruction: fix the punctuation, break the text into paragraphs, remove the obvious filler, and change nothing about the meaning. Seconds later the rough machine transcript reads like a document. Because that step also runs locally, the pipeline stays private end to end. Audio in, readable transcript out, nothing uploaded, nothing billed.

Formats: decide what you are making first

Whisper produces several shapes of result, and picking the right one up front skips a conversion step later. Plain text with no timing is what you want for a readable document. Timed segments suit subtitles. It writes standard subtitle file formats directly, ready to drop onto a video. And word level timing stamps every single word, which is what you need for captions that highlight along with the speech. I use that last mode for every video I publish: the voiceover goes through Whisper with word level timing and comes out as caption data ready to lay onto the footage, accurate enough that correction is rare.

What it does not do

Whisper transcribes words. It does not tell you who spoke them. Separating speakers is a different task, usually called diarization, and it needs an extra tool layered on top. Projects exist that combine the two, and they work, with added complexity and imperfect results. If your audio is a multi person conversation and you need names attached to lines, set that expectation before you start.

Languages are the opposite story, better than most people expect. Whisper handles many languages well, transcribing speech into text in the same language. It can also translate as it goes, turning speech in another language into English text. The translation is rough, a first pass rather than a polished rendering, and for understanding a recording in a language you do not speak, it is remarkably useful.

Wiring it into a pipeline

The setup stops being a chore and starts being infrastructure once you automate it. I keep a folder where audio files land. A script watches that folder, runs Whisper on anything new, passes the result through the cleanup model, and drops the finished transcript next to the original. Recordings go in, transcripts come out, and I touch none of it. At the volume I run, doing this against a per minute service would be a real recurring bill. Locally it is a few minutes of GPU time per file.

A few practical habits save headaches. Convert odd audio formats to something plain before transcribing, since clean input avoids strange failures. Trim long silence and music, which improves speed and accuracy at the same time. For very long recordings, keep the timed output so you can jump to any moment and verify it against the audio. And leave room on disk, because the larger models run to several gigabytes each.

The whole system in one paragraph

Install an optimized Whisper build and confirm it is using the GPU. Match the model size to the difficulty of the audio, medium for clean, larger for messy. Ask for timed output when captions are the goal. Pass the raw result through a local model to tidy it, and if you do this often, put a folder watcher in front so transcripts simply appear. After the hardware you already own, the marginal cost of each transcript is a slice of electricity, and no second of your audio ever leaves the machine.

Owning transcription instead of renting it has been one of the clearest wins in my local AI setup. More build notes like this one live on [the home page](/).

Get new guides and videos first — join the Telegram channel.