XavierFok
← all posts

Why I dictate more than I type

2026-08-10 · by Xavier Fok

The habit started as an accident

I didn't set out to build a "dictation workflow." I started talking to my phone instead of typing because my hands were full, and I never really stopped. A few months in, I noticed most of my Telegram replies, my WhatsApp messages, my rough article outlines, and even my commit messages were coming out of my mouth first and my keyboard second, if at all. That's not a productivity hack I'm selling you. It's just what happened once I paid attention to where the friction actually was.

Typing was never the real bottleneck

I can type fast. That was never the problem. The problem is that typing and thinking compete for the same channel. When I'm typing a message, part of my brain is managing spelling, punctuation, and where my fingers are on the keys. When I'm talking, that channel is free. I can hold a whole thought in my head and say it in one breath instead of building it word by word on a screen.

This matters more for certain kinds of output than others. A one-line Slack reply doesn't benefit much from dictation. A three-paragraph explanation of why a proxy rotation script is misbehaving, or a note to myself about what I want an AI agent to do next, benefits a lot. Talking lets me get the whole shape of the idea out before I start editing it, instead of editing sentence by sentence as I type.

What actually gets dictated in a normal day

Most of it is unglamorous. Replies to messages on Telegram and WhatsApp, first drafts of emails, notes to self that later become task lists, and rough article drafts like early versions of the piece you're reading now. I also dictate a lot of "status update" style writing, the kind where I'm summarizing what changed in a system or what I found while debugging something. That's the exact kind of writing where speaking out loud mirrors how I'd explain it to a person standing next to me, which usually makes it clearer than a typed version would have been on the first pass.

What I don't dictate: code, config files, anything with exact syntax, and anything where formatting matters more than content. Voice input is bad at brackets, indentation, and exact variable names. I'm not going to pretend otherwise. For that stuff I type, because typing is the right tool for precision and dictation is the right tool for getting a train of thought down quickly.

Where dictation breaks down

It's worth being specific about the failure modes, because a lot of the writing about voice input glosses over them.

Technical vocabulary is inconsistent. Product names, acronyms, and anything outside common English get mangled more often than plain conversational language does. If I'm dictating a note that mentions a specific tool or protocol name, I check it before I send it, because there's a real chance it came out wrong.

Long, unbroken dictation drifts. If I talk for two minutes straight without pausing to think about structure, I get a wall of text that reads like a transcript, not like writing. The fix isn't a tool, it's discipline: pause, restate the point, break it into chunks the way I would if I were talking someone through a plan.

It also doesn't remove the editing pass. Dictated text is a first draft, not a finished one. I still go back and tighten sentences, cut repetition, and fix the places where spoken language doesn't read well on a page. Anyone telling you voice input skips editing is selling you something.

Local versus cloud, and why it's not a simple choice

Since I run a mix of local and cloud AI tools for actual work, people assume I have a strong opinion on whether dictation should run locally or in the cloud. The honest answer is that it depends on what I'm dictating.

For anything with real content in it, client conversations, financial notes, anything I wouldn't want sitting on someone else's server, I want the audio processed on-device. That means accepting whatever the local option's accuracy ceiling is, and on older or weaker hardware that ceiling can be noticeably lower than what a cloud service delivers. For everyday, low-stakes dictation, a cloud-based option is usually faster to set up and more accurate out of the box, because those models are bigger and get updated more often than most people can run locally.

There's no universal right answer here. The tradeoff is the same one that shows up everywhere in AI tooling: local keeps your data closer to you and gives you more control over the failure modes, cloud gives you convenience and generally sharper accuracy. I use both, chosen per situation, not because I think one is categorically better.

How this connects to agents and automation

This is the part that actually matters for how I work, not just how I write. A lot of the automation I run day to day starts as a spoken note. I'll dictate something like "check on the render queue" or "draft a reply to this thread and flag it for review" while I'm doing something else, and that note becomes the input to a task I hand off, either to myself later or to an agent set up to act on it. Voice is a fast way to get an instruction out of my head and into a system, faster than opening a task manager and typing a structured entry.

That's the actual use case, not some vision of talking to a computer like it's a person. The dictation step is just an input method. What happens after, whether a human reads it, an agent picks it up, or it becomes an article draft like this one, is a separate decision with its own checks. Voice input doesn't make an agent smarter or more trustworthy, it just changes how the instruction got typed. I still review what agents do with dictated instructions the same way I'd review anything else, because a garbled transcription turning into a bad instruction is a real failure mode, not a hypothetical one.

What I'd actually tell someone trying this

Start with the low-stakes stuff. Messages, notes, rough drafts. Don't dictate anything you're not willing to proofread, because you will need to proofread it. Pick a mode, local or cloud, based on what you're comfortable having leave your device, not based on which one claims to be more advanced. And don't expect it to replace typing. It replaces the parts of writing that are really thinking-out-loud, and it leaves the parts that need precision exactly where they were.

If you want to see how this fits into the rest of how I run AI tools day to day, including where local models make sense and where they don't, take a look around the rest of the site [here](/).

Get new guides and videos first — join the Telegram channel.