When local AI is the wrong choice
# When local AI is the wrong choice
Most of my week runs on models I host myself. The video pipeline, the transcript cleanup, the summarizers, all of it sits on a GPU in this room. And I still pay for cloud AI every month, on purpose. People who follow my channel find that strange, since so much of what I publish is about getting off the subscription treadmill and owning your own stack. But I have shipped production work on both sides, and the honest position is that local is the right tool for a big pile of jobs and the wrong tool for a smaller pile that matters.
This post is about the smaller pile.
The quality ceiling is real
The models that fit on a home GPU are smaller than the best cloud models, and on the hardest tasks that gap is wide. Multi step reasoning. Code changes that thread through a large codebase. Careful analysis of a long, dense document. Deep questions in a narrow specialty. Work in that category punishes small models, and if the answer has to be right, forcing a local model to cope is a bad trade. You spend the savings twice over in time, fixing output a better model would have gotten right on the first pass.
So my first filter is a sorting question: is this task ordinary or genuinely hard? Ordinary covers summarizing, rewriting, classification, first drafts, plain questions. Local handles that whole tier well, at a cost close to zero. Hard covers anything where a subtle mistake burns real money or real hours, and that tier goes to the cloud without hesitation. Treating the choice as a loyalty test is the actual mistake. It is a routing decision, made one task at a time, and the setups that work best quietly use both.
Workload shape beats ideology
The second place local loses has nothing to do with model quality. It is hardware math. My GPU runs one job at a time, give or take. Ask it to process ten thousand documents by tonight and it grinds through them in a line, possibly for days. A cloud service can rent me a rack's worth of capacity for twenty minutes and chew through the same pile before lunch. For the rare giant burst, renting is simply the correct architecture. Buying enough hardware to absorb a spike that happens once a month would be absurd.
Flip the shape and local wins. Steady, high volume, ongoing work is local's home turf. A few hundred small tasks a day, forever, turns per request pricing into a bill that never stops, while the same work on my own card costs electricity. So the pattern holds across my whole stack. Steady drip, run it local. Rare burst, rent the cloud. One hard problem, rent the cloud. A constant stream of easy problems, local by a mile.
The costs nobody puts in the spreadsheet
Run AI locally and you are the operations team. You install the tools, you update them when they break, you work out why a model that was fine last week is suddenly slow, you manage disk space and drivers. I enjoy that work, so for me it barely registers as a cost. For most people it is a recurring tax on attention. The cloud's genuine virtue is that someone else pays that tax. Send a request, get an answer, never think about VRAM.
There is a second hidden cost, and it took me a while to name it: the hour you lose wrestling a model that was never going to manage the task. When a local model is slightly out of its depth, the failure is ambiguous. You assume the prompt is the problem, so you tweak and rerun, tweak and rerun, while the clock eats an afternoon. That time never appears in anyone's local versus cloud comparison. A cloud model that gets the task right on the first attempt is often the cheaper option once your own hours are counted honestly. The wrestling itself is a signal that you picked the wrong tool.
Privacy cuts both ways
Local AI keeps your data on your machine, and for genuinely sensitive material that is a strong, legitimate reason to run it. I mean that without reservation. The oversold version is the slogan: local equals private, cloud equals exposed. A reputable provider with a clear policy of no training on your data and deletion after processing can be perfectly appropriate for plenty of work. Meanwhile a sloppy local setup, exposed to your network without a second thought, can be riskier than the service you were avoiding. Privacy earns local the win whenever the data truly must not leave the machine. Below that bar, it is a judgment call per task rather than a blanket rule.
Speed and uptime land where you least expect
For a short request, local feels snappy because there is no network round trip. For a heavy request, the oversized specialized hardware behind a cloud service often finishes first even after the round trip. Neither side owns speed. Light requests favor local, heavy requests favor the cloud, and if responsiveness on hard tasks shapes your workflow, that is a quiet point for paying.
Uptime is less subtle. A local model is exactly as reliable as my machine, my power, my most recent software update. If the box is off, the model is off. A good cloud service is kept alive around the clock by people whose entire job is keeping it alive. Anything that must answer at three in the morning, regardless of whether my setup is healthy, runs against the cloud, and I consider that money well spent.
The platform around the model
The big services sell more than a bare model. Tool use, images and audio alongside text, large managed context, safety filtering, plus whatever shipped last week, all bundled and maintained by someone else. Rebuilding that surface at home is real engineering work. Even when a local model matches the raw quality a task needs, the surrounding machinery can make the cloud cheaper in practice, because recreating what the platform already does would burn weeks.
There is a mirror image on the cost side, and local sidesteps it completely: metered pricing plus a bug equals a surprise bill. A loop that runs away can rack up a startling charge before anyone notices, so every paid API in my stack sits behind a spending cap. Local cannot do that to me. Whatever my scripts get up to overnight, the worst case is a warm room.
Where local wins without argument
Anything that runs constantly in an automated loop belongs on hardware you own. A script calling a model hundreds of times a day, every day, turns per request pricing into a meter that never stops spinning, and every network hop adds latency the loop pays again and again. Local crushes that case. No meter, no quota, no round trip. The more a workload looks like machinery rather than occasional human use, the more decisively local is correct.
Lock in is the other honest reason to stay local even where the cloud is technically better. Build deeply on one provider and you depend on their pricing, their policies, their continued existence. Terms change, and your recourse is limited. A model on your own disk cannot be repriced or taken away. For anything I want still running in five years, I accept a little less capability now in exchange for never being at someone else's mercy later. I would rather call that a deliberate trade than pretend local is simply better.
Test in the cloud, run at home
My flow for new ideas uses both sides in sequence. I prototype against a top cloud model first, to learn what the best available model can do with the task. Once the task is proven, and a smaller model demonstrably handles it, the steady production version moves local and the recurring cost dies. The cloud tells me what is possible. Local makes it cheap to keep.
That order matters more than it looks. Prototype locally and you can never tell whether the task failed because it is impossible or because the model was too small. Prototype in the cloud and you know the ceiling before you spend a single evening on plumbing.
The four questions
For any given task, I ask these in order. Is it hard enough that a small model would fail in expensive ways? Cloud. Is it a rare giant burst? Cloud. Is it steady, high volume, and within a small model's reach? Local. Must the data stay on my machine no matter what? Local, full stop, even at some quality cost. Nearly every task falls out of those four questions on its own, and the agonizing stops.
The people getting the most out of AI did not pick a side. They built a setup with both and learned to route. Local carries the steady, private, high volume work. The cloud takes the hard problems, the bursts, the prototypes. I prefer local, I run a lot of it, and staying clear eyed about where it loses is what makes the whole system work.
If you want the build notes behind decisions like this one, the rest of my write-ups live on [the home page](/).
Get new guides and videos first — join the Telegram channel.