XavierFok
← all posts

Two years of local AI: what was hype and what was real

2026-08-16 · by Xavier Fok

# Two years of local AI: what was hype and what was real

I started running AI models on my own hardware two years ago, and I believed a good chunk of the sales pitch when I did. Some of those promises turned out to be understated. Others quietly fell apart once the novelty wore off. Since I push hundreds of small tasks a day through a card that is years past its prime, I can now score the pitch line by line against what actually happened. This is that ledger.

The privacy promise held up completely

When a model runs on my machine, my data stays on my machine. There is no fine print under that sentence. Nothing I feed the model leaves the building, and no policy update can quietly change the deal.

The ownership side mattered even more than the privacy. A tool I run locally cannot be discontinued, repriced, or gated behind a new plan while I sleep. My setup from last year still works this year because nobody else has a say in it. Of everything local AI promised, independence is what it delivered most fully, and it is the reason I will never go back to renting every capability I depend on.

The cost promise held up, with one sharp condition

Local AI genuinely is far cheaper than paid services for steady, high volume work. Hundreds of small jobs a day against a metered service would be a real monthly bill. On my own card, the marginal cost rounds to electricity. That held.

The condition the hype skips: the savings only exist at volume. The hardware and the power are fixed costs, and so is the time you spend on setup. Light use never earns them back. If I ran a handful of prompts a week, paying per use would beat owning the gear by a wide margin. I have watched people buy a graphics card specifically for AI, use it lightly, and spend far more than a pay per use service would ever have charged them. Calling local AI free is the most misleading claim in this whole space. Unmetered after a real upfront cost is the accurate version, and the difference matters.

The parity promise was oversold

The big claim was that local models had caught up with the cloud ones. On ordinary work, they mostly have. Summarizing, rewriting, classifying, drafting: a good local model handles all of it well enough that the gap barely registers. On genuinely hard work, the best cloud models remain clearly better. Complex reasoning, tricky code, and deep analysis all still favor the frontier, and that gap has been stubborn. It narrows for a stretch, then the cloud models improve and it opens back up.

So anyone telling you a small local model fully replaces the best cloud model is selling something. It replaces the easy majority of the work. The hard minority still lives elsewhere.

What believing the oversell cost me

Early on I burned real hours trying to force small models through tasks that were simply beyond them. Wrong answer, tweak the prompt, wrong answer again, tweak again, always convinced the magic phrasing was one edit away. It never was. Not once. The model was too small for the job, and no prompt fixes that.

The skill I eventually built was recognizing the ceiling quickly. When a task keeps failing after a couple of honest attempts, I stop fighting and route it to a bigger model. That habit has saved me more frustration than any tool I have installed since.

Progress was real, and so was the treadmill

What my old card runs today is dramatically better than what the same card ran two years ago. The models became more efficient and the tooling matured, so fixed hardware kept doing more. That was a genuine surprise on the pleasant side.

The honest half of the same observation: the top of the field moved just as fast. My local capability grew a lot in absolute terms and never gained an inch on the frontier. Both statements are true at once, and you have to hold both to see the situation clearly.

Nobody prices in the maintenance

Running AI locally means you are the operations team. Updates change behavior overnight. A model that ran fine last week turns sluggish. Drivers want attention at the worst possible moment, and when something breaks late at night before a deadline there is no support line to call, only you and a search box. I enjoy this kind of tinkering, so for me the cost is small and occasionally fun.

Plenty of people hate it, and I have watched them bounce off local AI entirely for that reason alone. They wanted a tool and found themselves holding a part time operations job. Sticking with the cloud is a completely valid response to that, and the pitch rarely admits it.

The reframe that made everything work

For the first stretch I treated local versus cloud as a contest with a winner. Dropping that framing was the most useful shift of the whole two years. Steady, high volume, easy work goes local, where it is private and effectively unmetered. Hard one off problems go to the cloud, where the quality lives. The people getting the most out of AI right now run both and know which tasks belong where. Once I wired that routing into my systems, the question of which side wins stopped mattering to me.

The cynics overcorrect

Honesty cuts both ways, so the doubters get a line in the ledger too. Some people describe local AI as a toy that trails the field by years. For the large majority of everyday tasks, a good model on a home GPU has been genuinely good enough for a while now. The truth sits in an unglamorous middle: very good at most things, second best at the hardest things. That middle position is exactly where the practical value lives.

The sleeper win: small models on narrow jobs

The hype pointed me at the biggest model I could run. The real workhorses of my setup turned out to be tiny ones. Classifying a message into a category, or deciding which of a few buckets a piece of text belongs in, or pulling one field out of a messy block: a small model nails jobs like these in a fraction of a second, costs nothing to run, and never gets bored of them. Chasing raw capability nearly made me miss where the volume value was, and that lesson has paid off more than any single big model I have downloaded.

Still a moving target

One promise aged badly: that setup would keep getting easier until this all became an appliance. The basic tools did improve. The ecosystem also fragmented and sped up, so there is always a new tool, a new format, a new best practice, a new gotcha, and staying current is its own ongoing effort. Things that worked one way last year work slightly differently now. Anyone calling local AI turnkey is glossing over how much churn still runs under the surface.

A quieter version of the same disappointment: the distance between a demo and a dependable system swallowed a lot of my early time. Getting a model to do something impressive once, on a clean example, is easy. Getting it to do the same thing reliably across messy real inputs takes far more work than the initial magic suggests. I learned to discount demos, including my own early excitement, and to judge every plan by how it survives the ugly middle of real data.

What overdelivered: the open ecosystem

Capable models are openly available, and so is the software that runs them. So is a flood of hard won knowledge from people who hit the same walls I did. I never had to ask anyone's permission to build any of this, and when I got stuck, the answer was usually already written up somewhere. The hype focused on the models themselves. The real gift was the open culture around them, and it is the reason a one person operation like mine is possible at all.

The habit that cut through the noise

Test on your own work. When a new model appears, I run my actual tasks through it and read the output. Charts and benchmark scores stopped mattering to me the day I started doing this, because results on my work kept disagreeing with results on paper. Occasionally a humble model quietly beats the celebrated one on exactly what I need. That habit, more than any purchase, is the first thing I would hand a beginner.

What I would tell someone starting today

Expect local AI to be excellent at the big pile of ordinary work, and expect real savings only if your volume justifies the fixed costs. Budget honestly for the hardware and the power, and count your own time as a real cost too. Keep a cloud option for the hard cases instead of forcing a small model to fail at them. Above all, treat this as building a system with two kinds of tools rather than picking a side in a fight.

The most honest summary I can offer is the shape of my own usage curve. The share of my AI work running locally climbed from almost nothing to the large majority of my daily tasks, then leveled off short of everything. It took over the bulk because the bulk is ordinary, and ordinary is where local shines. It stalled below the top because the hardest problems still belong to the cloud.

And the thing I got most wrong at the start: I treated this as a destination, a rig I would build once and be done with. It is a practice. The models keep changing, and so do my needs, so the right configuration from two years ago is wrong today. Go in expecting a finished product and you will be frustrated. Go in expecting a system you tend over time and it keeps paying you back. Two years in, I run more locally than ever and I still pay for cloud access, and that split is what actually works.

If this kind of no hype accounting is useful to you, there is more of it at [xavierfok.com](/).

Get new guides and videos first — join the Telegram channel.