Installing Ollama and running your first local model
# Installing Ollama and running your first local model
Running an AI model on your own machine sounds like an advanced project. It is a ten minute job. At the end of it you have a real model answering questions on your own hardware, with no account created, no card on file, and no internet connection required once the download finishes. This is the exact on ramp I put every beginner on before anything fancier, so let me lay it out start to finish, including the places where the easy path quietly hides things.
Set expectations before you start
I want to be straight about what you are getting, so nothing feels like a bait and switch halfway through. The models you run at home are smaller than the giant cloud ones, which makes them a little less sharp on hard tasks. For a large share of everyday work, though, they are properly good: answering questions, summarizing, rewriting, drafting, light coding help. And they are completely yours. Nothing you type leaves your machine. If what you want is a private, offline assistant with no recurring cost for ordinary tasks, you are in the right place.
Why Ollama
The tool for this is Ollama. Think of it as the thing that absorbs the annoying parts: it downloads model files and works out how to fit them onto whatever GPU you have, and it gives you a plain way to talk to the result. More powerful and more flexible tools exist, and you can graduate to them once this stops being enough. Ollama is where everyone should start because it removes most of the friction that makes people give up on local AI inside the first ten minutes. It runs on Windows, Mac, and Linux, and behaves almost identically on all three.
The install is boring, which is the point
Go to the Ollama website, download the installer for your operating system, run it. On Windows and Mac it is a normal installer you click through. On Linux it is a single command pasted into the terminal. No account creation, no email confirmation, no configuration files, no environment variables to set. When it finishes, Ollama sits quietly in the background as a service, waiting to be asked for something. That is the whole setup.
Pulling your first model
This step contains your one real decision. Open a terminal and type ollama run followed by a model name. I point beginners at a small, well rounded model in the three to eight billion parameter range, because it fits on almost any modern machine and responds quickly.
The first time you run the command, Ollama downloads the model, which is a few gigabytes, so give it a minute or two. After that the model lives on your disk permanently. Every future run starts instantly, and the download never repeats. One slow first pull buys you standing offline access.
The moment it clicks
When the download finishes, a prompt appears. You type a question, and the model answers, right there in the terminal, generated entirely on your own hardware. No request went to anyone's server. You can switch off your internet connection and it keeps working.
The first time people see that, it lands in a way no explanation manages. Until this point, AI was something you rented access to. This is the point where it becomes something you run, and that shift is what everything I build on this channel grows out of.
If it feels slow
Here is the biggest thing the easy path hides. Response speed depends almost entirely on whether the model fits inside your graphics card's memory. When it fits, words stream out faster than you can read them. When it is too big for the card, Ollama still runs it, spilling the overflow onto the CPU and system memory, and everything slows to a crawl.
So a sluggish first model nearly always means one thing: too big for your card. The fix is a smaller model, or a more aggressively compressed version of the same one. The sizing method I give beginners is simple. Start small and prove it runs fast, then step up in size until it stops being fast, then step back one. On a card with eight gigabytes of memory, good seven and eight billion parameter models run comfortably. More memory lets you reach larger ones. You do not need to understand compression levels on day one. Slow means go smaller; fast with quality to spare means you can try bigger.
The local server turns a toy into a tool
Chatting in the terminal is nice. The real power is that Ollama also runs a local server in the background, which means your own scripts and programs can send prompts to the model and read back answers, all on your machine.
It speaks a simple, well known interface, the same shape the big cloud AI services use. Code written against your local model looks almost identical to code written against a cloud one, which is a genuine gift: you can prototype everything against the free local model and later switch to a cloud model by changing almost nothing.
In practice this is a few lines of Python pointed at the local address where Ollama listens. No API key, no per request charge. You can run the script a thousand times in a loop and pay nothing but electricity. I have automations that hit my local model hundreds of times a day for small jobs, classification and text cleanup and first pass drafts, and doing that volume against a paid service would add up. Locally it is effectively free after hardware I already own.
Small things that trip up beginners
A few behaviors look like bugs and are normal. Models stay loaded in memory for a short while after use, then unload to free the card, so the first question after a pause feels slow while the model loads and the following ones are fast. You can keep several models installed and switch between them by name, which is great for experimenting, though a typical card only runs one at a time. And each installed model occupies a few gigabytes on disk, so a downloading spree deserves an occasional glance at your free space.
Five minutes of commands worth learning
Beyond the basic run command, a handful of others make the tool feel managed instead of poked at. There is a command to list every model you have installed, useful once you have collected a few. There is one to remove a model you no longer want, which matters at gigabytes apiece. And there is a way to inspect a model's size and settings. All short, all memorable. Learning them takes five minutes and swaps hoping it works for knowing what you have.
Making the model yours
One feature lifts this from generic chatbot to something personal, and people routinely miss it. You can write a short file containing a standing instruction, in effect telling the model what kind of assistant it is and how to respond, then build a customized version of the model from that file. Every conversation with the custom version starts already behaving your way, with nothing repeated.
I keep a few of these. One is tuned terse and technical for quick lookups. Another drafts in a particular voice. Setting them up costs nothing, and the difference in daily use is large.
You do not have to live in the terminal
Many of the popular graphical chat interfaces, and a number of code editor AI features, can point at a local Ollama instead of a paid cloud service, because they speak that same shared interface. Give the app your local address and the polished front end is suddenly running on your free local model.
The path I recommend: prove everything works in the terminal first, then put a friendlier interface on top once you know what normal looks like. You end up with the ownership and the nice experience together.
When something goes sideways
Something will wobble on day one, so here is the short list that covers almost everything. Slow responses: the model is probably too big for your card, pull a smaller one. A model fails to load: you may be short on either video memory or disk space, so close heavy programs and check free storage. The first response after a pause is slow while later ones are fast: that is the model loading, and it is normal. Anything weirder: restart the Ollama background service, which clears up a surprising number of odd states. None of these fixes require real expertise.
The honest limits
A local seven or eight billion parameter model will not match the best cloud models on the hardest work. Complex multi step reasoning and very long documents are where the gap is real, and deep niche knowledge runs thinner too. So do not throw your hardest problem at your first local model and write off the whole idea. The model is smaller, and it should be judged on the large pile of ordinary tasks where it shines. My own setup splits exactly this way: local for the common stuff, a cloud model kept around for the rare hard stuff.
That is the entire on ramp. Install, pull a small model, chat offline, go smaller if it drags, then wire your own code to the local server. Everything more advanced builds on top of exactly this, and getting one model running today already puts you ahead of most people who only ever talk about local AI.
Get new guides and videos first — join the Telegram channel.