XavierFok
← all posts

Building a private offline chatbot over your own documents

2026-08-15 · by Xavier Fok

# Building a private offline chatbot over your own documents

An AI you can question about your own files, running entirely on your own machine, with nothing uploaded anywhere. A couple of years ago that was a serious engineering project. Today it is a weekend build, and I run exactly this over my notes and records every day. This is how the pieces fit together, what makes the answers trustworthy, and where the honest limits sit.

Why build one

Privacy and permanence. If your documents are sensitive, personal records or business material or anything you would rather keep off other people's servers, a local setup keeps every word on your machine. And a setup you control keeps working for as long as your computer does, while cloud tools get repriced, redesigned, or discontinued underneath you. The goal is an AI that knows your material without you handing that material to anyone.

The core idea: look it up, then answer

A language model on its own knows nothing about your private files. It was trained on public data. The trick is to skip retraining entirely, since retraining on your documents would be slow and expensive, and let the model look things up instead. When you ask a question, the system first finds the most relevant pieces of your documents, then hands those pieces to the model along with the question, and asks it to answer from that material. The model supplies reasoning and language. Your documents supply the facts.

That lookup then answer pattern is the entire architecture.

The four pieces

You need surprisingly little. A reader that chops documents into manageable chunks, since whole books cannot go to the model at once. An embedding model that turns each chunk into a numeric fingerprint of its meaning, so search works by meaning instead of exact words. A store that holds those fingerprints and finds the closest ones fast. And a language model to read the retrieved chunks and write the answer. Every piece has a good local option, which is what keeps the whole thing private and free to run after the hardware.

Chunking is half the battle

How you split documents quietly determines answer quality, and most tutorials skate past it. Chop too small and each piece loses its context, leaving the model with fragments that mean little. Chop too big and the search goes blurry, because each piece covers too many subjects at once. The sweet spot is a few paragraphs per chunk, each holding one coherent idea with enough surrounding context to stand alone. Get this wrong and nothing downstream saves you. Time spent here pays off more than any model upgrade later.

Embeddings, or search that understands meaning

An embedding model reads a chunk and outputs a list of numbers representing what it means. Two chunks about the same topic end up with similar numbers even when they share no vocabulary. Ask about your vacation budget and the system can surface the chunk about trip spending, despite the word budget never appearing in the document. Embedding models run locally without strain, far lighter than the language model itself, so the semantic search stage stays on your machine too.

The store holding those fingerprints is called a vector store, and for a personal collection a simple local one is plenty. It sits on disk, stays private, and returns the closest handful of chunks in a blink. Nothing heavy is required at personal scale.

The model's actual job

The language model at the end can stay modest. It reads facts from the chunks you hand it rather than dredging them from its own memory, so a solid seven or eight billion parameter model on a home GPU handles the job well. Its task is reading a few relevant passages and answering a question about them in plain language. The whole stack fits comfortably on one card.

The main failure mode: retrieval misses

The system can only answer from what retrieval finds. If the right chunk never gets retrieved, the model never sees it, and it will either admit ignorance or guess. Most wrong answers trace back to retrieval handing over the wrong material rather than to a dumb model. Your question may use words so different from the document that the search never connects them. Or the answer may be scattered across many chunks with only some pulled in. The model takes the blame anyway. When quality disappoints, look at retrieval before you touch anything else.

Making the answers trustworthy

First, make the system show which chunks it used for each answer. A visible source lets you verify instantly and catches the times it guessed. Second, instruct the model explicitly to answer only from the provided material and to say it does not know when the material comes up short. That single instruction removes most of the confident invention. Third, when an answer is wrong, inspect what got retrieved before touching anything else. Nine times out of ten the right chunk simply never showed up.

Set expectations to match what the machinery does. It is excellent at finding and summarizing specific facts buried in your files and at pulling together what your notes say about a topic. It is weaker at questions spanning the entire collection at once, at counting, and at anything requiring inference across many scattered places. Used as a brilliant search and summarize layer over your own knowledge, it is genuinely transformative for finding things you half remember writing down.

Fixes that stack

Retrieval misses have concrete cures. Retrieve more chunks per question, so a right answer ranked third still reaches the model. Improve the chunking, since better pieces make better matches. And have the system rephrase the question a couple of ways before searching, casting a wider net so different wordings can connect. Each change is small, and stacked together they turn a system that misses half the time into one that reliably finds what you need.

Reranking adds one more level. The initial search is fast and a little blunt, grabbing chunks that are roughly relevant. A reranking step hands those candidates to a smarter, slower model that judges which ones genuinely answer the question, keeping only the best few. Cast a wide net cheaply, then choose carefully. It costs a moment per query, runs locally like everything else, and the jump in answer quality is usually worth it on a larger collection.

Keeping the index honest

Documents change. Notes get added, files get updated, old material gets deleted, and a chatbot answering from a stale index gives quietly wrong answers. The clean approach watches the documents folder and re embeds only the pieces that changed, patching the store instead of rebuilding it. For a modest personal collection, even occasionally re indexing everything is fine. What matters is a habit or an automation that keeps the index matching reality.

Scanned pages and other formats

Plenty of what you might want to search lives outside text files: scanned pages, photographed documents, PDFs that are really pictures. Those need a text recognition step first, which also runs locally, and the extracted text then flows into the same chunk, embed, store pipeline as everything else. The architecture stays identical. You bolt a converter onto the front and end up with one private chatbot over the whole pile regardless of what format anything started in.

The pull-the-cable test

Genuinely private means every component is local. The embedding model, the vector store, the language model, the text recognition, all of it. If one piece quietly calls a cloud service, your documents are leaving the machine and the point of the exercise is gone. Some convenient tools default to cloud embeddings or a cloud model unless told otherwise, so check each link in the chain. The proof at the end is satisfying: pull the network cable and watch the whole thing keep answering. That one test settles it.

Scale, honestly

People imagine pointing this at a million documents on day one, and mostly they have nowhere near that. Personal notes, important records, reference material: usually a modest pile that a simple setup handles with ease, in the hundreds or low thousands of documents. Very large collections eventually demand more careful engineering to keep the search fast. Start with what you actually have, get it working well, and reach for heavier machinery only if you genuinely outgrow the simple version. Most people never do.

The build, end to end

Chunk your documents thoughtfully, since that is half the quality. Embed each chunk with a local embedding model so search works by meaning. Keep the fingerprints in a simple local vector store. At question time, retrieve the most relevant chunks, hand them to a local language model, and instruct it to answer only from that material while citing which chunks it used. Everything runs on your machine, so the system is private by construction and free to run after the hardware. And when an answer goes wrong, look at retrieval first.

A private AI that knows your own documents and never phones home is one of the most useful things local AI can give you. More builds like this one, with the limits spelled out, live on [the home page](/).

Get new guides and videos first — join the Telegram channel.