How I monitor my servers with AI summaries
# How I monitor my servers with AI summaries
At some point I stopped reading my own monitoring output. The emails still arrived, two hundred lines of log excerpts at a time, and I glanced at them and closed them, day after day, until closing them became automatic. Then something actually broke, and while digging through history I realised I had been skimming past the early signs for weeks. Everything had been technically fine. The alerts were technically quiet. And the information that would have saved me was sitting in output I had trained myself to ignore.
The fix I landed on was to put a model in charge of writing the summaries and to keep every other part of the system exactly as dumb as it was before. The line between what the model touches and what it never touches is the entire design, so this post is mostly about that line.
Monitoring output is written for machines
If you have taken server monitoring seriously at any point, the failure pattern is familiar. You set up metrics and log shipping, you build a dashboard, and the dashboard goes unwatched. The alerts fire too often, so you tune them. They fire too often again, so you silence a few. Eventually the check that would have caught your next outage is disabled because it cried wolf one time too many.
None of that means the tooling is bad. The real issue is that raw monitoring output is machine shaped. A wall of metric charts or a page of log lines tells you everything and communicates almost nothing. Somebody has to translate it into a sentence a human can act on, and no small operation has a person sitting around reading logs all day. In practice nobody reads them until the service is already down, at which point you are doing archaeology instead of prevention.
The fleet this has to cover
My setup is small and ordinary. A handful of Linux servers: a couple doing web and proxy work, one running automation jobs, and a database host. They emit the usual stream of CPU load, memory, disk io, network throughput, plus application level signals like request counts, error rates, and job durations.
The scale matters more than the specifics. The full picture of any thirty minute window across this fleet is a few hundred log lines and a one page metric snapshot, which fits comfortably inside a model context window. If you run a hundred node cluster producing terabytes of logs a day, this pattern does not transfer directly. For a small fleet it fits without any summarisation tricks at all.
Collection stays deterministic
Every thirty minutes a small script pulls the latest system metrics from each host, grabs recent log lines from the services I care about, and writes everything into one plain text file. There is no database and no streaming pipeline. Just a text blob representing what the last half hour looked like.
Two deliberate choices live in that script. First, it contains no model. It gathers data and writes it down, nothing else. Second, it collects only the things I would check by hand if I suspected a problem: disk space and io rates, memory, CPU, network, job queue depth, error counts per service. Collecting everything possible was tempting and I skipped it on purpose. Every extra field is noise the summary has to wade through.
There is a bonus property I did not plan for. If the collection script dies, the summaries stop arriving, and I notice the silence within a few hours. The script breaking is its own health check, and catching it requires no intelligence anywhere in the system.
The prompt asks for very little
The collected text goes to a model with a short prompt: read this block of logs and metrics from the last cycle, write three to five plain language sentences about what changed and what looks unusual, and flag anything trending the wrong way.
I am deliberately not asking for a diagnosis. No recommended fixes, no severity ratings, no opinion on whether anyone should be paged. The job is translation, from a wall of numbers into something I can read in about fifteen seconds while making coffee. The output lands in my chat and that is the whole interaction.
Getting the prompt right took a few weeks. My first version asked the model to summarise and advise, and it produced four confident paragraphs ending with suggestions like consider investigating the network configuration. Useless. Once I pinned down the length and the tone explicitly, the output became short, direct, and free of false confidence. Models do this job far better when you tell them exactly what shape the answer should take.
The model never fires an alert
Here is the line, and I hold it firmly. The model writes summaries. Alerting is a separate system made of dumb deterministic rules in a config file. Disk above ninety percent, alert. Memory sustained above eighty percent for ten minutes, alert. Process not responding, alert. No model anywhere near it.
The summary can mention that disk usage has been climbing all week, and that is useful context. Whether that observation becomes a page at 2am is decided by a number crossing a threshold, never by a sentence.
The reasoning is about auditability. When a threshold fires, I know exactly why: disk is at ninety two percent, there is the number, and I can decide whether the threshold was set well. When a model fires, I have to trust its reading of that particular context window, with that particular prompt, on that particular run. That is far too much surface area for the mechanism that wakes me up. I want alerting boring and fully testable, and I want summaries smart and readable. Two different jobs, two different tools.
What it has actually caught
Two saves stand out. Once the summary noted that a service had restarted seven times in two days, above its weekly average, with the count climbing gradually over four days. No threshold had fired because my rules watched for a process being down, and the process kept coming back. The cause was a slow memory leak, and I fixed it before the service fell over for real.
Another time it mentioned that disk write latency on the database host had been running higher than usual for several cycles. Nothing alarming, just different. A backup job had drifted to the wrong time and was competing with production writes. Again, no rule was breached. The summary made the pattern readable enough that I went and looked.
Both catches share a shape: a slow drift below every threshold, visible only if someone reads the data daily. The summary layer is what makes daily reading sustainable.
Where it fails
The model misses things. If a thirty minute window happens to not surface the right signal, the summary reads clean while a problem keeps growing. The deterministic thresholds still do the real safety work underneath, and I treat the summaries as a readability layer on top of proper monitoring rather than a replacement for it.
It also over flags, calling out metrics that are a touch higher than the last cycle when the difference is noise. I read those with a light touch and move on unless something else points the same direction.
The subtlest failure took me a while to spot. When the same anomaly repeats daily, the model starts treating it as baseline and quietly stops mentioning it. A nightly cleanup job of mine ran slow for about a week, appeared in the summaries at first, then vanished from them once slowness had become the model's idea of normal. I only caught it by looking at raw numbers directly. The model cannot tell seasonal variation from a genuine trend, and it does not remember what normal used to look like. That judgment stays with me.
Cost, and who this fits
The economics are small. One summarisation call every thirty minutes, a few hundred lines in, five sentences out. Even around the clock it comes to a few dollars a month on a metered API, and the collection script runs on a server I already pay for. The real investment was a few afternoons of building and prompt tuning, and it has run unchanged for months since. I will say plainly that it is a paid model call. Cheap is the honest word, free is a lie, and you should price your own call frequency before copying the design.
As for fit: a serious production operation with SREs and real observability tooling is not missing this. A single personal server probably does not justify the overhead. The sweet spot is exactly where I sit, a few servers doing real work, no dedicated ops person, and monitoring that technically functions while producing more noise than signal.
The frame that keeps the whole thing honest is that the model is a translator. It turns machine shaped output into human readable sentences, and it does that well. Deciding whether something is a real problem, whether to roll back, whether anyone gets woken up, all of that stays with me and with rules I can read in full. Better information from the model, decisions from deterministic code and a human. That is the only version of this I would run on infrastructure I care about.
If you want to see more builds like this, with the tradeoffs left in, [there is more on xavierfok.com](/).
Get new guides and videos first — join the Telegram channel.