XavierFok
← all posts

The dashboard I built to watch every automation I run

2026-08-15 · by Xavier Fok

# The dashboard I built to watch every automation I run

The point of automating something is that you stop watching it. That same property is what makes a broken automation dangerous. A script that fails while you are sitting in front of it gets fixed in ten minutes. A scheduled job that dies on a quiet afternoon can stay dead for days, because nobody is looking, and the whole arrangement was designed so nobody had to look. You usually find out when something downstream goes wrong, or when someone else notices before you do.

I run dozens of scheduled jobs and agents across my machines, and I hit that wall hard enough that I built one small tool to deal with it: a single local dashboard that shows the health of everything I have running.

The gap between the jobs

Each job on its own was fine. It ran on its schedule, did its work, and logged the details somewhere. The missing piece sat between the jobs: no single place could answer the only question I actually cared about day to day, which is whether everything is okay right now.

Without that view, answering the question means opening the logs of every job, one by one, every day. That is a manual chore, and eliminating manual chores was the reason I automated in the first place. So the checking never happened. Skipping it felt fine for weeks at a stretch, right up until the day it was catastrophic.

What the page shows

The dashboard is deliberately plain. One local page, one row per automation. Each row shows when the job last ran, whether that run succeeded, and whether the job has reported recently enough that I should have heard from it by now. Healthy rows are green and quiet. Anything broken or overdue is loud.

I open it, I get one answer, and I move on. On a normal day that takes a few seconds. That is the entire user experience, on purpose.

Jobs push their status; the page only reads

The one architectural decision that matters most: the dashboard never reaches out to check on anything. I considered a design where it probes each automation to see if it is alive, and dropped the idea quickly, because the dashboard would then need to understand the internals of every different job, and every new job would mean new probing logic.

So the flow runs the other way. When a job finishes a run, it writes a tiny record to one shared place: which job it is, when it ran, and whether it succeeded. Three facts, nothing more. The detailed logs stay wherever they always lived. The shared place is a roll call, and the dashboard does nothing except read that roll call and render it.

The smallness of the record is load bearing. Adding reporting to a job costs a line or two at the end of it, so every job actually does it. If reporting were a burden, jobs would quietly skip it, and a dashboard that covers half your automations hands you false comfort about the other half. Cheap reporting is what keeps the coverage honest, and it also means a new automation shows up on the page the first time it runs, with zero dashboard changes.

The staleness trap

Here is the failure that separates a real monitoring tool from a decorative one. Imagine two jobs. The first ran an hour ago and errored, and its record says so. The second was supposed to run hourly, and its schedule quietly broke three days ago, so it has not run at all since. Its most recent record, from three days back, says succeeded.

A naive dashboard flags the first job and shows the second as healthy green, because it only looks at the result of the most recent run. The second job is the worse failure by a wide margin. It is completely dead, and the page is reassuring you about it. A monitoring tool that comforts you while things are broken is the worst version of a monitoring tool.

The fix is to treat silence as a failure. Every job declares how often it expects to run, and the page flags anything that has gone quiet past that window, regardless of how its last run ended. A job that should have reported by now and has not deserves the same alarm as a job that ran and errored. Arguably more, since the job that errored at least tried. Catching the silent ones is the feature that earns the dashboard its trust.

Boring when healthy, loud when broken

Everything else on the page serves one design rule: make broken loud and obvious, and make healthy quiet and dull. On a good day, which is nearly every day, the page should be almost uninteresting. A glance, all calm, done. On a bad day, the problem should grab me instantly, with no scanning and no interpreting.

A dashboard that is busy and colorful all the time trains you to stop reading it, and once you stop reading it you are back to the false comfort you started with. The contrast between a dull healthy state and a screaming broken state is where all the value lives.

Pairing it with a push

A dashboard is passive. It only helps if I actually open it, which means the staleness trap can repeat one level up: I stop opening the page, and a red row sits glowing on a screen nobody looks at.

So for the failures that genuinely matter, I add a push. When something crosses from healthy to broken, a message lands on my phone whether or not I was looking. The dashboard serves the deliberate glance; the alert covers the thing that cannot wait for one.

The alert channel has to obey the boring-when-healthy rule even more strictly than the page does. If my phone buzzes every time something minor twitches, I learn to swipe the buzzes away, and the one alert that mattered goes out with the noise. I am ruthless about what is allowed to interrupt me. Genuine, actionable problems get through. Everything else waits quietly on the page for the next time I choose to look. A handful of alerts I trust completely beats a hundred I have trained myself to ignore.

What it changes, and what it does not

I want to be careful about the claim here, because monitoring is easy to oversell. The dashboard fixes nothing. A job is exactly as likely to break with the page running as without it. Reliability comes from other machinery entirely: the retries, the queues, the validation inside each job.

What the dashboard changes is when I find out. A silent failure I would have discovered days later becomes a visible failure I catch the same day. That is the entire unlock, and it turns out to be enormous, but it is awareness rather than prevention, and you need both kinds of machinery.

The page is also one more component that can itself fail, which is a slightly uncomfortable thought: the thing watching my automations is an automation. The honest answer to that is simplicity. Mine reads records and renders them, and does nothing else, precisely so it stays the most reliable piece in the whole setup. The less it does, the less there is to break.

Build the view before you regret not having it

The broader lesson I took from all this: you cannot comfortably walk away from systems you cannot see. People imagine that the endpoint of automation is setting things up and never thinking about them again. Real confidence comes from somewhere else, from knowing that when something breaks, and it will, the break becomes visible immediately. That property has a name, observability, and it is what lets the amount of automation grow without the anxiety growing to match. Every job I add reports to the same place and lands on the same page, so more automation adds nothing to my watching burden.

If you are running more than a couple of automations, build the view early. It needs no polish. One shared place every job reports to, one plain page that reads it, broken made loud, healthy made dull, and silence treated as a failure. That is the difference between automation you nervously hope is working and automation you can leave alone with a clear head. Building the jobs was never the scary part. Running them blind was.

Get new guides and videos first — join the Telegram channel.