XavierFok
← all posts

Single points of failure in a one person business

2026-08-10 · by Xavier Fok

A single point of failure is any part of your setup that, if it stops working, stops everything downstream of it. In a company with a team, that role usually gets covered by someone else when a person is out sick or quits. In a one person business, there's no one else. Every single point of failure you have is a bet that nothing goes wrong on a specific day, with a specific machine, or with a specific account.

I run content automation, local rendering, and a handful of AI agents wired into real tools every day. Most of what breaks in that setup isn't dramatic. It's small, boring failures that cascade because nobody was watching for them. This is what I've learned about where those failures actually sit, and what's realistic to do about them.

What counts as a single point of failure

Not everything that could break is a single point of failure. If a script fails and you notice within the hour and rerun it, that's a hiccup. A true single point of failure is something where the failure is silent, or where there's no fallback path at all. Two conditions, not one: it has to be able to break, and there has to be nothing else that catches you when it does.

That distinction matters because it tells you where to spend effort. You don't need a backup for everything. You need one for the things that fail quietly and have no second path.

The obvious ones

Some single points of failure are easy to name because they're physical or contractual. One laptop. One bank account. One domain registrar. One hosting provider. One phone number tied to two-factor authentication on everything else you own. If you've been running a solo operation for more than a year, you've probably already had at least one of these bite you, whether it was a dead hard drive or a payment processor freezing an account without warning.

These get talked about a lot, so I won't spend more time on them here. What gets talked about far less is the layer that's specific to running AI in production, because that layer is newer and the failure modes are less familiar.

The AI-specific ones that actually get people

The first is the API key. If your content pipeline, your agent, or your automation script authenticates against one cloud AI provider with one key, that key is a single point of failure in three different ways: it can be revoked, it can hit a rate limit at the exact moment you need it, or the provider can have an outage. None of these are rare events. They're the normal operating conditions of depending on infrastructure you don't control.

The second is the model itself. If your workflow is built around the specific behavior of one model version, an unannounced update to that model can quietly change your output quality without any error being thrown. Nothing crashes. The automation keeps running. It just starts producing worse work, and because there's no failure signal, you might not catch it until a human reviews the output days later.

The third is the automation script that has no one else who understands it. This is the one people underestimate most. If you wrote a script that chains a few AI calls together with some local logic, and you're the only person who's ever read that code, then you personally are the single point of failure for maintaining it. If you're unavailable for two weeks, the pipeline still runs, but the moment it breaks, it stays broken until you're back.

Local vs cloud: different failure shapes, not different amounts of risk

There's a temptation to think local AI is inherently safer than cloud AI because you're not depending on someone else's servers. That's true for uptime, but it swaps one failure mode for another. A local model on your own hardware doesn't go down because a provider had an outage. It goes down because your GPU dies, your power goes out, or the machine it's running on needs a restart for something unrelated and you forget to bring the service back up.

Cloud AI's failure mode is external and out of your hands. Local AI's failure mode is internal and entirely your responsibility. Neither one is automatically the safer bet. What matters is whether you've actually thought through which failure mode you're exposed to, and whether the workflow can tolerate it. A same-day content draft that goes out on a fixed schedule needs different protection than a batch job that can run whenever the machine is free.

Agents and MCP add more moving parts before they remove any

Wiring an AI agent into real tools through something like MCP means the agent isn't just generating text anymore. It's calling a search tool, reading a file, sending a message, maybe touching a database. Each of those connections is a dependency. If one tool in that chain changes its interface, goes offline, or returns something the agent doesn't expect, the whole task can fail partway through, sometimes leaving things in a half-done state that's worse than not starting at all.

This is the part that gets oversold. People talk about agents as if adding more tools makes the system more capable and more resilient at the same time. In practice, every tool you connect is one more thing that has to keep working correctly, with no team member to notice when it quietly stops. More integrations mean more surface area for a small change somewhere else to break your workflow without warning. That's not a reason to avoid agents. It's a reason to know exactly which tools in the chain are load bearing and which ones are optional.

What actually helps, in practice

I don't believe in trying to eliminate every single point of failure. That's not realistic for one person, and chasing full redundancy on everything is its own way of never shipping anything. What's worked for me is being selective about where redundancy actually pays off.

For anything that runs on a schedule without me watching it, I want a way to know it failed, not just a way to prevent failure. A silent failure is worse than a loud one. If a job doesn't run, or runs and produces something obviously wrong, I want that surfaced somewhere I'll actually see it, not buried in a log file I only check when something feels off.

For anything tied to one external account, whether that's a cloud provider, a payment processor, or a hosting account, I try to know in advance what the manual fallback looks like. Not build it out fully in every case, that's not always worth the time, but know what I'd do in the first hour if it went down, so I'm not figuring that out for the first time during an outage.

For scripts and automations, the test I use is simple: could someone else, or a future version of me who's forgotten the details, understand what this does well enough to fix it at 2am. If the answer is no, that script is a single point of failure even if it's never crashed once.

The one you can't automate away

Here's the part that doesn't get said enough. Even with good redundancy, good monitoring, and clean automation, you are still the single point of failure for judgment calls. Automation can catch a failed job and retry it. It can't decide whether a piece of content is actually good, whether a business decision is the right one, or whether something that looks fine on the surface is actually a problem. That decision-making layer sits with you, and no amount of AI tooling changes that. The goal of good automation isn't to remove yourself from the business. It's to make sure the parts that don't need your judgment stop demanding it, so the parts that do get your full attention.

If you're building something similar and want to see how this plays out in practice, more of it is on the [home page](/).

Get new guides and videos first — join the Telegram channel.