XavierFok
← all posts

How I keep AI agents reliable enough to trust

2026-08-15 · by Xavier Fok

# How I keep AI agents reliable enough to trust

An AI agent that demos beautifully and an AI agent you can leave alone with a job are two different machines. The demo runs once, on a clean example, with a person watching. The trustworthy version runs a thousand times, on inputs nobody anticipated, while you are asleep, and it either does the right thing every time or says clearly that it could not.

Closing that gap is most of the real work of building agents, and almost none of it involves getting a smarter model. I run agents against my real tools every day, and I have watched them fail in roughly every way an agent can fail. What follows are the failure modes I actually see and the specific constraints that keep my agents reliable enough to trust.

The four ways agents fail

You cannot guard against what you refuse to name, so let me name them.

The first is wandering. You hand an agent a goal and some tools, and instead of taking the short path it meanders. It calls a tool it did not need, revisits a decision it already made, spends three steps on something that needed one. It usually arrives eventually, slowly and expensively, and sometimes it talks itself into a small loop instead.

The second is grabbing the wrong tool. The agent misreads the situation and reaches for something that does not fit, usually because a description was unclear or the context was ambiguous.

The third is the confident mistake. The agent does the wrong thing while being completely sure it is right. This one is dangerous precisely because nothing in its behavior signals that anything went wrong.

The fourth is the quietest and the worst: silent partial failure. The agent completes most of the job, fails on one piece, and reports success anyway. You find out much later that a step never actually happened.

Better models do not close the gap

Here is the uncomfortable part that took me a while to accept. Waiting for a smarter model fixes none of this. Better models wander less and misjudge less, but they still do all of it, just less often, and less often is a different thing from never.

What actually makes an agent reliable is constraint. You make the agent's world smaller and clearer until the ways it can go wrong shrink toward zero. That means giving up some of the autonomy that made the agent exciting in the first place, and that trade is almost always the correct one. Reliability is something you build by limiting how much rope the agent has.

Constraint one: narrow the scope

An agent pointed at one well defined job is far more reliable than an agent told to handle whatever comes up. So I do not build one big agent that does everything. I build small agents, each with a tight goal. A narrow agent has fewer decisions to make, which means fewer decisions it can get wrong.

When I catch myself writing a sprawling mission statement for an agent, that is usually the signal to split it into two smaller agents with clearer jobs.

Constraint two: cut the tools

This one runs against instinct, because more tools feels like more capability. But every extra tool is another choice the agent has to make correctly, and another chance to pick wrong. An agent with three carefully chosen tools and a clear goal behaves dramatically better than the same agent with twenty tools and the same goal. The twenty tool version has more ways to wander and more wrong things to grab.

So each agent gets the smallest tool set that lets it finish its actual job, and nothing extra just in case. Just in case is where wandering comes from.

Constraint three: validate with code that has no opinions

I never accept an agent's result just because it was returned confidently. A language model delivers every answer in the same fluent, assured tone whether it is right or completely wrong, so the tone carries no information. Confidence is not evidence.

The practical shape of this: define ahead of time what a valid result looks like, then reject everything else with plain deterministic code. If the agent should return one of five categories, accept only those five and flag anything outside them. If it returns a quantity, check that it is a number in a plausible range. If it returns a date, check that it parses. And if it claims to have completed three steps, go and verify in the real world that all three actually happened, instead of taking its word.

That last check is what catches silent partial failure, the sneakiest one, because everything looks done until you confirm that it is.

The trick is to keep the validation dumb on purpose. A second agent checking the first agent gives you two fuzzy things and no solid ground. What you want as the final judge is boring code with no creativity, a fixed set of checks that pass or fail. The agent brings the judgment. The validation brings the certainty. The model proposes, deterministic code disposes.

Constraint four: gate the irreversible

This is my hard rule, and I do not bend it. An agent never takes an action that would be painful to undo without a human approving it first. Deleting things. Sending messages to other people. Spending money. Publishing anything public.

For all of those, the agent proposes the action and waits. I look at the proposal and approve it, or I do not. The agent still gets to be smart and suggest the move. I keep the final say on everything with real consequences. This single rule is what lets me hand agents genuinely useful capabilities without lying awake, because the worst an unattended agent can do is the reversible things. The dangerous things wait for me.

Constraint five: log the choices, not just the outcomes

People skip logging until the first time an agent does something baffling. Early on, I logged that an agent ran and whether it finished, which felt sufficient right up until a strange result appeared and my log explained nothing.

What I actually needed was the decision trail. Every tool the agent called, the inputs it passed, what came back, and what it chose to do next. With that trail, a misbehavior becomes a story I can read from start to finish, and I can usually point at the exact step where things went sideways. Without it, the agent is a black box that occasionally does something odd for no visible reason.

The rule I settled on: log the agent's choices, not merely its outcomes, because the choices are where the failures live. It costs a little storage and repays it many times over.

The trail did something else I did not expect. Watching an agent's decisions across hundreds of runs, seeing it mostly choose sensibly and catching the rare odd call, is what let me genuinely believe it was safe to leave running. You do not trust an agent because someone says the model is good. You trust it because you watched it behave, run after run. Trust is earned through visibility, and visibility is something you build.

What this looks like assembled

Take the agent that triages incoming messages for me. Its scope is narrow: read a message, decide what kind of thing it is. Its tools are few: read the message, file it into a category, flag it for my attention. It has no tool that can reply to anyone, because replying to a real person is exactly the kind of irreversible action that should never happen unattended. Its output is validated against a fixed list of allowed categories, so it cannot invent a nonsense one. And every decision it makes lands in the log, so a miscategorization comes with its own explanation.

That agent is genuinely useful, and I can leave it running. The model behind it is fallible. The box around it is what makes it safe.

The price is autonomy

I want to be straight about the cost, because it is real. A constrained agent is less autonomous and less impressive than the open ended version. The wide open agent with twenty tools and a vague mission does flashier things and makes a better demo.

I am not building demos. I am building things I can leave running, and the price of trust is autonomy. I will take the narrow agent that does its one job reliably and asks before anything risky over the brilliant generalist that occasionally surprises me in a way I cannot undo. I make that trade every time, without regret.

Ask whether it needed to be an agent at all

One last point that sits underneath all of this. Once you have constrained an agent down to a narrow job, a few tools, and validated output, it is worth asking whether it needed to be an agent in the first place, or whether a plain script with one model call would have done the same work. Often the script wins. The agent shape earns its place only when the path genuinely has to be decided at runtime, based on messy input. When that is true, these constraints are how you make it safe. When it is false, the most reliable agent is the one you replaced with deterministic code.

The summary is short, even though the work is not. An agent becomes trustworthy when you make its world small and clear. Narrow the scope. Cut the tools. Validate the output with code that has no opinions. Gate the irreversible behind your own approval. Log every choice. Do that, and you can walk away while it runs. Skip it, and you own a clever demo that will eventually do something you cannot take back, at a moment when nobody is watching.

Get new guides and videos first — join the Telegram channel.