XavierFok
← all posts

When an AI agent beats a plain script (and when it does not)

2026-08-15 · by Xavier Fok

# When an AI agent beats a plain script (and when it does not)

I keep seeing the same mistake, and I have made it myself more than once. A task shows up that a ten line script would handle cleanly, and the builder reaches for an AI agent instead. A language model gets wired up with tools and a reasoning loop, and everyone feels very current. Then the job that used to finish in half a second takes thirty, bills a small amount of money on every run, and once in a while does something nobody can explain.

I run both kinds of system in a real daily operation. Scripts move my files and reconcile my records. Agents read messy text and make judgment calls where no fixed rule survives. After living with both for a while, the choice between them has collapsed into a single question for me, and I want to lay that question out properly, because the current hype pushes everyone toward agents for everything, and that path ends with a slow, expensive version of a job code had already solved.

Pinning down the two words

The word agent has gone fuzzy, so let me define both sides before arguing about them.

A plain script is code that follows a fixed path. Step one, then step two, then step three, in the same order on every run. Its entire behavior is decided before it executes.

An agent is a language model that decides its next move while the work is happening, usually with a set of tools it is allowed to call. Nobody wrote its path down in advance, because the path depends on what it finds along the way.

That is the distinction that matters. A script knows its route before it starts. An agent works its route out as it goes. Every other difference between the two follows from that one line.

The rule

If I can write the steps down in advance, I write a script. The script will be faster and cheaper than any agent doing the same work, and when it breaks, it will break in a way I can read.

An agent becomes the right tool only when the correct next step depends on something I cannot know until the input is in front of me. When no rule I could write ahead of time survives contact with the data, the model's judgment is the thing I am paying for. Everywhere else, that judgment is a cost with no matching benefit.

What the agent actually costs

The tradeoffs are concrete, and they are larger than people expect.

Speed first. A script completes a simple transformation in a few milliseconds. An agent has to ship the request to a model, wait for it to reason, receive the reply, possibly run a tool, send the tool result back, and wait again. Each round trip is measured in seconds. A task with five decision points can take half a minute where the script version would have finished before your finger left the enter key.

Money second. Every model call has a price. A fraction of a cent means nothing on one run. Multiply it across a task that fires ten thousand times a day and the agent is suddenly thirty or forty dollars daily, plus hours of cumulative waiting, for work a script performs free and instantly. Before shipping an agent on any recurring job, I want to know that daily number, because the per run cost looks harmless right up until you remember the run count.

Then there is the texture of failure, which is the cost people discover last. A script fails loudly. It reaches a line, throws an error, and the error names the line. I read it, I fix it, done. An agent fails softly. It makes a slightly wrong call, misreads a field, picks the wrong tool, or does the wrong thing with total confidence. Nothing crashes. The output just quietly stops being trustworthy, and tracing why takes far longer than reading a stack trace.

Where an agent earns its keep

Three situations genuinely justify the cost.

The first is branching on input you cannot predict. Picture triage on a stream of incoming messages, where some are urgent, some are junk, and some need routing somewhere specific. Human writing is too varied for a clean rule set. You can try keyword rules, and what you get is a growing tangle that snaps on the first message phrased in a way you never anticipated. Reading a messy message and judging what kind of thing it is happens to be exactly what language models are good at, and no script matches them there.

The second is turning loose text into structure. Scripts are excellent with tidy input. Hand one a clean table and it transforms it perfectly forever. Hand it a page of free flowing prose that happens to contain the three facts you need, and the extraction rules you write will be miserable to build and brittle in production. A model reads the paragraph and fills the fields. Anywhere the input is loose human text and the output needs to be structured, a model step has a real claim on the job.

The third is tool choice. When there is a single tool and it always gets called, a plain function call covers it. When there are five tools and the right one depends on what the request turns out to be, letting a model examine the request and pick among them is the agent pattern doing the job it was built for.

Where a script should have stayed

The reverse cases are just as clear.

Fully deterministic work belongs in code, always. Move these files, rename them to this pattern, upload them on this schedule. No judgment exists anywhere in that path, so a model contributes nothing except delay, expense, and a fresh way to be wrong.

High volume simple work belongs in code even when a single agent run would have been acceptable, because volume multiplies every weakness. I get twitchy whenever I see an agent wrapped around a frequent, simple task.

Work where mistakes are expensive and the rules are actually clear also belongs in code. When you can write the rule and errors genuinely hurt, predictability is worth more than flexibility.

Volume multiplies everything

The volume point deserves its own section, because it changed how I build.

A good script is deterministic. Same input, same behavior, every time, forever. An agent is probabilistic by nature, so even a strong one will occasionally choose differently on similar input. On a task that runs once, that variability is invisible. On a task that runs ten thousand times a day, a tiny rate of odd decisions turns into dozens of strange outcomes daily, each of which someone now has to notice and handle. The more often a job runs, the more the boring predictability of code is worth, and the more the fuzziness of a model costs in surprises.

The hybrid I actually build

Most of my real systems are neither pure script nor full agent, and the middle shape is badly underused.

Almost every real task is a deterministic skeleton with one or two genuinely fuzzy joints. Fetching, looping, storing, scheduling, all of that is known path, so it gets written as plain code. The model gets called at the single step that needs judgment, and nowhere else. The scraper I run is built exactly this way. The pipeline is an ordinary script from end to end, and a model touches one step only, the parsing of loose text that no rule handled well. I pay for intelligence at the joint that needs it, and everything around that joint stays quick and predictable.

Nobody has to choose between a dumb script and a fully autonomous agent. A script with a model bolted into the one spot that needs judgment beats both, far more often than either side of the argument admits.

A decision I watched myself make

One concrete case, so this stays honest. I had incoming items that needed sorting into a handful of buckets. Sorting feels like judgment, so my first instinct said agent. Then I asked the actual question. Could I write the rule down?

For most items, I could. The large majority landed in obvious buckets based on simple properties a script can check without any judgment at all. Only a small slice was genuinely ambiguous, where the right bucket really did depend on reading messy text. So the system became a script that handles the clear majority deterministically, with the ambiguous slice routed to a model call. The result was mostly a cheap, reliable script with a small expensive smart part sitting exactly where the smartness was needed.

Had I followed the instinct and built a full agent, I would have paid model prices and accepted model fuzziness on the ninety percent of items that required neither.

The question that settles it

Before reaching for an agent, the question worth asking is never whether the task is hard or interesting. It is whether I can write the rule down. A rule I can write, even a long and fiddly one, belongs in code, because code following a written rule runs free at the margin and behaves the same way every single time. The model's territory begins only where no written rule captures the decision, which usually means understanding messy human input. Even then, as the sorting example showed, only a slice of the task tends to live in that territory, while the rest is mechanical work a script should own.

An agent is a powerful tool for one specific shape of problem, the shape where the path stays unknown until you look at the input. For every other shape, boring deterministic code is the correct engineering choice, and it is the one I reach for first.

Get new guides and videos first — join the Telegram channel.