The work I delegated and took back
Delegation is not a switch you flip once
I run AI agents against real production systems every day: content pipelines, rendering queues, provisioning scripts, inbox triage. The pitch you hear about agents is usually "hand it off and walk away." That is not how it works in practice. Delegation is a setting you keep adjusting, and some of the settings I tried had to be walked back.
This is not a story about AI failing. Most of what I handed off is still handed off. This is about the specific jobs where I gave an agent more rope than the job could tolerate, watched what happened, and pulled the rope back in. If you are building your own agent workflows, the interesting information is in the failures, not the wins.
The job: batch content authoring
I run a content engine that spins up articles across several brands in parallel. Early on, I let an agent handle the full loop for one brand: pick a topic from an inventory, draft the article, and push straight into the publish queue.
The problem showed up downstream, not in the drafting step itself. The agent was fine at writing to spec. What it was not good at was catching when a topic had already been covered under a slightly different title, or when a batch of drafts drifted out of order relative to the queue they were supposed to feed. A silent reordering between the authoring stage and the refill stage meant articles could sit staged without ever reaching the publish step, and nothing in the pipeline flagged it as an error because from the agent's point of view, the file existed and the job was done.
I took the "draft to publish" pipeline back to a staged model: agents draft, but nothing crosses into the live queue without a batch review where I check ordering and check for topic overlap against what is already published. The drafting stayed delegated. The gate before publish did not.
The job: proxy and modem fleet health checks
I manage a fleet of physical modems and proxy servers, and for a while I had an agent polling status and auto-restarting anything that looked "down." This is the kind of job that sounds perfect for automation: repetitive, rule-based, high volume.
It went wrong because "down" was ambiguous in a way I had not accounted for. A modem showing offline in the monitoring layer was sometimes actually online, and the monitoring layer itself was reporting stale data because of how it queried external IP detection rather than the actual customer-facing path. An agent restarting servers based on that signal was, in effect, cycling healthy hardware because it trusted the wrong column of data.
I did not remove automation from fleet monitoring. I removed the auto-restart action and kept the detection. The agent still flags anomalies, but a restart against physical hardware now requires me to look at the real traffic path first, not just the tool's own status page. The lesson was not "don't automate infrastructure." It was "don't let an agent take an irreversible action on a signal that has a known false-positive mode you have not fixed yet."
The job: partner and vendor emails
I tried delegating first-pass replies to partnership and vendor emails: an agent reads the inbox, drafts a response, and sends anything routine. This one I walked back fastest, within a week.
The failure mode was not that the drafts were bad. Most of them were fine. The problem was tone and commitment. An agent replying to a partner pitch will, if you are not careful with the system prompt, imply agreement on things that were never actually agreed, because it is optimizing for a helpful-sounding reply in the moment rather than for what I am actually willing to commit to weeks later. I had one case where a draft reply implied I had accepted terms on a partner swap that I had only verified in part.
Now agents draft, but nothing with a partner, vendor, or dollar figure attached to it sends without me reading it first. Routine internal notifications, sure, those go straight out. Anything external and negotiated does not.
What actually stayed delegated
It is worth being specific about what I did not take back, because the point of this article is not "AI does not work." Local rendering queues, where the job is deterministic (take this script, produce this video, follow this spec) have stayed fully automated for a long time without incident, because the failure mode is visible immediately: either the file renders correctly or it does not, and there is no ambiguity for an agent to misjudge.
Routine status logging, low-stakes internal drafts, and first-pass code review on scripts I am going to read anyway also stayed delegated. The pattern across everything that stayed automated is that the cost of being wrong was low and immediately visible. The pattern across everything I took back is the opposite: the cost of being wrong was either delayed (a queue silently stalling for days before anyone notices) or hard to reverse (a partner email implying a commitment, a physical restart on hardware).
The actual rule I use now
Before I hand a task to an agent unattended, I ask two questions. First, if the agent gets this wrong, will I find out immediately, or will it surface days later as a mystery? Second, if it gets this wrong, can I undo it cleanly, or does it touch something external, like a person's inbox or a piece of physical hardware, that does not roll back.
Anything that fails both questions gets a human gate before the consequential step, even if the agent does the actual work. Anything that passes both stays fully automated, because the worst case is cheap and visible.
This is a mundane rule, and it is not specific to AI. It is the same rule you would apply to any junior team member you were onboarding: give them the reversible, visible work first, and keep a review step on anything that is not. The difference with an agent is that it will not tell you when it is unsure, and it will not flag an edge case it does not recognize as an edge case. It will just execute, confidently, on whatever signal it was given. That confidence is useful for throughput and dangerous for judgment calls, so the judgment calls stay with me.
What I'd tell someone setting up their first agent workflow
Start every new automation with a human gate on the consequential step, even if that feels slow. Once you have watched it operate for a while and you understand its actual failure modes on that specific task, not failure modes in general, decide whether to remove the gate. Do not remove gates based on how the tool performed on a different task. A model that handles content drafting well tells you nothing about how it will handle ambiguous monitoring data or externally facing commitments, because those are different failure surfaces entirely.
None of this is about whether the underlying models are capable. It is about matching the blast radius of a mistake to how much unsupervised room you give the thing making decisions.
If you want to see how this plays out across an actual multi-brand content and automation operation, [xavierfok.com](/) is where I write it up as it happens.
Get new guides and videos first — join the Telegram channel.