The messages I never let AI answer
I run AI in production every day. It writes first drafts of content, sorts my inbox, flags what needs attention, and handles a good chunk of routine customer messages before I ever see them. So when people ask if I let AI handle customer support, the honest answer is: mostly, yes. But there's a specific list of messages I never let it touch, and the reasons come from watching what actually goes wrong, not from some general fear of the technology.
This isn't a doom post. AI customer replies work well for a large share of what comes in. The point of this piece is where I draw the line and why.
Where ai customer replies genuinely work
Most support volume is repetitive. "How do I reset my password." "What's included in the plan." "Is this compatible with X." These questions have a stable, correct answer that doesn't change based on the specific person asking. A language model drafting a reply to these is working from a fixed set of facts, and if you've given it the right reference material, it can produce something accurate and fast.
The way this actually works under the hood matters. The model isn't "knowing" the answer the way a support agent who's been doing the job for years knows it. It's predicting the next most likely words given the prompt and whatever context you fed it. When that context is a clean FAQ or a documented policy, the output tends to track the source closely. When the context is thin, ambiguous, or missing, the model still produces a confident-sounding sentence anyway. That gap between "sounds confident" and "is grounded in something true" is the whole story of where this technology helps and where it hurts.
For routine, low-stakes, well-documented questions, that gap rarely matters because there isn't much room to be wrong. For everything else, it matters a lot.
The messages I flag out immediately
Here's the actual list, not a hypothetical one:
Refund and billing disputes. Anything involving money that's already changed hands. A customer saying they were charged twice, or that a renewal happened without consent, or asking for money back. These messages need someone who can pull the real transaction record, understand the account's specific history, and make a judgment call that the business will actually stand behind. An AI-drafted reply can sound reasonable while committing to a refund policy that doesn't match what happened, or promising a timeline nobody can guarantee.
Anything angry. Not mildly annoyed, angry. When someone writes in frustrated, the reply needs to acknowledge the specific thing that went wrong for them, not a generic version of "we're sorry for the inconvenience." A model without the actual incident details will smooth over the tone but miss the substance, and customers can tell the difference immediately. That mismatch tends to make things worse, not better.
Legal-adjacent language. Words like "chargeback," "dispute," "lawyer," "report," or anything that references terms of service violations. These aren't support conversations anymore, they're liability conversations, and the reply has consequences beyond that one thread.
Account security and access issues. Someone locked out, someone claiming unauthorized access, someone asking to change ownership or transfer a paid account. These require verification steps that an AI reply generator has no way to actually perform. Getting this wrong isn't a bad customer experience, it's a security incident.
Anything where I don't have current, accurate state. If a server is down, or a delivery is delayed, or a feature is temporarily broken, the model doesn't know that unless I've told it in that exact moment. Left alone, it will draft a plausible, polite, entirely wrong reply because it has no way to know the ground truth has changed since its instructions were written.
Why "sounds right" isn't the same as "is right"
This is the part that took me longest to internalize, even running this stuff daily. A well-drafted AI reply and a correct AI reply look identical on the page. Both are grammatically clean, on-brand in tone, and structured like a real answer. The only way to tell them apart is to check the underlying facts, and by the time you're checking every fact in every reply, you've removed most of the time savings that made automation worth doing in the first place.
That's the actual tradeoff, and it's not really about AI being untrustworthy. It's about which categories of message have a small blast radius if the answer is subtly wrong, and which categories have a large one. A wrong answer to "does this work on Android" costs a follow-up message. A wrong answer to "you charged me twice" costs trust, and possibly a real dispute.
I treat that blast radius as the actual filter, not some blanket "AI can't be trusted with customers" rule. The categories above are the ones where being wrong is expensive, so those are the ones a person handles.
What I actually do instead
For the flagged categories, AI still does work, just not the final reply. It sorts the message into the right bucket, pulls relevant history if I've wired that context in, and sometimes drafts a rough starting point that I rewrite before sending. The model is doing preparation, not decision-making, on anything with real stakes attached.
For everything else, the reply goes out with a human reviewing it before send, at least for now, on the categories where the account is unusual or the phrasing suggests something the model hasn't seen before. Routine questions with clean, unambiguous answers go out with less oversight, because the cost of being wrong there is genuinely low.
The dividing line isn't "is this a customer message," it's "does this message involve money, anger, legal risk, security, or a fact that changes faster than my documentation does." If yes, a person handles it. If no, the model earns its keep.
The lesson wasn't really about AI
The interesting part of this whole exercise wasn't discovering that AI has limits. Everyone building with this stuff knows it has limits. The interesting part was realizing I needed a written rule for where those limits sit, because without one, the temptation is to let automation creep into categories it shouldn't touch simply because it's already handling everything next to them. Automation expands to fill the space you don't explicitly fence off.
So the fence is the actual system. Not a vague sense of caution, a specific list: money, anger, legal language, account security, and anything where my documentation might be stale. Everything on that list gets a person. Everything off it can run through the pipeline like the rest of the routine traffic.
If you're setting up ai customer replies for your own operation, that's the exercise worth doing before anything else. Not "can the model write a good sentence," it almost always can. The real question is which categories of message you're willing to let it answer unsupervised, and which ones need a human name attached to the send button.
If you want to see how I actually wire this stuff together, from local rendering to agents connected into real tools, you can find more of it on the [xavierfok.com homepage](/).
Get new guides and videos first — join the Telegram channel.