XavierFok
← all posts

Who Maintains the Code an AI Wrote?

2026-08-10 · by Xavier Fok

The question nobody asks until it breaks

The first time an AI generated script fails at 3am, you find out fast who actually maintains the code. It's not the model. It's not the chat window you closed three weeks ago. It's whoever is on call, staring at a stack trace, trying to reconstruct the intent behind a function nobody on the team remembers writing, because technically nobody did.

I run AI agents daily across content automation, local rendering pipelines, and a handful of MCP-connected tools that touch real infrastructure. The code generation part is the easy part. It's fast, it's often correct on the first pass, and it saves real hours. Maintaining what comes out of it is a different job entirely, and it's one that doesn't go away just because a model wrote the first draft.

AI writes fast, but it doesn't own anything

An AI model has no stake in whether the code it wrote still works in six months. It doesn't get paged. It doesn't remember the edge case it handled last Tuesday unless that context is fed back in. Once the session ends, the reasoning behind a design choice is gone unless a human wrote it down somewhere durable, like a commit message, a comment, or a doc.

This isn't a criticism of the tools. It's just how they work. A coding agent operates inside a session with a context window. When that session closes, its understanding of your codebase closes with it. The next session starts cold, rebuilding context from whatever files, memory systems, or prompts you give it. That's fine for generating new code. It's a real problem if you're relying on the agent to remember why a piece of logic exists.

Ownership doesn't transfer to a tool that has no continuity. It stays with whoever is responsible for the system running in production, which in most setups is still a person.

What "maintenance" actually means here

Maintenance isn't just fixing bugs. It's four separate obligations, and AI generated code doesn't reduce any of them:

Understanding. Someone has to be able to read the code and explain what it does and why, not just that it currently passes tests. If nobody on the team can explain a piece of logic without re-running the prompt that created it, that's a liability, not a convenience.

Review. Generated code still needs a human pass before it ships, the same as code written by a junior engineer would. AI models make plausible-looking mistakes: an off-by-one in a loop, a race condition in async code, a permissions check that looks right but checks the wrong field. These don't always throw errors. They just quietly do the wrong thing.

Testing. Tests don't write themselves just because the code did. If you let an agent generate both the implementation and the tests in the same pass, you can end up with tests that validate the implementation's assumptions rather than the actual requirement. Someone still has to check that the tests are testing the right thing.

Updating. Dependencies change. APIs get deprecated. Business logic shifts. Code that was correct the day it was generated can become wrong six months later, and the AI that wrote it has no idea that happened unless you ask it again.

None of this is unique to AI generated code. It's the same maintenance burden that exists for any code, written by anyone. The difference is that AI generated code gets produced faster and in higher volume, so the maintenance backlog can grow faster than a team expects if nobody is watching for it.

Where this shows up when you run agents daily

In my own pipelines, the failure mode is rarely the model writing broken syntax. It's usually one of these:

A script worked when it was generated, but a dependency shifted its output format and nobody updated the parsing logic that assumed the old shape. The agent that wrote the original parser isn't around to notice the drift. A human has to notice, because the failure shows up as a silent wrong result, not a crash.

A piece of automation was generated to solve a narrow problem, then got reused somewhere it wasn't designed for. The original context, what it was supposed to handle and what it wasn't, lived in a prompt that's long gone. Whoever picks it up six months later has to reverse engineer the intent from the code alone.

Two agent sessions touch the same file at different times with different assumptions about the current state, and the result is a merge that looks fine but breaks something subtle. This is closer to a normal concurrency problem than an AI problem, but it happens more often when generation is fast and cheap enough that people skip the coordination step they'd normally do by habit.

In every one of these cases, the fix wasn't "ask the AI to fix it again." It was a person reading the code, understanding the actual system state, and making a judgment call. The AI can help with that judgment call if you feed it the right context, but it can't make it unsupervised, because it doesn't know what it doesn't know about your production environment.

The review step you can't skip

The teams that get burned by AI generated code are usually the ones that treat code review as optional once the AI is "good enough." It never is, for the same reason a human's first draft isn't good enough to ship without review: correctness at the moment of writing doesn't guarantee correctness in context.

The review has to answer questions the model can't answer about itself: does this fit our actual data shapes, does this handle our actual failure modes, does this match how the rest of the codebase does things, and will the next person who reads this understand it without the original prompt. A model can get you a strong first draft of the code. It can't tell you whether that draft is the right code for your system, because it doesn't have full visibility into your system's history, your team's conventions, or the incidents that shaped the tradeoffs you've already made.

This is also where you catch the quiet stuff: a generated script that works but is quietly slow, a function that duplicates logic that already exists elsewhere in the codebase, an error handler that swallows an exception instead of surfacing it. None of these will show up as failures right away. They show up as maintenance debt six months out.

MCP and the illusion of autonomy

Model Context Protocol and similar agent-tooling setups make it easy to let an AI take actions directly, editing files, running commands, calling external services. That's genuinely useful, and it's most of what makes agent workflows worth running in production instead of just chatting with a model in a browser tab.

But giving an agent the ability to act doesn't give it accountability. If an agent modifies a file, sends a message, or triggers a deploy, the responsibility for that action still sits with the person who set up the permissions and the person who reviewed the change, not with the agent. This matters more as agents get access to more systems, not less, because the blast radius of an unreviewed action grows with the scope of what the agent can touch.

The practical implication is that autonomy and maintenance responsibility move in opposite directions. The more autonomously you let an agent operate, the more deliberate your review and rollback process needs to be, because you're compressing the window where a human would normally catch a mistake before it ships.

A practical rule I use

Every piece of AI generated code that goes into a production system gets a named human owner, the same as if a contractor had written it. That person is responsible for understanding it, reviewing changes to it, and knowing when it needs to be revisited. If nobody is willing to take that ownership, the code doesn't ship, regardless of how clean it looks.

This isn't about distrust of the tools. It's about the fact that maintenance is a human process attached to a system that runs for years, and AI generation is a fast way to produce a first draft of that system, not a substitute for the people who have to live with it afterward.

If you want to see how this plays out in an actual automation setup, from local rendering to agent tooling, [take a look around the site](/).

Get new guides and videos first — join the Telegram channel.