XavierFok
← all posts

Video SOPs rot faster than written ones

2026-08-10 · by Xavier Fok

The re-record cycle nobody budgets for

I run automation across a handful of tools that update on their own schedule, not mine. A rendering pipeline, a couple of admin panels, an MCP client. Every one of them ships UI changes without asking me first. When I had video walkthroughs of how to operate any of these, I was re-recording every few weeks. Not because the process changed. Because a button moved, a menu got renamed, a dashboard got a new layout. The steps were identical. The video was wrong.

A written standard operating procedure survives that same change with one sentence edited. That gap in maintenance cost is the whole argument for this post, and it is not a small one once you are running more than two or three processes that other people, or other AI agents, need to follow.

Why a video breaks and a document does not

A video is a recording of a specific state of a specific interface at a specific moment. Every pixel in that recording is now evidence of how things looked, not just what to do. When the interface changes, the video is factually inaccurate the second that update ships, even if the underlying logic of the task has not moved at all.

A written SOP, by contrast, is an abstraction over the interface. "Click the export button in the top right" survives even when the button moves to the top left, as long as it still says "export." "Open the render queue and confirm the job status shows complete" survives a full visual redesign of that queue, because the instruction refers to state and intent, not to coordinates on a screen.

This is the same reason code that hardcodes pixel positions breaks constantly while code that queries by label or by name is comparatively durable. Text-based instructions bind to concepts. Video binds to appearance.

Where this actually bites in an automation setup

I feel this most sharply with tools that sit between me and an AI agent. When I wire an agent into a real system through something like MCP, I am not just documenting steps for a future version of myself. I am often documenting steps that an agent needs to interpret and execute, or that I need to hand to someone else running the same pipeline on a different machine.

A video of "here is how you configure the connector" is useless to an agent. It cannot watch a video and extract the sequence of actions. It needs the written version: the exact field names, the exact order of operations, the exact conditions that mean success versus failure. Every time I have tried to skip writing the SOP because "I already have a video walkthrough," I have ended up transcribing that video into text later anyway, usually under time pressure, usually because something broke and I needed the steps in a format I could paste somewhere or feed into a process that needed to reason over it.

So the video did not save time. It deferred the writing and added a re-watching tax on top.

What video is actually good for

None of this means video has no place. Video is good at showing judgment calls that are hard to put into words: what a "healthy" render queue looks like at a glance versus one that is quietly backing up, what a UI actually feels like to navigate the first time, the pacing of a multi-step task where timing matters. It is good for onboarding, for building intuition, for showing someone the shape of a system before they touch it.

What video is bad at is being the reference someone returns to six weeks later to execute a specific step correctly. By then the video is describing a system that no longer exists in that exact form, and the viewer has to do mental translation work between what they see on screen and what they see in front of them. That translation step is where mistakes get introduced, especially under any kind of time pressure.

The split that actually holds up

The pattern I have settled on, after too many wasted re-recordings, is to treat video and text as different tools for different jobs rather than two formats of the same document.

The written SOP is the operational source of truth. It lists the steps in plain language, references UI elements by their function rather than their position, and states the expected outcome after each step so a person or an agent can verify they are on track. It gets updated the moment a step changes, because the edit is small: swap one sentence, not re-shoot and re-edit a video.

The video, when I make one at all, is a companion piece. It shows the process once, at a point in time, mainly to build confidence and context for someone new to the task. I do not treat it as something that needs to stay accurate forever, and I say so explicitly if I publish it, because pretending a screen recording will age well is how you end up with a library of confidently wrong tutorials.

Writing SOPs that agents can actually use

This matters more once agents are involved, because an agent following an SOP does not have the human ability to notice "oh, this screenshot is outdated, the button probably moved but I know what they meant." An agent executing steps through a tool call either has the right label, the right endpoint, the right field name, or it fails. Vague or stale instructions do not degrade gracefully with automation the way they sometimes do with a patient human.

Practically, that pushes SOP writing toward being explicit about state rather than appearance. Instead of "you'll see a screen like this," write "the response will include a status field, and success means that field equals complete." Instead of describing where a setting lives visually, describe what the setting is called and what changing it does. That style of writing is more work up front, and it is also exactly what holds up when the interface underneath it changes, or when the "user" reading the SOP is an agent rather than a person.

I want to be clear this is not a claim that AI agents remove the need for someone to write and maintain these procedures, or that they replace the judgment of the person running the pipeline. They do not decide what the right process is. They execute steps that a person defined, and they are only as reliable as the instructions they are given. Writing a solid SOP is still work a person has to do, and rewriting it when a real process changes is still work a person has to do. The only thing that changes is how much of that maintenance burden a small interface tweak costs you, and text costs a lot less than video.

A rough rule for deciding which to make

If I am documenting something I expect to change in the next few months, meaning almost anything touching a third-party UI, a dashboard, or a tool that ships regular updates, I write it. If I am documenting something meant to build intuition once, like showing a new operator the general feel of a system before they start, I might also record it, knowing it has a short shelf life and treating it accordingly.

What I try never to do anymore is treat a screen recording as the durable reference. It is not, no matter how thorough it is at the moment I make it. The interface underneath it is not mine to keep static, and neither is yours.

If you want more on how I structure automation and agent workflows in practice, you can find the rest of what I write about at [xavierfok.com](/).

Get new guides and videos first — join the Telegram channel.