XavierFok
← all posts

I built an AI research agent to do my grunt work

2026-08-15 · by Xavier Fok

# I built an AI research agent to do my grunt work

Early on, this agent handed me a brief that looked perfect. Clean structure, confident tone, every field filled in. It was built on a source about the wrong subject. The agent had searched, found a page that matched the words of my query while missing the meaning, read that page carefully, and extracted facts that were accurate about something I never asked about. Nothing in the output gave it away. The brief read beautifully.

I open with that failure because it shaped the whole build. The fix was never a smarter model. The fix was making every claim in the output checkable, and making myself actually check the ones that matter. Everything else in the design follows from that.

The job I handed over

The work itself used to be the most boring part of my week. Take a name or a topic. Open a dozen tabs. Read pages of material. Copy out the handful of facts that matter. Shape them into a short brief I can act on. Two hours of that produces roughly one paragraph of useful output. The work is repetitive and light on judgment, which makes it close to the ideal shape for handing to an agent.

So the agent runs exactly that loop. I give it a target, either a topic I want background on or a name I want a quick profile of. It searches, pulls back a set of sources, reads them, extracts the specific facts I care about, and returns a brief in a fixed format. Same fields, same order, every time, so I can scan it in seconds. A couple of minutes of run time replaces a couple of hours of mine.

Four parts, none of them clever

The build has four pieces.

First, a search tool. The agent has one way to find things, and that tool is scoped to searching. It does not get a general-purpose browser to wander around with.

Second, a fetch step. For sources worth reading, the agent pulls the full page text so it works from real content.

Third, an extract step. This is where the model earns its keep, reading that text and pulling out the fields I asked for.

Fourth, a fixed output shape. The agent fills a defined structure rather than writing freeform prose.

That is the entire architecture. The decisions that matter live inside two of those pieces, so let me take them in turn.

Read the page, skip the snippet

The lazy version of a research agent works from search snippets, those two-line previews under each result. It is fast and cheap because nothing gets fetched. It also produces shallow briefs that miss the point, because a snippet can make a page look relevant when the full text is about something else, and a snippet never contains enough material to extract real facts from.

My agent fetches the actual page for anything worth reading. That costs more, since there is more text for the model to process, and it takes longer. The quality gap justifies it many times over. A brief assembled from full source text is something I can rely on. A brief assembled from snippets is a guess wearing a suit. If I had to name the single decision that separates this agent from a toy, it is this one, and most people skip it because the snippet version ships in an afternoon.

Structure as a checking tool

The fixed output shape does more work than it appears to.

When an agent writes freeform prose, its mistakes hide inside fluency. A smooth paragraph can carry a wrong fact, and the smoothness gives you no warning. Force the same content into defined fields and the problems surface. A date field holding a vague phrase is a flag. An empty field is a flag. A value outside the plausible range is a flag. Checking stops being careful reading between the lines and becomes scanning a form for entries that look off, which is a far easier job for me and a tractable job for validation code.

The structure also carries sourcing. Every fact in the brief arrives with where it was found. This is what made that early failure catchable. A brief built on the wrong page reads exactly like a brief built on the right one, so the prose itself will never warn you. Glancing at the source takes seconds and reveals the mismatch immediately. Once the sourcing became visible, that whole class of failure went from invisible to routine to catch.

"Not found" is a good answer

There is a limit I had to design around. The agent is only as good as what it can reach. When the information genuinely is not out there in a form it can access, no amount of searching produces it. The dangerous part is what a language model does under pressure to fill a field. It will sometimes produce a plausible value rather than admit it came up empty, and an invented answer looks identical to a found one.

So the agent is explicitly allowed to return "not found". I treat that as a perfectly good result, better than a confident guess by a wide margin. When I review a brief, facts with no clear source behind them get extra suspicion, because those are the ones most likely to have been conjured. Making honesty acceptable, and treating completeness as optional, removed a whole class of quiet fabrication before it could reach me.

Where my ten minutes go

The old manual version cost about two hours per target. The agent version costs a couple of minutes of run time plus a few minutes of my attention, call it ten minutes total. Those ten minutes are the most important part of the process, and I consider them mandatory.

I split my checking effort deliberately. Facts that would drive a real decision get verified against their sources. Facts that barely matter ride through unchecked, because the cost of a small error there is tiny. Re-researching everything would defeat the purpose of the agent. Checking nothing would hand my decisions to a machine that is occasionally wrong with full confidence. So the attention goes exactly where being wrong is expensive, and the agent's speed carries the rest.

What I gained is time and a consistency I did not expect to value as much as I do. The briefs arrive in the same shape every run, which makes them faster to use than the slightly different formats I produced by hand depending on my mood. What I kept is the thinking. The agent took over the tab-opening and the copying. The deciding stayed with me, where it belongs.

The rule that holds it together

One rule makes this system safe. The agent gathers, I decide. Nothing downstream acts on the agent's raw output automatically. The brief is a fast first pass. I read it, verify what matters, and make the call myself.

Blur that line and you inherit the failure I opened with, at full speed and without the glance that catches it. An agent that gathers and also decides and also acts will eventually execute on a confident mistake, and you find out after the damage. I would keep this division of labor even if the model got twice as good, because it is the correct split. The slow gathering goes to the machine, where it saves real hours. The judgment stays with the person, where being right actually matters.

That points at what I think agents are genuinely for right now: acceleration rather than autonomy. The useful agent is the one that does the slow gathering so that when you sit down, the grunt work is already finished and your job is the deciding. A machine that drafts plus a human who decides beats either one working alone, and that pattern keeps showing up in everything I build.

If you want to build one of these for your own repetitive research, here is the order I would follow. Define the output fields first and force the agent to fill them. Give it the minimum tools to search and to read full sources. Require a source on every fact. Then build your own habit of checking whatever would drive a real decision. That gets you an assistant that saves real hours every week. Skip the sourcing and the checking and you get a confident machine that will eventually hand you a beautiful brief about the wrong thing.

More breakdowns like this are on the [home page](/).

Get new guides and videos first — join the Telegram channel.