I built a scraping system that runs itself, and the scraper is the smallest part
# I built a scraping system that runs itself, and the scraper is the smallest part
Nearly every scraper you see in a tutorial works exactly once. It runs on the presenter's machine, pulls the data, the table looks clean, everyone nods. Then you build your own version, it works for a day, and on the first quiet morning when nobody is watching it falls over. The site renamed a class. The page started loading its content with javascript. A rate limiter kicked in and handed your parser a block page instead of data.
I have built a fair number of these systems now, the kind that feed a real pipeline every day, and the gap between a scraper that works once and one that runs itself has very little to do with the scraping code. This is a walk through where the actual work lives.
The shape of the system
The whole thing is a loop on a schedule. It fetches a set of pages, extracts structured data, stores results in a database, checks each item against everything already seen, and acts only on the new ones. When something new and interesting appears, it messages me. Fetch, parse, store, dedupe, notify.
Worth noticing immediately: AI appears nowhere in that core loop. A model earns a place only at the edges, on messy pages where plain parsing collapses, and I will show exactly where that boundary sits.
Fetching is the easy fifteen percent
Plain HTTP requests handle most of my sources. A normal library asks for a URL, gets raw HTML back, and for simple sites the data sits right there in the markup for a parser to walk.
Trouble starts when a site declines to put data in the HTML at all. Many modern pages ship a nearly empty shell and load the real content with javascript afterwards, so a plain request sees the shell and nothing else. Two ways forward from there. You can find the underlying data request the page itself makes, usually a tidy JSON endpoint hiding behind the scenes. Or you can drive a real browser in the background, which is slower, heavier, and breaks more often. I reach for the hidden endpoint first every time. Driving a full browser to scrape one number is renting a truck to carry a letter.
The endpoint trick deserves a concrete walkthrough, because it is the biggest time saver I know and beginners almost never reach for it. Open the page in your browser with the developer tools network tab recording. When the listings you want appear a second after load, clearly fetched separately, look through the captured requests. One of them will usually return clean structured data, already in the exact shape the page uses to draw itself. Call that endpoint directly and the parsing problem largely disappears. It is faster, it sidesteps javascript rendering entirely, and it survives redesigns better, since data endpoints change far less often than visual layouts. I check for this on every new source before writing a single line of parsing.
Parsers must suspect themselves
Parsing is where homemade scrapers rot. A selector that grabs the price or the title works perfectly against today's page, then a redesign ships three weeks later, the structure shifts, and the selector grabs nothing or grabs the wrong thing entirely.
The dangerous version is the quiet one. A scraper that crashes loudly is annoying and honest. A scraper that keeps running while storing garbage is a real problem, because by the time anyone notices, the database is full of nonsense with no marker showing where the rot began.
So my parsing layer is built to be suspicious of itself. A field it expected but cannot find never becomes a silent blank. It becomes a flag and a message to me. I would rather field one false alarm a week than excavate a month of corrupted rows after the fact.
The database exists for one question
Early on I dumped results into text files and spreadsheets, and it worked until the hundredth run, when I could no longer tell what I had already seen. An automated scraper runs over and over, and each run mostly finds the same items as the last run plus a handful of new ones. Telling new from old is the entire reason the database exists.
Every item gets a stable identifier built from the fields that make it unique. Before storing anything, the system checks whether that identifier already exists. Seen before, skip. Genuinely new, store it and flag it for attention. That single check converts a noisy scrape that shouts about everything into a quiet system that speaks only when something changed.
Getting blocked is usually self inflicted
Scrape anything at real volume and you will meet the defenses. Sites do not love being scraped, and hammering one with a hundred requests a second from a single address earns you a rate limit or a block, deservedly.
My honest answer here is politeness. Space requests out with delays. Match your request frequency to how often the data actually changes, because scraping an hourly schedule against a site that updates daily is rude and gains nothing. I set a normal user agent so requests look like a regular browser rather than an anonymous script, and when a site says it does not want to be crawled, I respect that. Gray areas exist in scraping and pretending otherwise would be dishonest, but most of the blocking people complain about traces straight back to greedy request volume.
Retries absorb the network's bad moments
The network is imperfect in ways that have nothing to do with you. A momentary blip, a briefly overloaded server, a dropped connection. Without retries, every tiny hiccup registers as a failure and leaves the run incomplete.
So a failed fetch waits a short moment and tries again, two or three times, with a growing gap between attempts. Most transient failures clear on the second try. The ones that persist are probably real, a page that genuinely moved or a site that is genuinely down, and those get logged clearly instead of retried forever. A broken scrape pounding a dead site five hundred times helps nobody.
A scraper you have to remember is a chore
Mine runs with no input from me. On this Windows machine that means a registered scheduled task; on a Linux server it would be a cron entry, and the tool matters not at all. What matters is that the job fires on time whether I am awake, asleep, or away for a week, and that if it stops running I learn about it from an alert rather than from noticing stale data. Once a scrape is scheduled and reporting on itself, it leaves my head completely.
Where the AI actually sits
Now the promised boundary. Fetching, scheduling, deduplication, and retries need no language model, and inserting one there would only add cost and latency for nothing.
The model earns its place on messy unstructured sources, the ones that offer no clean table, just a paragraph of free text with the three facts I want buried inside prose. Writing brittle rules to extract structured fields from loose human writing is miserable work that breaks constantly. Handing the paragraph to a model and asking for those specific fields as clean structured data is exactly what models are good at. On those sources, the scraper fetches raw text the normal way and a small AI step turns mess into structure. That is the model's entire footprint in the pipeline.
The step carries real tradeoffs. It is slower than a local parser, since it is a network call. It costs money on a metered API, a fraction of a cent per page that adds up across thousands of pages. And it can misread a field or hallucinate a value that never appeared on the page, so nothing it returns is trusted blindly. Extracted prices must be numbers in a sensible range. Extracted dates must parse as real dates. Anything failing those checks gets flagged for my eyes instead of stored as truth. The model does the hard reading; boring validation code keeps it honest.
Notifications people actually trust
The notify step turns a database filler into something that serves me. New rows nobody looks at accomplish nothing. When a run finds something genuinely new and worth knowing, it sends the few new items, alone, to a channel I actually watch, the kind of alert that lands on my phone.
Genuinely new is the load bearing phrase, and it is why deduplication matters so much. Without dedupe, every run would announce the same hundred items and I would mute the channel within a day. With it, a buzz means something changed. A notification you can trust to fire only on real news beats ten dashboards you never open.
Maintenance is the cost tutorials omit
A scraper is never build and forget. Sites change, parsers break, and someone fixes them. Across all my sources something breaks every few weeks, rarely a big fix, usually a selector to update or a moved endpoint, and I go in expecting it. Suspicious parsing is what makes this bearable: because missing fields raise flags instead of storing blanks, a broken source surfaces within a run or two rather than a month later. Good design turns maintenance into a quick known chore instead of archaeology through corrupted data.
The part that matters least is the part that feels like the project
The fetch and parse code is maybe a day of work. The system around it, the database tracking what is new, the schedule firing reliably, the retries surviving blips, the validation catching bad data, the alerts announcing breakage, holds all the real engineering. I have rebuilt the scraping core of these systems many times as sites changed. The surrounding system has barely needed touching, because it was built to survive the scraper breaking.
If I could hand one piece of advice to someone starting a first real scraping project: assume the page will change without warning, and build everything around that assumption. Make the parser suspicious, make storage check for what is new, make runs report on themselves, and keep a human in the loop for anything that looks off. A scraper built that way bends when a site changes. One built on the assumption that nothing ever changes is a time bomb with a database attached.
More real systems explained with their tradeoffs intact are on [the home page](/).
Get new guides and videos first — join the Telegram channel.