The Un-Obfuscated Crawl: What AI Bots Are Actually Doing to Your Origin Server
Last October, my phone started buzzing at 3:14 AM. A technical documentation portal I manage for a client had completely fallen over. It was running on a modest 4GB RAM instance on Hetzner—nothing fancy, but normally more than enough to handle its daily run rate of 15,000 page views with plenty of caching headroom.
When I SSH'd in, the load average was sitting at 48. Nginx was throwing 502s because PHP-FPM had exhausted every worker pool, swapping out memory until the Linux OOM killer murdered MySQL. My first instinct was a layer-7 DDoS or a script kiddie running sqlmap. I yanked the tail of the raw access log, stripped out static asset requests, and watched the stream fly past.
It wasn't a malicious attack. It was an un-obfuscated, completely unthrottled crawl from two major AI training bots running in parallel: GPTBot and ClaudeBot. In exactly 94 minutes, they had generated 42,810 requests against faceted search queries, filtered archives, and deep pagination links that had never been warmed in the cache.
The Polite Crawler Is a Myth
The documentation published by LLM labs paints a rosy picture. They tell you their crawlers respect robots.txt, identify themselves cleanly via explicit User-Agent strings, and back off when origin servers get warm. That might be the intention in their engineering meetings, but in production, their crawl engines are optimized for speed and ingestion volume, not your monthly infrastructure budget.
Here is what an un-obfuscated crawl looks like when you bypass third-party dashboards and look directly at your server's access files:
First, they hit routes your regular users never touch in sequence. A human reads an article, clicks a related post, maybe scans a category. An AI crawler discovers a URL structure like /docs/search?topic=auth&sort=date&page=12 and proceeds to loop through every permutation of your query parameters. If you have five filter dropdowns, they will hammer thousands of edge cases that force cold queries straight to your database.
Second, their retry logic is ruthless. When my client's server started throwing 503 Service Unavailable, the scrapers didn't back off for an hour. They queued retries within seconds from different IPs in the same ASN subnet. The server never got a breath of fresh air to clear its thread pool.
Third, they consume massive upstream bandwidth on repeat visits. I watched one crawler fetch the same 1.8MB uncompressed JSON schema document 210 times in three hours because each of its distributed worker nodes evaluated the route independently without a shared cache.
Why Dashboards Hide the Damage
A lot of teams think they understand their bot landscape because they check Google Analytics or peek at Cloudflare's summary tab. That is a mistake.
Google Analytics is completely blind here. These crawlers do not execute client-side JavaScript. They pull the raw HTML, parse the text tokens, and leave. If you rely on GA4 to estimate bot load, you are looking at zero percent of the picture.
Edge firewalls and CDN proxies are slightly better, but they introduce their own blind spots. They group traffic into wide, opaque buckets like "Automated Traffic" or hide deep log inspection behind enterprise contracts that cost five figures a year. More importantly, CDN dashboards show you smoothed averages over time. A five-minute spike of 60 requests per second will melt a small origin server, but on a 24-hour aggregate chart, it looks like a harmless little blip.
To see what is actually happening, you have to look at the un-obfuscated raw web server logs: the Nginx, Apache, or Caddy access files stored right there on your disk.
The Data Privacy Trap of Log Analysis
When clients realize they have a crawler problem, their first impulse is usually to sign up for an external SaaS log aggregator. They pipe their access logs out to an external cloud bucket and run queries through a managed dashboard.
I advise against this almost every time.
Your access logs are full of sensitive footprint data. They contain raw visitor IP addresses, referrers from authenticated intranet sessions, and occasionally sensitive parameters that sloppy junior developers left in GET requests. Shipping gigabytes of un-redacted server history to a third-party server just to see if Perplexity is scraping your blog creates an unnecessary compliance risk under GDPR and basic security hygiene.
You do not need an external data pipeline to audit your crawlers. The file is already on your machine. All you need is a local parser that runs right in your browser session, evaluates the User-Agent tokens, matches the IP ranges against verified crawler ASN blocks, and shows you the volume without your logs ever leaving your own device.
If you suspect scrapers are chewing through your compute or scraping your proprietary content without permission, take a look at what we built at GuardLabs. I work as an ai bot traffic analyzer freelance specialist helping teams regain control of their origin infrastructure. You can drop your raw log files directly into our browser-based utility—it processes everything client-side without sending a single byte of your data over the wire: Анализатор логов сайта: какие ИИ-боты и сканеры его посещают. It will give you an exact breakdown of which models are reading your site, how frequently they hit your dynamic routes, and whether it is time to rewrite your access rules.