Your Robots.txt is Leaking Value to LLMs, and Googlebot Won't Save You
I used to believe in the open web. I really did. For over a decade as a developer, my default setup for any new client site was a lazy, optimistic couple of lines: User-agent: * followed by Allow: /. It felt clean. It felt collaborative. It was the digital equivalent of leaving your front porch light on for neighbors.
Then last October happened.
I was managing a high-end technical blog for a boutique engineering client. We noticed their site was sluggish, almost gasping for air under a sudden surge of traffic. When I dug into the raw Nginx logs, I found something that made my blood run cold. Exactly 41% of our total server bandwidth that month was consumed by a relentless horde of undocumented scrapers. They weren't indexing our pages to send us traffic. They were vacuuming up our deeply researched, proprietary case studies to train LLMs.
A week later, I found our exact paragraphs being spat out by a popular chat assistant. No attribution. No link. Zero referral traffic. Just our hard work, repackaged as a zero-click answer.
That was the day the open web died for me. And it’s why your current robots.txt file is probably a liability.
The Transaction Has Changed
For twenty years, the relationship between website creators and search engines was a fair trade. We gave Googlebot and Bingbot access to our content. In return, they put us in their index and sent us human visitors. Everyone won.
But the new breed of AI crawlers doesn't play by those rules. When GPTBot or ClaudeBot crawls your site, they aren't looking to index you for search. They are downloading your intellectual property to train a model that will eventually compete with you. They want to absorb your knowledge so their users never have to visit your website again.
This isn't search engine optimization anymore. This is asset protection.
If you do any kind of ai bots robotstxt freelance development or consulting, you need to realize that treating all bots the same is a recipe for disaster. Your clients are losing their competitive advantage every single day their site remains wide open to LLM scrapers.
How to Draw the Line
You cannot just block everything. If you block Googlebot, your business dies because you disappear from search results. But you also can't just let everyone in. You have to learn how to segment your traffic.
The solution lies in understanding the specific user-agents of the companies doing the scraping. For example, Google now separates its search crawler from its AI training crawler. Googlebot is what you want to allow. Google-Extended is what they use to train Gemini. You can block the latter without hurting your search rankings.
When you are configuring your robots txt GPTBot ClaudeBot directives, you have to be highly specific. Here is what a defensive, modern robots.txt file actually looks like when you want to keep your search traffic but protect your data from being used as training fodder:
User-agent: Googlebot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: PerplexityBot
Disallow: /
This looks simple, but keeping up with this list is a nightmare. New AI startups launch every week. They invent new user-agents. Some of them disguise themselves. Some of them ignore the robots exclusion protocol entirely, though the major players (OpenAI, Anthropic, Google) generally respect it to avoid massive copyright lawsuits.
Why Manual Maintenance is a Trap
I tried managing this manually for my freelance clients. I had a master text file that I would copy-paste across thirty different servers. But within a month, it was outdated. I missed the announcement of a new crawler, or a client accidentally overwrote the file during a WordPress update.
You shouldn't be writing these files by hand anymore. The landscape moves too fast, and the stakes are too high. One missed user-agent can mean your entire database of original content gets scraped overnight and ingested into a model that will render your traffic obsolete.
We built a simple, reliable tool to solve this exact headache for our own projects, and we keep it constantly updated as new AI crawlers emerge.
If you want to stop letting tech giants train their multi-billion-dollar models on your hard work for free, you need to lock down your configuration today. We can help you generate a watertight, up-to-date file in about thirty seconds: Генератор robots.txt для ИИ-краулеров (GPTBot, ClaudeBot и др.) → https://guardlabs.online/care/. It’s a small step, but it’s the only way to take back control of your own data.