Featured on SaaSBison Featured on Toolfio Listed on Bowora Featured on Uneed Featured on ToolPilot

AI crawler traffic tripled after GPT-5. Rate-limiting the wrong bot can quietly cost you citations

AI crawler traffic has roughly tripled since GPT-5 launched, and some sites are throttling or blocking bots outright to cope with the server strain. The problem is that OpenAI's own documentation shows GPTBot, OAI-SearchBot, and ChatGPT-User do very different jobs, so a blanket rate limit can choke off the citations a site was actually trying to protect.

OpenAI's crawler traffic roughly tripled in the eight months after GPT-5 launched in August 2025, according to a Botify analysis of seven billion log events across enterprise sites. OAI-SearchBot, which retrieves pages for ChatGPT's search feature, logged 3.5 times more requests. GPTBot, which gathers training data, logged 2.9 times more. That surge is colliding with a separate trend: site owners under real server strain from AI bots are reaching for rate limits and outright blocks. The two trends don't mix cleanly. Throttle the wrong bot, and a site can quietly lose the citations it was trying to protect.

How much AI crawler traffic has actually grown

Botify's dataset, covering November 2024 through March 2026 across its enterprise client base, found OpenAI's search-to-training crawl ratio flipped after GPT-5: roughly 0.95 before the launch and 1.14 after, meaning search retrieval now generates more requests than training ever did. OAI-SearchBot activity rose fastest in healthcare, up about 740%, and in media and publishing, up about 702%, year over year. Even with that growth, OpenAI's bots remain a fraction of Googlebot's footprint: in the most recent 30-day window Googlebot logged 18.2 billion events against 887 million for OpenAI's combined crawlers, about 4%, up from roughly 1.38% a year earlier, according to Search Engine Journal's analysis of the Botify data. The growth rate matters more than the absolute share: a crawler going from negligible to meaningful in a year is exactly the kind of traffic a server team notices first and a marketing team notices second.

Why server strain is pushing some sites toward outright blocking

Fastly's threat research team found AI crawlers made up nearly 80% of AI bot traffic across 6.5 trillion monthly requests between mid-April and mid-July 2025, with Meta alone generating 52% of that crawling, more than Google and OpenAI combined.

OpenAI's bots told a different story on the real-time side: fetcher requests, the kind a chatbot sends out mid-conversation to check a specific page, came almost entirely from OpenAI, at 98% of fetcher volume, and peaked above 39,000 requests a minute against single sites. That strain has already forced extreme responses elsewhere. A widely read report from TechCrunch described AmazonBot ignoring robots.txt and spoofing IP addresses while repeatedly knocking a Git-hosting server offline in early 2025, which pushed developer Xe Iaso to build Anubis, a proof-of-work gate that blocks bots before they ever reach the server. A Fedora system administrator blocked an entire country, Brazil, after similar bots wouldn't respect standard defenses either. Measuring whether a bot's traffic is worth that strain is what the crawl-to-refer ratio is for: it compares how often a bot crawls a site against how often it sends a human visitor back, and a badly skewed number is often the first sign a bot deserves scrutiny before a block.

You can't secure what you can't see, and without clear verification standards, AI-driven automation risks are becoming a blind spot for digital teams, said Arun Kumar, Fastly's senior security researcher.

Not every OpenAI bot does the same job, and treating them alike is the mistake

GPTBot, OAI-SearchBot, and ChatGPT-User are three separate bots with three separate jobs, and OpenAI's developer documentation treats them as independent switches rather than one setting. Rate-limiting all three identically risks throttling the one bot responsible for citations while leaving the one that only feeds model training untouched.

GPTBot gathers content that may train future models, and disallowing it opts a site out of training without touching search visibility at all, per OpenAI's own guidance. OAI-SearchBot is the one that determines whether a page can show up in ChatGPT's search answers, so blocking it removes a site from citations even if GPTBot is left untouched. ChatGPT-User is different again: it fires only when a person asks ChatGPT to visit a specific page mid-conversation, and OpenAI says it isn't governed by robots.txt the same way the other two bots are, since the request is user-initiated rather than automated crawling. We've covered the training side of this split before: disallowing GPTBot for training doesn't block ChatGPT's live citations, because OAI-SearchBot keeps running regardless of what GPTBot is told. The reverse mistake, rate-limiting all three bots at the server or CDN level because one of them is spiking traffic, has the opposite effect. It can choke off OAI-SearchBot's live retrieval requests along with GPTBot's training crawl, which is the one combination that actually costs citations.

Why crawl-delay and blanket rate limits don't work the way you'd hope

Crawl-delay, the robots.txt directive many site owners reach for first, isn't part of the formal standard that governs robots.txt, and nothing requires an AI crawler to honor it. RFC 9309, the IETF standard published in 2022, intentionally left crawl-delay out because there was no consistent real-world behavior worth codifying.

That leaves robots.txt doing what the standard itself calls advisory guidance, not an access-control mechanism. A well-behaved crawler reads the Disallow lines and complies, and a crawler that ignores the file entirely, which several AI crawlers have been documented doing, pays no penalty for doing so. A hard rate limit enforced at the server or CDN level is more reliable than a polite request in a text file, but it is also blunt: a rule written against a user agent string or an IP block tends to catch every bot wearing that identity, whether it's the training crawler nobody needs urgently or the search bot a page is actively trying to get cited by. We've already seen what happens when a CDN's default settings block AI crawlers without anyone intending it; the same blunt-instrument problem applies to rate limits set in a hurry during a traffic spike. A fuller breakdown of what robots.txt actually stops for each major AI bot is worth reading before writing a single Disallow line.

A practical approach: verify the bot, then throttle by purpose

The fix isn't choosing between unlimited access and a blanket block. It's verifying which bot is actually making the request, then setting limits that match what that bot does rather than what it calls itself.

Start with the logs, not the robots.txt file: analyzing AI crawler traffic in server logs shows which bots are actually hitting a site hardest and when, which is the only way to know whether GPTBot, OAI-SearchBot, or something spoofing either one is behind a spike. User agent strings are trivial to fake, so match requests against the IP ranges vendors publish instead. A growing share of traffic hitting sites under a familiar bot's name turns out to be fake crawlers impersonating the real ones, and rate-limiting a verified search bot while a spoofed one sails through defeats the purpose entirely. Once the traffic is verified, the limits can be asymmetric on purpose: throttle or cache training-only crawlers hard, since GPTBot's output has no deadline, and leave more room for OAI-SearchBot and similar fetch-on-demand bots, whose requests are usually gone within the few seconds it takes to answer the question that triggered them.

Frequently Asked Questions

What's the difference between GPTBot and OAI-SearchBot?

GPTBot crawls content OpenAI may use to train future models, and disallowing it opts a site out of training only. OAI-SearchBot retrieves content so ChatGPT's search feature can find and cite a page in real time. OpenAI's developer documentation treats the two as independent settings, so a site can block GPTBot for training while still allowing OAI-SearchBot for citations, or the reverse.

Does crawl-delay in robots.txt actually slow down AI crawlers?

Not reliably. Crawl-delay was intentionally left out of RFC 9309, the IETF standard that formalized robots.txt in 2022, because there was no consistent behavior worth standardizing. Several AI crawlers have been documented ignoring robots.txt directives altogether, so a crawl-delay line functions as a polite request rather than an enforceable limit.

Will rate-limiting AI bots hurt my ChatGPT or Perplexity visibility?

It depends which bot gets throttled. Limiting a training-only crawler like GPTBot has no effect on whether ChatGPT cites a page, since citations come from OAI-SearchBot and similar real-time retrieval bots. Throttling those retrieval bots indiscriminately, especially during the brief window right after a user asks a question, risks timeouts that keep a page out of the answer.

How do I tell a real AI crawler from a spoofed one before rate-limiting it?

Match the request's IP address against the ranges each vendor publishes, since user agent strings alone are easy to fake. OpenAI and other vendors publish current IP lists, and checking server logs against those ranges, rather than against the user agent header, is the most reliable way to confirm a request actually came from the bot it claims to be.

Why are some sites blocking AI crawlers entirely, even by country?

Server strain, not citation strategy. A 2025 TechCrunch report described aggressive AI crawlers pushing open-source maintainers into drastic measures, including one Fedora system administrator blocking all of Brazil, after standard defenses like robots.txt and user agent blocking failed against bots that spoofed their identity and rotated IP addresses.