GPTBot, ClaudeBot, and PerplexityBot have never published a single sentence confirming they read the meta robots noindex tag. OpenAI's own documentation for its crawlers, Anthropic's documentation for Claude's three bots, and Perplexity's crawler reference all describe one control mechanism only: robots.txt. Thousands of sites still add noindex to pages they want kept out of AI answers, assuming it works the same way it does for Google. It doesn't, because noindex was built for a search engine's index, not for an independent AI assistant deciding what to fetch and cite. What actually governs these three crawlers is narrower, more fragile, and worth understanding before you rely on it.
What the noindex tag was actually built to do
Noindex is an instruction a search engine reads after it has already fetched a page, telling that engine not to show the page in its own results. Google's robots meta tag documentation defines it precisely: the directive tells Google to “not show this page, media, or resource in search results.” That only works because Googlebot checks every page it crawls for the tag, and Google's own ranking pipeline is built to honor it when found.
An independent AI assistant is a different pipeline entirely, built by a different company, with no obligation to check for a tag Google defined for its own search results.
None of the three big AI assistants document reading it
OpenAI's crawler documentation names three separate agents: GPTBot, which gathers data for model training; OAI-SearchBot, which indexes pages for ChatGPT Search; and ChatGPT-User, which fetches a single page because a person asked ChatGPT a direct question. The page explains, agent by agent, how to allow or block each one through robots.txt. It says nothing about meta tags, noindex included.
Anthropic runs the same three-way split for Claude: ClaudeBot collects training data, Claude-SearchBot indexes pages for search, and Claude-User fetches a page live, on a user's behalf. Anthropic's support page states only that its bots “respect 'do not crawl' signals by honoring industry standard directives in robots.txt,” and a Search Engine Journal review of that documentation in February 2026 found the added granularity was entirely about which user-agent string a site blocks, not about any HTML tag.
Perplexity draws an identical line between its two agents. Its crawler documentation recommends allowing PerplexityBot, which indexes content, in robots.txt, and separately notes that Perplexity-User, which fetches a page only because a person asked about it, generally ignores robots.txt altogether, since the request behaves more like a person clicking a link than a bot indexing a site. A meta tag appears nowhere in either description. Three vendors, three sets of agents, and the same gap in every one of their published rules.
Robots.txt is what actually governs them, and even that isn't absolute
The one signal all three vendors do document is robots.txt, the plain text file a site publishes at its root that lists which user-agents may fetch which paths. A closer look at what each bot's robots.txt rule actually covers shows how differently those rules apply once an agent is making a live fetch instead of an autonomous crawl.
OpenAI is explicit about the limit of that file for one of its own agents. Its documentation says plainly that ChatGPT-User's fetches aren't autonomous crawling, and for that reason:
Because these actions are initiated by a user, robots.txt rules may not apply.
TollBit's State of the Bots report for the first half of 2026 measured how often that gap shows up in practice: about 15% of identified AI page fetchers reached URLs a site's robots.txt had explicitly disallowed. ChatGPT-User, Bytespider, and Youbot each still reached blocked content on close to half of the European sites that disallowed them, compared with 9% for Claude-User and 13% for Perplexity-User in Europe, rising to 26% for both bots in North America.
A bypass rate is a different problem from a crawler lying about its identity, which is worth checking for separately. It's also a different problem from opting a brand out of AI training, since a Disallow rule written for GPTBot does nothing to stop Claude-SearchBot, PerplexityBot, or any other agent reading the same page for a different purpose. A site that wants real coverage has to write out each agent by name and revisit the list whenever a vendor adds one.
Google's AI Overviews are the one place a meta tag does apply, just not noindex
Google's current robots meta tag documentation, last revised in March 2026, draws a line the three independent assistants never drew. It states that the nosnippet directive “will also prevent the content from being used as a direct input for AI Overviews and AI Mode,” and that max-snippet limits “how much of the content may be used as a direct input” in those same features. Noindex still works the way it always has: removing a page from Google's index removes it from Google AI Overviews by extension, since AI Overviews draws on indexed pages, but Google never frames noindex itself as an AI-specific control. The levers it names specifically for AI Overviews and AI Mode are nosnippet and max-snippet, and both apply only inside Google's own AI features, not to ChatGPT, Claude, or Perplexity.
Google's documentation also covers content that isn't HTML. A PDF, for instance, has no head section to carry a meta tag, so Google reads the equivalent instruction from an X-Robots-Tag HTTP header sent with the file instead. None of GPTBot, ClaudeBot, or PerplexityBot's documentation mentions that header either, which means the same gap that applies to noindex on an HTML page applies to a disallowed PDF or image: robots.txt is still the only lever any of the three vendors name.
What to actually check if you want AI crawlers out
Start with a server log rather than a settings page. A log shows which named agents are actually requesting a page and what response code they got, which is the only way to confirm a robots.txt rule took effect instead of assuming it did. Write Disallow rules against the exact agent string meant to be blocked (GPTBot is not OAI-SearchBot, and ClaudeBot is not Claude-User), and recheck them whenever a vendor splits a bot in two, the way Anthropic and OpenAI both have. A noindex tag left on those same pages is doing nothing for any of the three, and it's still worth removing once you have a robots.txt rule that actually works, since a stray noindex can quietly stop a page from reaching Google and its AI Overviews even after the AI-crawler question is settled.
Cloudflare's bot-management settings now sort crawlers into three categories, Search, Agent, and Training, as of September 2026, and that sorting decides whether a request reaches the server before any robots.txt file or meta tag is read at all.
Frequently Asked Questions
Does adding a noindex tag stop ChatGPT, Claude, or Perplexity from citing a page?
No. OpenAI's, Anthropic's, and Perplexity's own crawler documentation describes robots.txt as the only control signal for GPTBot, ClaudeBot, PerplexityBot, and their related agents. None of the three mentions reading the meta robots noindex tag. Noindex governs whether a page appears in a search engine's own index, like Google's, not whether an independent AI assistant's crawler fetches or cites it.
What actually controls whether GPTBot, ClaudeBot, or PerplexityBot can access a page?
Robots.txt, and specifically the exact user-agent string each vendor publishes. GPTBot, OAI-SearchBot, and ChatGPT-User are three separate agents at OpenAI; ClaudeBot, Claude-SearchBot, and Claude-User are three separate agents at Anthropic. A Disallow rule has to name the specific agent a site wants to block, since blocking one doesn't block the others reading the same page for a different purpose.
Does robots.txt reliably stop these crawlers?
Not entirely. TollBit's State of the Bots report for the first half of 2026 found about 15% of identified AI page fetchers reached URLs their site's robots.txt had disallowed, with ChatGPT-User, Bytespider, and Youbot reaching blocked content on nearly half of the European sites naming them. Robots.txt is a published request a crawler is expected to honor, not a technical barrier it cannot cross.
Why does ChatGPT sometimes still access pages disallowed in robots.txt?
Because ChatGPT-User, the agent that fetches a page when a person asks ChatGPT about it directly, is explicitly documented by OpenAI as behaving differently from GPTBot or OAI-SearchBot. OpenAI states that because those fetches are initiated by a user, robots.txt rules may not apply to them, since the request resembles a person clicking a link rather than an autonomous crawl.
Does Google's AI Overviews respect noindex?
Indirectly. Noindex removes a page from Google's index, and AI Overviews draws on indexed pages, so a noindexed page won't surface there either. But Google's own robots meta tag documentation doesn't frame noindex as an AI Overviews control. The directives it names specifically for AI Overviews and AI Mode are nosnippet and max-snippet, both of which limit how much content those features can use.
What should a site actually do if it wants to keep specific AI crawlers out?
Check server logs first to see which named agents are actually requesting pages, since analytics tools miss crawler traffic entirely. Write robots.txt Disallow rules against the exact agent string, not a wildcard. For a guarantee that doesn't depend on voluntary compliance, block at the server or CDN level, the way Cloudflare's Search, Agent, and Training bot categories do.



