OpenAI's documentation for GPTBot, Anthropic's support article for ClaudeBot, and Perplexity's documentation for PerplexityBot all explain what the bot does and how to block it. None of the three says a single word about rel=canonical. For a signal that has shaped how Google indexes duplicate content for more than a decade, that's a conspicuous silence, and it means the canonical tag sitting in a page's head might be doing nothing at all for the engines now deciding whether ChatGPT, Claude, or Perplexity ever mentions a brand.
What rel=canonical actually tells Google
A canonical tag tells Google which URL, among several near-identical pages, should represent that content in search results. Google's own documentation on canonicalization describes the process as choosing the page that, based on signals collected during indexing, is objectively the most complete and useful for search users, weighing HTTPS status, sitemap inclusion, redirects, and the rel=canonical annotation itself.
Google is explicit that a stated canonical is a hint, not a rule. The company can and does override it when its own signals disagree with the webmaster's preference. That caveat matters here too: if Google treats its own best-documented duplicate-content signal as negotiable, there was never a strong reason to assume a large language model treats an undocumented one as binding.
What GPTBot, ClaudeBot, and PerplexityBot's documentation says about it
Strip away the branding and the three leading AI crawlers' documentation covers the same narrow ground: what the bot is for, and how robots.txt stops it. Duplicate content, near-identical URLs, and canonicalization don't come up in any of them.
GPTBot's developer documentation states that OpenAI uses the crawler "to crawl content that may be used in training our generative AI foundation models," and that disallowing it in robots.txt removes a site from that pipeline. Anthropic's support article for ClaudeBot, Claude-SearchBot, and Claude-User covers the same ground for its three bots: what each one does, and which robots.txt directive stops it. PerplexityBot's documentation is thinner still, focused almost entirely on allow and block rules. Not one of the three documents what happens once a crawler meets five URLs serving the same paragraph.
A canonical tag is a request Google is willing to negotiate with. It's a request the AI crawlers reading a site have never agreed to receive.
How AI retrieval actually decides between near-duplicate pages
Instead of reading a canonical tag, the retrieval systems behind ChatGPT, Claude, and Perplexity break a page into passages through chunking, generate an embedding for every passage, and compare those embeddings to the query. When several pages on one domain produce near-identical embeddings, the model treats them as a single entity and keeps only one.
Bing's Fabrice Canel and Krishna Madhavan confirmed exactly this behavior in a December 2025 post on the Bing Webmaster Blog: near-duplicate URLs get clustered, and the model "may select a version that is outdated or not the one you intended." VizibleAI covered what that clustering costs a brand in its piece on content cannibalization and AI citations, and the same mechanic applies whenever duplicate pages differ only by a canonical tag rather than by anything the retrieval system can actually read.
What predicts citation instead, according to Ahrefs' 1.4-million-prompt study
Ahrefs analyzed 1.4 million ChatGPT prompts from February 2025 and found the model cites only about half the URLs it retrieves per query, selecting based on how closely a page's title and URL match the question being asked, not on which domain or page a webmaster marked canonical.
Cited URLs matched a query's underlying sub-questions at a cosine similarity of 0.656, against 0.484 for URLs that were retrieved but not cited. Pages with natural-language URL slugs saw an 89.78% citation rate, against 81.11% for pages without them. VizibleAI's own breakdown of query fan-out explains why that sub-question matching step carries so much weight: most of the dozens of fan-out queries one search spins off don't affect citation at all, but the few that do are won on phrasing, not on canonical signaling. Ahrefs' analysis never needed to measure canonical tags, and that absence is itself informative: a study built on 1.4 million real prompts had every reason to flag canonicalization as a factor if it had turned out to be one.
Should a brand still add canonical tags for AI crawlers?
Yes, but as table stakes for classic SEO rather than as a lever over AI citations specifically. A canonical tag still consolidates a domain's Google index, still costs nothing to maintain, and still matters for any tool built on Google's data. It just won't make GPTBot, ClaudeBot, or PerplexityBot collapse duplicate pages the way it nudges Google.
Two adjustments matter more than the tag itself. If the canonical annotation is injected by client-side JavaScript rather than present in the initial HTML, assume several crawlers never see it at all, since none of the three bots' documentation confirms they render client-side code before reading a page. And wherever a duplicate page has no reason to stay live, use an actual 301 redirect instead of relying on the tag alone, since a redirect removes the ambiguity a canonical only hints at.
Pair whichever page is kept with structured data that names the entity explicitly. The access question is separate from the citation question; VizibleAI's rundown of what every major AI crawler actually obeys in robots.txt covers that side of the same consolidation. A canonical tag with no retrieval-level signal behind it is just a note left for an engine that was never confirmed to be reading notes.
Frequently Asked Questions
Do AI crawlers like GPTBot, ClaudeBot, or PerplexityBot respect canonical tags?
Their own documentation never mentions canonical tags. GPTBot's and ClaudeBot's docs describe crawling purpose and robots.txt compliance; PerplexityBot's docs focus on allow and block rules. None describes using rel=canonical to resolve duplicate pages, and independent research into what predicts AI citations, including Ahrefs' 1.4-million-prompt study, doesn't find canonical signals among the factors that matter.
How do AI engines choose between duplicate or near-identical pages if not by canonical tag?
They chunk each page into passages, generate an embedding for every chunk, and compare those embeddings to the query. When multiple pages on one domain produce near-identical embeddings, the model clusters them and keeps a single representative, a pattern Bing's webmaster team confirmed publicly in December 2025.
Should canonical tags be removed since AI crawlers ignore them?
No. Canonical tags still consolidate a site in Google's index and cost nothing to maintain, so keep using them for classic SEO. Just don't expect the tag to control which duplicate an AI engine cites; use 301 redirects for true duplicates and structured data to reinforce which page is authoritative instead.
Does Google treat canonical tags the same way AI crawlers do?
No. Google documents canonicalization as a weighted decision using signals including HTTPS status, sitemaps, redirects, and a stated rel=canonical, calling that preference "a hint, not a rule." AI crawlers' own documentation doesn't describe an equivalent process at all.
What actually determines which page an AI engine cites when duplicates exist?
Ahrefs' analysis of 1.4 million ChatGPT prompts found citation correlates with how closely a page's title and URL match a query's underlying sub-questions, plus naturally phrased URL slugs and relative freshness, not with canonical signaling or domain-level authority alone.



