Google's documentation for site owners says there are "no additional requirements to appear in AI Overviews or AI Mode". It adds that you do not need to create "new machine readable files, AI text files, or markup" to be included.1 OpenAI's guidance for ChatGPT search is almost as short: allow OAI-SearchBot to crawl your site, and make sure your host or content delivery network (CDN) accepts traffic from OpenAI's published crawler IP addresses.2 So a site missing from AI answers usually has an access problem, not a content problem. The fault is often in infrastructure that the people responsible for content never see.

This article covers what the operators themselves document about how ChatGPT search, Google AI Overviews, AI Mode, Microsoft Copilot and Claude find and cite web pages. It covers the robots.txt and CDN settings that decide whether their crawlers can reach you, what makes a page easy to cite once they can, and how to measure the result. It reflects documentation available on 26 September 2026, including a change to Cloudflare's default AI crawler settings that took effect on 15 September. It sits within our software coverage, alongside the engineering work of making a site fast and crawlable.

How each AI search system finds pages

The major systems do not share one index. Each finds pages its own way, and each is controlled by a different crawler name.

SystemWhat a page needsControlled byFirst-party reporting
Google AI Overviews and AI ModeIndexed in Google Search and eligible to show a snippet1Googlebot, plus snippet controls such as nosnippetSearch Console Performance report, Web search type, not broken out1
ChatGPT searchCrawlable by OAI-SearchBot from OpenAI's published IP ranges2OAI-SearchBot3None found in OpenAI's documentation
Microsoft Copilot and Bing AI answersIndexed by BingBing's crawler; Bing respects robots.txt4Bing Webmaster Tools AI Performance report4
ClaudeIndexed by Claude-SearchBot, or fetched on request by Claude-User5Claude-SearchBot, Claude-UserNone found in Anthropic's documentation

Google and OpenAI both describe a step that changes how content should be written. Google says AI Overviews and AI Mode may use "query fan-out": "issuing multiple related searches across subtopics and data sources" to build a single response.1 OpenAI says ChatGPT search typically rewrites a question into one or more targeted queries and sends them to search partners. Its example turns a researcher's question about cancer drugs into "CCR8 immunotherapy drug development 2025", then follows up with narrower queries.2 OpenAI names Microsoft among those partners.2

Search crawlers and training crawlers are separate

The most common reason sites disappear from AI answers is an opt-out aimed at AI training that also caught AI search. The operators have split these functions into separate robots.txt tokens:

OperatorSearch and citationModel trainingFetches triggered by a user
OpenAIOAI-SearchBotGPTBotChatGPT-User (robots.txt "may not apply")
GoogleGooglebot (covers AI Overviews and AI Mode)Google-Extended (also covers grounding in Gemini Apps and Vertex AI)–
AnthropicClaude-SearchBotClaudeBotClaude-User

Sources: OpenAI,3 Google,6 Anthropic.5

OpenAI states that "each setting is independent of the others". A site can allow OAI-SearchBot to appear in search results while disallowing GPTBot.3 Sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers", although they can still appear as navigational links.3 Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal".6

A robots.txt that allows AI search and blocks AI training looks like this:

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Allow: /
Disallow: /admin

Sitemap: https://example.com/sitemap.xml

That file has a trap. Under Google's documented interpretation of the Robots Exclusion Protocol, a crawler follows only the most specific group that matches its name. User-agent-specific groups and the * group "are not combined".7 In the example, OAI-SearchBot and Claude-SearchBot match their own groups, so the Disallow: /admin rule under * no longer applies to them. If a path should be off-limits to every crawler, repeat the rule in each named group.

The layer most sites forget: CDN and firewall rules

Google's list of practices for AI features includes "ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure".1 OpenAI asks site owners to confirm their host or CDN allows traffic from its published search crawler IP addresses.2 Both point to the same fact: a robots.txt file that allows a crawler is useless if a firewall rejects the crawler first.

The largest CDN changed its defaults this month. From 15 September 2026, Cloudflare applies new AI bot defaults to new domains. Bots it classifies as Training or Agent are blocked on pages that display ads, and bots classified as Search remain allowed. Crawlers that combine search and training are blocked by every configuration that blocks AI training, including the legacy "Block AI bots" switch.8 Cloudflare defines Agent as "automated activity acting in real time on a person's behalf", including "chat fetch bots".9

Blocks created in Cloudflare's AI Crawl Control are enforced as WAF custom rules on the zone, and return 403 Forbidden or 402 Payment Required.10 Geographic restrictions, ASN blocks and generic bot-fighting rules can stop verified crawlers in the same way, and they do not appear in robots.txt at all.

How to test access properly

A common first check is to fetch your robots.txt with the crawler's user-agent string:

curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
  https://example.com/robots.txt

A 200 here proves only that no rule blocks that user-agent string from your own network. Real crawlers arrive from their operators' networks, which are usually in another country, so rules based on IP address, ASN or country will not fire during your test. A complete check has three more steps:

  1. Read your CDN's security event log for the crawler's user agent or its operator's network, and look at the action taken and the status code returned.
  2. Compare the IP addresses in your logs with the operators' published lists: OpenAI's searchbot.json,3 Google's common-crawlers.json,6 and Anthropic's bots.json.5 Google warns that user-agent strings can be spoofed.6
  3. Use the URL Inspection live test in Search Console, which fetches from Google's own infrastructure and reports failures such as an unreachable robots.txt.

What a failing robots.txt does to Google crawling

Google documents exactly how it reacts when it cannot fetch robots.txt, and the rules are strict.7

  • 5xx errors, timeouts and network failures: for the first 12 hours, Google stops crawling the site while it retries robots.txt. After that it uses the last good copy for up to 30 days. If there is no cached copy, it assumes no restrictions.
  • 4xx errors other than 429: Google treats the site as having no robots.txt, so there are no restrictions.
  • More than five redirect hops: treated as a 404.
  • Caching: robots.txt is generally cached for up to 24 hours, and longer when a refresh fails. Google may adjust the cache lifetime based on Cache-Control: max-age.

AI Overviews and AI Mode draw on Google's index,1 so a robots.txt endpoint that intermittently returns server errors can pause crawling of the whole site. Serve robots.txt as a static, cacheable file that does not depend on application code or a database. Make sure it is reachable from networks outside your own country.

What makes a page easy to cite

Once crawlers can reach a page, the operators' advice is consistent and unexotic. Google's list for AI features is its ordinary SEO list: internal links that make content findable, a good page experience, important content available as text, supporting images and video, structured data that matches the visible text, and up-to-date Business Profile and Merchant Center information.1

Microsoft is more specific about citation. Its guidance, published with the AI Performance report, recommends:4

  • clear headings, tables and FAQ sections, which "make content easier for AI systems to reference accurately"
  • examples, data and cited sources to support claims
  • regular updates, so AI systems cite the current version
  • consistent descriptions of the same entities, products or concepts across text, images and video

Render content on the server. Google says important content should be available "in textual form".1 OpenAI's crawler documentation does not say whether OAI-SearchBot executes JavaScript.3 Content that appears only after client-side rendering is therefore a risk for at least one major AI search system. Server-side rendering or static generation removes that risk.

Signal freshness. Microsoft recommends IndexNow, a protocol for telling participating search engines when a URL is added, updated or deleted.4 Submitted URLs are shared with all other participating engines.11 An accurate lastmod in your XML sitemap does the same job for crawlers that read sitemaps.

Keep structured data honest. Google says no special schema.org markup is needed for AI features.1 Structured data still helps search engines identify entities such as your organisation, authors and products, but it is not a lever for AI citations.

What does not help, or is unproven

  • llms.txt. Google says AI text files are unnecessary for its AI features.1 None of the crawler documentation we reviewed from OpenAI, Anthropic or Microsoft lists it as an input to search or citation.354 It does little harm, but it will not fix an access problem.
  • Treating training opt-outs as a visibility setting. As shown above, they are separate switches.

Measuring AI search visibility

Google. Clicks and impressions from AI Overviews and AI Mode are counted in the Search Console Performance report, under the Web search type, as part of overall search traffic.1 There is no separate filter, so you cannot isolate AI feature traffic in Search Console. Google says clicks from results pages with AI Overviews are "higher quality", meaning visitors spend more time on the site.1 That is Google's claim about its own product, and we have not seen the data behind it.

Microsoft. Bing Webmaster Tools launched an AI Performance report in public preview on 10 February 2026. It shows total citations in Microsoft Copilot, Bing's AI-generated summaries and "select partner integrations", along with average cited pages per day, citation counts per URL, and a sample of the "grounding queries" the AI used when it retrieved your content.4 Microsoft added intent, topic, citation share and comparison views in June 2026.12 It is the most detailed first-party citation reporting we found.

OpenAI and Anthropic. We found no publisher-facing reporting in either company's crawler or search documentation. The practical signal is referral traffic in your own analytics, together with crawler activity in your server or CDN logs. Cloudflare's AI Crawl Control shows which AI crawlers request your pages and which ignore your robots.txt.10

A checklist, in order

  1. Make robots.txt dependable. It should return 200 quickly, as a static file, from networks outside your country.
  2. Allow the search crawlers you want to be cited by: Googlebot, Bing's crawler, OAI-SearchBot, Claude-SearchBot. Decide on training crawlers separately.
  3. Repeat shared Disallow rules in every named user-agent group. Named groups do not inherit rules from *.
  4. Audit the CDN and firewall. Check the AI bot policies (on Cloudflare, the defaults from 15 September 2026), plus country, ASN and bot-fighting rules. Confirm that verified crawlers are allowed.
  5. Verify with the operators' tools: the Search Console URL Inspection live test and Bing Webmaster Tools. Compare crawler IP addresses in your logs with the published lists.
  6. Keep pages indexable and snippet-eligible. nosnippet and max-snippet limit what AI Overviews can show.1
  7. Render important content on the server.
  8. Write for extraction: answer first, one question per section, tables for comparisons, dated and sourced figures.
  9. Signal changes with accurate sitemap lastmod values and IndexNow.
  10. Measure monthly in Search Console, Bing's AI Performance report and your own referral data.

Steps 1, 4 and 7 are infrastructure work, not copywriting. If you are commissioning a new site, put them in the brief: our guide to choosing a software development partner covers how to write requirements a supplier can be held to. For an existing stack, this kind of audit is part of Revere Group's performance tuning work.

Limitations and open questions

  • Nobody guarantees placement. OpenAI says placement in ChatGPT search results "is not guaranteed", and Google says indexing and serving are not guaranteed even when every requirement is met.21 Neither publishes how it chooses which sources to cite.
  • Google does not separate AI feature data. Until it does, the effect of AI Overviews on a site's traffic can only be inferred.
  • Crawler classifications change. Cloudflare changed its defaults this month, and the operators revise their crawler lists. Check your own dashboard instead of relying on the state described here.
  • Evidence for many popular "GEO" tactics is thin. We have limited this article to what the operators themselves document. Claims that particular phrasing, schema types or files raise citation rates are, as far as we could find, not supported by primary sources.

Frequently asked questions

Do I need an llms.txt file to appear in ChatGPT or Google AI Overviews?
Google says no: you do not need new machine-readable files, AI text files or special markup to appear in AI Overviews or AI Mode. None of the crawler documentation we reviewed from OpenAI, Anthropic or Microsoft lists llms.txt as an input to search or citation either.
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI's robots.txt settings are independent. GPTBot covers content that may be used to train its models. OAI-SearchBot controls whether your site appears in ChatGPT search answers. You can block GPTBot and still allow OAI-SearchBot.
Does blocking Google-Extended stop my site appearing in AI Overviews?
No. Google states that Google-Extended does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are controlled by Googlebot and by snippet controls such as nosnippet. Google-Extended does control grounding in Gemini Apps and Vertex AI, which is a separate AI answer surface.
Is there special schema markup for AI Overviews or ChatGPT?
Not for Google. Its documentation says no special schema.org structured data is needed for AI Overviews or AI Mode, though any structured data you use should match the visible text on the page. OpenAI's crawler documentation does not mention structured data.
How can I see whether my site appears in AI Overviews or Copilot?
Google includes AI Overviews and AI Mode traffic in the Search Console Performance report under the Web search type, but does not report it separately. Bing Webmaster Tools has an AI Performance report that counts citations of your pages in Microsoft Copilot and Bing's AI-generated answers.
How long does a robots.txt change take to affect ChatGPT and Google?
OpenAI says it can take about 24 hours from a robots.txt update for its search systems to adjust. Google generally caches robots.txt for up to 24 hours, and may keep a cached copy longer if it cannot fetch a fresh one.

Sources

  1. Google Search Central, "AI features and your website", last updated 10 December 2025, accessed 26 September 2026. https://developers.google.com/search/docs/appearance/ai-features ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14

  2. OpenAI Help Center, "Searching the web with ChatGPT", accessed 26 September 2026. https://help.openai.com/en/articles/9237897-chatgpt-search ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  3. OpenAI, "Overview of OpenAI Crawlers", OpenAI API documentation, accessed 26 September 2026. https://developers.openai.com/api/docs/bots ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7

  4. Microsoft Bing, "Introducing AI Performance in Bing Webmaster Tools Public Preview", Bing Webmaster Blog, 10 February 2026. https://blogs.bing.com/webmaster/2026/2/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview/ ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  5. Anthropic, "Does Anthropic crawl data from the web, and how can site owners block the crawler?", Claude Help Center, 7 April 2026. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler ↩ ↩2 ↩3 ↩4

  6. Google, "List of Google's common crawlers", last updated 14 July 2026, accessed 26 September 2026. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers ↩ ↩2 ↩3 ↩4

  7. Google, "How Google interprets the robots.txt specification", last updated 31 August 2026, accessed 26 September 2026. https://developers.google.com/crawling/docs/robots-txt/robots-txt-spec ↩ ↩2

  8. Cloudflare, "Block AI Bots", Cloudflare documentation, last updated 1 July 2026, accessed 26 September 2026. https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/ ↩

  9. Cloudflare, "Bots: AI bots", Cloudflare documentation, last updated 1 July 2026, accessed 26 September 2026. https://developers.cloudflare.com/bots/concepts/bot/#ai-bots ↩

  10. Cloudflare, "Manage AI crawlers", AI Crawl Control documentation, last updated 28 July 2026, accessed 26 September 2026. https://developers.cloudflare.com/ai-crawl-control/features/manage-ai-crawlers/ ↩ ↩2

  11. IndexNow, "Documentation", accessed 26 September 2026. https://www.indexnow.org/documentation ↩

  12. Microsoft Bing, "New AI Visibility Insights in Bing Webmaster Tools: Intents, Topics, Citation Share, Compare", Bing Search Blog, 16 June 2026. https://blogs.bing.com/search/2026/6/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare/ ↩