AI Visibility Audit: ChatGPT, AI Overviews, Copilot
By Shawn Michael Thomas II, Founder of Simple Easy Transactions · Published ·
What is an AI visibility audit, and how do I check whether ChatGPT, Google AI Overviews and Copilot can read my site? An AI visibility audit checks whether the crawlers behind ChatGPT, Claude, Perplexity, Google AI Overviews and Microsoft Copilot may read a site, and whether each page can be indexed and understood once they do. It reads robots.txt for each crawler (OAI-SearchBot, GPTBot, Claude-SearchBot, PerplexityBot, Google-Extended, Bingbot and others), noindex rules, canonical links, JSON-LD structured data and llms.txt. It shows what blocks a page; no audit can make an engine cite one.
AT A GLANCE
Search and answer crawler vs Training or grounding control.
| Question | Search and answer crawler | Training or grounding control |
|---|---|---|
| OpenAI (ChatGPT) | OAI-SearchBot: surfaces sites in ChatGPT search; opted-out sites are not shown in ChatGPT search answers | GPTBot: crawls content that may be used to train OpenAI's models; set independently of OAI-SearchBot |
| Anthropic (Claude) | Claude-SearchBot indexes for search; Claude-User fetches pages when a person asks Claude | ClaudeBot: collects content that may contribute to model training |
| Perplexity | PerplexityBot: surfaces and links sites in Perplexity search; Perplexity-User fetches for a person | None: Perplexity says PerplexityBot is not used to crawl content for AI foundation models |
| Google (AI Overviews, AI Mode) | Google Search's own crawling: a page must be indexed and eligible for a snippet | Google-Extended: Gemini training and grounding; no effect on Google Search |
| Microsoft (Copilot) | Bing's index: content tagged NOARCHIVE is left out of answers | NOCACHE and NOARCHIVE also limit use in training Microsoft's models |
| Apple (Siri, Spotlight, Safari) | Applebot | Applebot-Extended: opt out of training Apple's generative models |
| Common Crawl | None | CCBot: an open repository of web crawl data, free for anyone to use |
DEFINITIONS
What an AI visibility audit is, and what AIO and GEO mean.
An AI visibility audit, also called an AI search audit or an AI SEO audit, asks a narrow question: can the systems behind AI answers read this page, and does the page say clearly what it is? It covers who may crawl (robots.txt and firewall rules), whether the page may be indexed (noindex and canonical), and how well a machine can understand it (headings, text, JSON-LD structured data, and sometimes llms.txt).
AIO is optimization for AI overviews in search results, such as Google's AI Overviews and AI Mode. GEO, generative engine optimization, is a term from a 2023 research paper, “GEO: Generative Engine Optimization,” which describes it as a way to help content creators improve their visibility in generative engine responses. SEO, AIO and GEO overlap almost entirely at the technical level: Google says there are no additional requirements to appear in AI Overviews or AI Mode beyond being indexed and eligible for a snippet.
What an audit cannot do is decide what an engine shows. Google's own page says that meeting every requirement “doesn't mean that Google will crawl, index, or serve its content.”
WHAT EACH ENGINE READS
What ChatGPT, Claude, Perplexity, Google and Copilot actually read.
OpenAI separates search from training. In OpenAI's words, “OAI-SearchBot is used to surface websites in search results in ChatGPT's search features,” and sites opted out of it “will not be shown in ChatGPT search answers, though can still appear as navigational links.” GPTBot crawls content that may be used to train its models, and OpenAI says each setting is independent: you can allow OAI-SearchBot and disallow GPTBot. ChatGPT-User fetches a page when a person asks, and OpenAI says robots.txt rules may not apply to it.
Anthropic lists three bots. ClaudeBot collects content that may contribute to training; Claude-User fetches pages when a person asks Claude; Claude-SearchBot indexes for search, and Anthropic says disabling it “may reduce your site's visibility and accuracy in user search results.” Anthropic says its bots honor robots.txt and the non-standard Crawl-delay.
Perplexity says PerplexityBot “is designed to surface and link websites in search results on Perplexity” and is not used to crawl content for AI foundation models. Its user-triggered fetcher, Perplexity-User, “generally ignores robots.txt rules” because a person requested the fetch.
Google's AI Overviews and AI Mode draw on Google Search itself. To be a supporting link, a page “must be indexed and eligible to be shown in Google Search with a snippet,” and Google says “You don't need to create new machine readable files, AI text files, or markup to appear in these features.” Google-Extended is a separate control for Gemini training and grounding, and Google says it “does not impact a site's inclusion in Google Search nor is it used as a ranking signal.” To limit what Search's AI features show, Google points to nosnippet, data-nosnippet, max-snippet and noindex.
Microsoft Copilot answers from Bing. Bing's AI Performance report in Bing Webmaster Tools shows how often a site is cited in Microsoft Copilot and Bing's AI-generated summaries. Bing's September 2023 announcement, written when the product was called Bing Chat, says content tagged NOARCHIVE “will not be included in Bing Chat answers,” while NOCACHE limits answers to the URL, title and snippet, and its example addresses Bing's crawler as bingbot. Bing also recommends IndexNow so AI answers use the current version of a page.
Apple says publishers can opt out of training Apple's generative models by disallowing Applebot-Extended, and that content stays discoverable through Spotlight, Siri and Safari. Common Crawl's CCBot builds an open repository of web crawl data that anyone can use, and Common Crawl warns that some crawlers falsely identify themselves as CCBot.
THE CHECKLIST
The AI visibility checklist, in the order I run it.
Robots.txt is the first thing to read, and it's easy to misread. RFC 9309, the Robots Exclusion Protocol, says a crawler obeys the group that names its own token and falls back to the * group only when none matches, and that within a group the most specific (longest) matching rule wins. So a site that allows everything under * but has an old group for GPTBot is following the GPTBot group for GPTBot, whatever the * group says. The RFC also says the rules “are not a form of access authorization”: they are requests, which is why user-triggered fetchers can ignore them.
- robots.txt: for each token (OAI-SearchBot, GPTBot, ChatGPT-User, Claude-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Googlebot, bingbot, Applebot-Extended, CCBot), find the group that applies and the rule that decides the page.
- The edge: a firewall, CDN bot rule or host block can refuse a crawler that robots.txt allows. Google lists allowing crawling “by any CDN or hosting infrastructure” as a fundamental, and Perplexity's docs explain allowing its bots in a WAF.
- noindex: a robots meta tag or X-Robots-Tag header removes the page from Google. Google warns the page must not be blocked in robots.txt, or the crawler never sees the noindex.
- Canonical: the rel="canonical" link should point to the URL you want shown. Google calls it a strong signal, alongside redirects and sitemap inclusion.
- Text, not pictures of text: Microsoft's guidance warns against key information only in images or PDFs, and against hiding important answers where AI systems may not render them.
- JSON-LD: structured data should parse and match the visible page. Google recommends JSON-LD and says there is no special schema.org markup needed for its AI features; Microsoft says schema helps AI systems understand content. Schema.org is the shared vocabulary.
- llms.txt: optional. It is a proposal from llmstxt.org, not a standard any engine has committed to read, and Google says no AI text file is needed for AI Overviews.
- Freshness: a sitemap for every engine, and IndexNow for Bing and the engines that take it.
# Search and answer crawlers: allowed User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot Allow: / # Training controls: opted out User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot Disallow: / # Everyone else, including Googlebot and bingbot User-agent: * Allow: / Sitemap: https://example.com/sitemap.xml
LLMS.TXT
llms.txt: what it is, and what it isn't.
llmstxt.org proposes “adding a /llms.txt markdown file to websites to provide LLM-friendly content”: a short, structured file at the site root with background, guidance and links to more detailed markdown. It's human-readable and easy for a program to parse.
It is a proposal. Google says you don't need AI text files to appear in its AI features, and none of the crawler pages quoted above says its crawler reads llms.txt. I publish one anyway because it costs little, it gives any agent that does read it a clean summary of what SET sells, and it keeps me honest: when the summary and the pages disagree, a test fails.
RUN THE CHECK
How to run it: free, $1 for an agent, or the full audit.
The free AI Visibility Checker at /ai-visibility-check reads one page's robots.txt, llms.txt and HTML, then reports which AI crawlers robots.txt allows there (OAI-SearchBot, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended and CCBot) and the rule that decided each, whether the page asks not to be indexed, where its canonical link points, whether its JSON-LD parses, and whether the site has an llms.txt. It stores nothing.
Agents can run the same check for $1.00: POST /api/agent/ai-visibility with {"url": "…"}, paid by card or Link through a Stripe Shared Payment Token, or in USDC on Base over x402 or MPP, or with a prepaid SET Agent Credits key. The page is checked before the payment is taken, and a check that fails is not charged. The same check is the ai_visibility_check tool on SET's MCP server at /api/mcp.
For the fixes rather than the findings, the SEO / AIO / GEO Audit is a one-time written audit of one site with a prioritized fix list, $497, bought on the Custom Build page. The Search and AI Visibility Review is hands-on work at $99 an hour, capped at 5 hours, with a written report and a review call. Neither guarantees rankings, traffic, indexing or AI citations, because no one outside the engines can.
curl -i -X POST https://simpleeasytransactions.com/api/agent/ai-visibility \
-H "Content-Type: application/json" \
-H "Idempotency-Key: my-check-0001" \
-d '{"url":"example.com"}'
# 402 Payment Required: pay with an MPP or x402 client and retry,
# or add: -H "Authorization: Bearer <SET Agent Credits key>"FIELD NOTES
What I've learned building this.
- This site's robots.txt allows every crawler on its public pages, AI crawlers included, and says so with a Cloudflare Content-Signal line: search=yes, ai-input=yes, ai-train=yes. It disallows /api/ and the private operator routes, so the paid agent endpoints are found through /openapi.json and the agent registries instead of a crawl.
- The first time I pointed my own checker at my own domain, it refused, and when I looked closer the domain's Custom Domains still pointed at an old version of the site. A check is only as good as what the crawler actually reaches, so I now read the live response, not the code.
- Some large sites answer a checker with 403. That's a firewall decision robots.txt can't show, and the checker reports it rather than guessing.
- Every page's JSON-LD comes from one schema module, after per-page graphs once gave one ID two different service names and declared the same FAQ twice. The guides publish their FAQ through the same builder, so the questions you see are the ones in the markup.
- llms.txt and llms-full.txt are checked by tests: the full corpus must name every guide and mirror every published FAQ answer, so the file an agent reads can't drift from the page a person reads.
- Each production deploy submits the site's URLs to IndexNow after the regression suite passes.
- The free checker leaves out ChatGPT-User and Perplexity-User on purpose: their vendors say they may not follow robots.txt, so reporting them as allowed or blocked would mislead.
HOW TO CHOOSE
How I'd choose.
- Decide search and training separately. Every major vendor now gives you a separate token for each, and blocking a training crawler doesn't remove you from that vendor's search.
- Fix the basics before anything AI-specific: robots.txt groups, edge rules, noindex, canonical, real text, and structured data that matches the page.
- Use the free check for one page, the $1.00 endpoint when an agent needs the answer as JSON, and the $497 audit or the hourly review when you want a prioritized list of fixes.
QUESTIONS
Questions people ask.
What is an AI visibility audit?
It checks whether the crawlers behind ChatGPT, Claude, Perplexity, Google's AI Overviews and Microsoft Copilot may read a site, and whether its pages can be indexed and understood: robots.txt for each crawler, firewall rules, noindex, canonical links, JSON-LD structured data and llms.txt. It reports what blocks a page; it cannot make an engine cite one.
How do I check if my site is visible to ChatGPT?
Check that robots.txt allows OAI-SearchBot, which OpenAI says surfaces websites in ChatGPT's search features, and that no firewall or CDN rule blocks it. Sites opted out of OAI-SearchBot are not shown in ChatGPT search answers. SET's free AI Visibility Checker reads the rule for one page.
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI says each setting is independent: GPTBot covers content that may be used for training, and OAI-SearchBot covers ChatGPT search. You can allow OAI-SearchBot and disallow GPTBot.
Does Google-Extended affect Google Search or AI Overviews?
Google says Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal; it controls use in Gemini training and grounding. AI Overviews and AI Mode use pages that are indexed and eligible for a snippet. To limit what they show, Google points to nosnippet, data-nosnippet, max-snippet and noindex.
Do I need an llms.txt file?
No engine in this guide requires one. llms.txt is a proposal from llmstxt.org, and Google says you don't need AI text files or new markup to appear in AI Overviews or AI Mode. It can still help agents that choose to read it.
What do AIO and GEO mean?
AIO is optimization for AI overviews in search results, such as Google's AI Overviews. GEO, generative engine optimization, comes from a 2023 research paper and means improving a page's visibility in generative engine answers. Both rest on the same basics as SEO: crawlable, indexable, clearly written pages.
How do I get my site into Microsoft Copilot answers?
Copilot answers from Bing's index, so the page must be crawlable by Bing and indexed. Bing says content tagged NOARCHIVE is left out of its AI answers. Bing Webmaster Tools' AI Performance report shows how often your pages are cited, and Bing recommends IndexNow to keep answers current.
Does structured data (JSON-LD) help with AI search?
Google says no special schema.org markup is needed for its AI features, but structured data should match the visible text. Microsoft says schema helps AI systems understand content. Structured data helps a machine read the page; it doesn't guarantee a rich result or a citation.
Can robots.txt stop every AI fetcher?
No. RFC 9309 says robots.txt rules are not access authorization, and OpenAI and Perplexity say their user-triggered fetchers, ChatGPT-User and Perplexity-User, may not follow robots.txt. Blocking needs firewall rules, which can also block crawlers you want.
Why would a crawler be blocked when robots.txt allows it?
Because a firewall, CDN bot rule or host can refuse it first. Google lists allowing crawling in robots.txt and by any CDN or hosting infrastructure as a basic, and Perplexity documents how to allow its bots through a web application firewall.
What does SET charge for an AI visibility check or audit?
The AI Visibility Checker is free for one page. Agents can run it for $1.00 a call at POST /api/agent/ai-visibility or as the MCP tool ai_visibility_check. The SEO / AIO / GEO Audit is a $497 written audit with a prioritized fix list, and the Search and AI Visibility Review is $99 an hour, capped at 5 hours.
Can an AI SEO audit guarantee citations in ChatGPT or Google?
No. Search and AI engines decide what they index, rank and cite. Google says meeting every requirement doesn't mean it will crawl, index or serve a page. An audit shows what you control and what to change.
SOURCES
Where these facts come from.
Sources checked on . Providers change their products and rules, so the linked pages are the final word.
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity Docs: Perplexity crawlers
- Google Search Central: AI features and your website (AI Overviews and AI Mode)
- Google Crawling Infrastructure: Google's common crawlers (Google-Extended)
- Google Search Central: Introduction to structured data markup
- Google Search Central: Block Search indexing with noindex
- Google Search Central: How to specify a canonical URL
- Bing Webmaster Blog: Introducing AI Performance in Bing Webmaster Tools (February 2026)
- Bing Webmaster Blog: New options for webmasters to control usage of their content in Bing Chat (September 2023)
- Microsoft Advertising: Optimizing your content for inclusion in AI search answers (October 2025)
- Apple Support: About Applebot
- Common Crawl: CCBot
- RFC 9309: Robots Exclusion Protocol
- llmstxt.org: The /llms.txt file
- Schema.org: Getting started with schema.org
- Cloudflare Docs: Managed robots.txt and Content Signals
- arXiv: GEO: Generative Engine Optimization (Aggarwal et al.)