On this page
Being cited by an AI answer engine, whether that is ChatGPT search, Perplexity, Copilot, or Google's AI Overviews and AI Mode, is not something you can guarantee with a checklist. What you can control is two separate decisions: which of your pages are worth improving first, and how each page is written and marked up once you get to it. Most teams blur those two questions together and end up either rewriting everything at once or picking pages at random. This article is about the first decision, prioritization. A companion piece on this site, about making content quotable to AI search tools, covers the second: how a page should be structured and marked up so a retrieval system can pull a clean passage out of it.
Before either decision matters, confirm the technical basics are in place. A page that a crawler cannot read cannot be prioritized into anything.
Know which crawler does what
AI companies run separate crawlers for separate jobs, and conflating them is one of the most common mistakes in this space. OpenAI documents distinct user agents: GPTBot, which gathers content to train its models, OAI-SearchBot, which crawls and indexes pages so ChatGPT's search feature can surface and cite them, and ChatGPT-User, which fetches a specific page live when a person's prompt asks ChatGPT to read it. OpenAI states plainly that each setting is independent: you can allow OAI-SearchBot for citations while disallowing GPTBot for training, and the reverse is also a valid choice.
Source: OpenAI: Overview of OpenAI Crawlers
Perplexity draws the same line: PerplexityBot builds the index behind Perplexity's answers, and Perplexity states it is not used to collect content for model training. Perplexity-User is a separate agent that fetches a page live in response to a specific user request. The practical point is that blocking one of these is a different decision from blocking the other, so check which agent a rule in your robots.txt actually affects.
Source: Perplexity: Perplexity Crawlers
Google is a different case again, and worth calling out on its own. AI Overviews and AI Mode are features inside Google Search, built on Google's regular Search index and ranking systems rather than a separate AI-only index. Google's own guidance for site owners says the established SEO best practices remain what matters, and that there are no additional requirements or special optimizations needed to appear in these features. Google-Extended is a distinct control that governs whether crawled content can be reused to train Gemini and other generative models; it is not the setting that determines whether a page can appear in AI Overviews or AI Mode.
Source: Google Search Central: AI features and your website
Microsoft's Copilot and Bing draw on Bing's own crawl and index, through Bingbot, with a separate submission path through IndexNow for fast re-crawling of updated pages. As with Google, there is no evidence of a distinct citation-only crawler separate from the crawler that builds the regular search index. Because each vendor documents this differently and updates it over time, check the current published list for each platform you care about rather than relying on a single rule of thumb.
The technical baseline, before any prioritization
- Confirm robots.txt isn't accidentally blocking the search or citation crawlers you want, for OpenAI, Perplexity, and any other engine you care about.
- Avoid content that only appears after client-side JavaScript executes; a crawler that cannot render it cannot cite it.
- Check for stray noindex, nosnippet, or max-snippet:0 tags that would quietly block an otherwise good page from being shown or quoted in a snippet or an AI feature.
- Check X-Robots-Tag response headers too; they can override what is in the page's meta tags without anyone noticing.
- Keep sitemaps current and submitted to Google Search Console and Bing Webmaster Tools, and confirm priority pages return a clean 200 status.
A prioritization method that isn't a guess
Once the technical basics are sound, score your existing pages on three factors and work down the list in order, rather than starting from whichever page you last thought about.
| Factor | What to check | Why it matters |
|---|---|---|
| Existing search performance | Is the page already getting impressions or clicks for its target query in Search Console, or is it invisible? | A page with some existing performance in normal search gives you a signal to build on; a page with none is a longer, riskier bet |
| Commercial or informational value | Does this page answer a question your actual buyers ask, not just one you would like them to ask? | Improving a page nobody searches for wastes the effort regardless of technical quality |
| Clarity of the current answer | Is the answer to the core question buried in paragraph six, or missing entirely? | A page with a real answer needs restructuring; a page with no answer needs to be written first |
Pages that score reasonably on the first two factors but poorly on clarity are usually your best return on effort: you are not building a page from zero, just making an existing answer easier to extract. Pages with no search performance at all are a longer play. Fix the technical and structural basics on them too, but set expectations carefully, because a brand-new page has less for any retrieval step to work from, and no vendor publishes what its own retrieval requires.
Track visibility with the right metrics
Traditional rank tracking alone will not tell you whether an AI engine is quoting you. Track these as well, on their own cadence, separate from classic SEO reporting:
- Citation rate: how often your domain gets referenced across a defined, repeated set of prompts you check by hand or with a monitoring tool.
- Mention position: whether you are the primary quoted source or one of several sources buried lower in the answer.
- AI referral traffic: segment it distinctly in your analytics rather than lumping it in with general referral traffic.
- Search Console performance for the same queries, since Google's AI features are built on that same index and ranking layer.
Where teams waste the effort
- Publishing more pages before confirming crawlers can read what is already live.
- Treating an AI-drafted answer as finished without checking whether it is accurate for the business today.
- Blocking a training crawler and assuming that also blocks the crawler used for citations, or the reverse, without checking current documentation.
- Prioritizing by instinct or by whichever page was last discussed internally, instead of scoring the list.
- Judging the whole effort after a few weeks. Technical fixes can show up in logs quickly, but building the kind of search performance these systems draw on takes longer and is not guaranteed on any specific timeline.
None of this is a formula that outputs a guaranteed citation. It is a way to stop guessing about where to start: confirm crawlers can read your site, score pages on visibility, value and clarity instead of picking at random, and measure the right things so you know whether the effort is working. What to actually change on the page once you have picked it, structure, headings, the shape of an FAQ, is a separate piece of work, covered in the companion article on making content quotable to AI search tools.
Reviewing pages, prioritizing the list and keeping the technical basics current is exactly the kind of recurring, expert-managed work the monthly cycle is built to handle.
See how the membership works