GEO Readiness: How AI Assistants Find and Cite You

Jillian Oco, CMO

A table of open books and documents with a narrow light beam illuminating only a few pages.

AI assistants like ChatGPT, Perplexity, and Google's AI Overviews find your content through dedicated crawlers separate from their training crawlers, then decide what to cite using retrieval and ranking systems that increasingly ignore your organic search position. Getting cited now depends on how clearly a page answers one specific question, not on domain authority alone.

How do ChatGPT, Perplexity, and Google's AI Overviews actually find your pages?

Each platform runs more than one crawler, and they do different jobs. OpenAI documents three separate ones: GPTBot, which gathers content to train its models, OAI-SearchBot, which is the crawler that actually surfaces your pages inside ChatGPT's search results, and ChatGPT-User, which fetches a live page in the moment a person's question triggers a real-time lookup. Perplexity runs a similar split, with PerplexityBot building the index its answers cite from and Perplexity-User handling on-demand fetches. Perplexity has said directly that PerplexityBot is not used to train its foundation models, it exists only to index content for search and citation.

That separation matters for an MSP site because blocking the wrong bot in robots.txt can quietly cut you out of citations while leaving your regular search visibility untouched, or the reverse. Each of these crawlers can be allowed or blocked independently, and a site that never updated its robots.txt for anything beyond Googlebot may be blocking a citation crawler without knowing it. A firewall rule added to stop aggressive scraping, or a security plugin set to block unrecognized bots by default, can catch OAI-SearchBot or PerplexityBot in the same net as a bad actor. Checking your current robots.txt and firewall allowlist against the current crawler names, not a list from a year ago, is a five-minute task worth doing before you touch anything else on this list.

CrawlerCompanyWhat it actually does
GPTBotOpenAITrains foundation models, not tied to citations
OAI-SearchBotOpenAIIndexes pages for ChatGPT search citations
ChatGPT-UserOpenAILive fetch when a user's question triggers a real-time lookup
PerplexityBotPerplexityIndexes pages for search and citation, not training
Perplexity-UserPerplexityLive fetch for a user's in-session query

Google is the outlier here. AI Overviews draw from Google's existing search index rather than a separate crawl, which the prior post on this site covered in more depth. What changed is how that index gets used once a query comes in, and that's the part worth understanding next.

Why ranking number one doesn't guarantee an AI citation anymore

Being the top organic result used to be a strong predictor of getting cited in an AI Overview. It no longer is. An Ahrefs analysis of roughly 863,000 keywords and about 4 million AI Overview citation URLs found that only 38 percent of cited pages also ranked in the top 10 organic results for that same query, down sharply from 76 percent in an earlier version of the same study run about a year prior.

The reason is a process called query fan-out. Instead of answering your exact search phrase, Google's AI system splits it into several related sub-questions, runs each one separately, and then pulls citations from whichever pages show up most consistently across that whole cluster of sub-queries, not just the one query you typed. A page that ranks first for the literal phrase but doesn't cover the adjacent questions a buyer would naturally ask next has a weaker shot at citation than a page that answers the whole topic cluster, even from a lower ranking position.

For an MSP marketing site this means a single, narrowly-targeted service page optimized for one keyword phrase is a weaker citation candidate than a page structured to answer the full set of questions a prospect would actually ask around that topic, one after another, in the same session. A page on managed detection and response that only defines the service misses the fan-out. A page that also covers what it costs, how it differs from a basic antivirus tool, and how fast a team actually responds to an alert is answering the whole cluster an assistant is likely to generate, and that's the page more likely to survive being cited across several of those sub-queries at once.

Does llms.txt actually get you cited?

No, not on Google, and there's no evidence it changes anything on the other major assistants either. Google's John Mueller was asked directly whether the appearance of llms.txt files on some Google properties amounted to an endorsement of the format, and he said no. No Google Search system reads or acts on it. The file showed up on a handful of Google's own developer docs because of an unrelated content management platform change, not because Search started consuming it.

If you've had a developer or an agency spend time generating an llms.txt file for your site, that work isn't harmful, but it also isn't the lever you were told it was. The lever is still the content itself: whether a page states its answer plainly enough, early enough, and in a structure clean enough for a retrieval system to lift a passage out with confidence that it answers the question on its own.

What's a ghost citation, and why should you care?

A ghost citation is when an AI assistant links to your page as a source but never actually names your brand in the answer text the reader sees. Semrush's research with analyst Kevin Indig, built on 3,981 domain appearances across 115 prompts in 14 countries and four major AI search engines, found that 61.7 percent of citations fall into this category. The reader gets your information without ever learning it came from you.

That gap isn't random. The same research found that comparison-style content, the kind that directly weighs one option against another, earns brand mentions at 2.4 times the rate of purely informational content. An assistant summarizing "what is remote monitoring and management" has little reason to name a source. An assistant answering "how does an RMM tool compare to endpoint protection for a 40-person office" has a much stronger reason to attribute a specific claim to a specific name, because the comparison itself is the value being delivered.

This is the practical stake for an MSP marketing site. A citation with no name attached protects nothing about your reputation and drives no direct traffic. Structuring key pages around genuine comparisons and named tradeoffs, rather than only generic explainers, is one of the few levers you actually control here. If your service pages read like neutral definitions instead of a specific point of view on how your approach differs from the obvious alternative, you're producing exactly the kind of content that gets summarized without ever getting named.

What actually earns the citation, then

Three things show up consistently across how these systems are described to work: a clear direct-answer opening that a retrieval system can lift as a standalone passage, content organized around the specific sub-questions a buyer would ask in sequence rather than one broad keyword, and enough named specificity, real comparisons, real numbers, real tradeoffs, that an assistant has something concrete worth attributing to you by name instead of paraphrasing generically.

None of that is a one-time fix. The crawler rules change, the fan-out behavior changes, and platform citation rates move by the month, which is exactly why treating this as a monitored process rather than a project matters. ActiScore scores a site against current AI-citation and structural signals on an ongoing basis, so a page that quietly stopped earning citations because a competitor restructured around a comparison you never wrote gets flagged instead of discovered by accident three months later.

Is this worth building a real process around?

Yes, and the traffic shift backs that up. Similarweb's tracking shows referral traffic from ChatGPT to sites nearly tripled following a platform update earlier this year, and the split of overall generative AI traffic is shifting too, with ChatGPT's share of that traffic sliding as Gemini and Claude pick up meaningfully more of it. An MSP site optimized only for one assistant's current behavior is optimizing for a moving target that's already moved once this year.

If you're weighing where a tool like this fits against everything else you already resell, the stack builder walks through how AI-search readiness pairs with the rest of your offering in a few guided questions. And if your own site is still running on last year's SEO checklist while you pitch clients on staying current, that's worth fixing before the next pitch. See the full stack.

Sources: OpenAI Help Center, "Publishers and Developers FAQ" | Perplexity, "Perplexity Crawlers" documentation | Ahrefs, "Google AI Overview Citations From Top-Ranking Pages Drop Sharply" (via Search Engine Journal reporting) | Search Engine Roundtable, "Google Search Team Does Not Endorse LLMs.txt Files" | Semrush, "The Ghost Citations Study" | Similarweb, "ChatGPT Referral Traffic Near Triples Overnight."