AI systems can crawl your pages without citing them, cite your research without naming your brand, and recommend your products without linking to your website.
The problem with citation-led AEO reporting isn’t that citations are irrelevant. It’s that we’re measuring the final output of a process we can barely observe.
Consider a scenario.
You publish an original research report containing proprietary industry data. Over the following weeks, your server logs record thousands of requests from AI-related crawlers. Your content is accessible, Google indexes your pages, and several industry publications reference your findings.
Yet your AI visibility dashboard reports a disappointing citation share. What should you conclude?
Perhaps your content isn’t being selected during retrieval. Perhaps it is retrieved but loses out during source selection. Perhaps an AI system uses information from a third-party article discussing your research. Or perhaps your content is encountered by a training crawler, making its subsequent contribution to any particular answer practically impossible to establish.
Your citation dashboard can’t distinguish between these possibilities.
Now consider the reverse. Your citation share increases, but the answers citing your research never mention your brand. Your content helps an AI assistant explain an industry problem, while the assistant recommends three competing products.
Technically, your citation performance has improved. Commercially, the outcome is considerably less clear. This is the problem with treating citations as the primary measure of answer engine optimization (AEO).
We’ve taken one observable output from a complex information-retrieval process and started using it as a proxy for discovery, visibility, authority, and sometimes even commercial influence.
Those are four different things. Understanding the differences requires us to look at how AI systems actually encounter and process web content.
We’re measuring the output without understanding the pipeline
Anyone who has worked extensively in SEO should recognize this problem.
A page can be discoverable without being crawled. Crawled without being indexed. Indexed without ranking. And ranked without generating meaningful traffic.
We’ve spent decades developing diagnostic approaches to distinguish between these stages.
Yet much of AEO reporting skips directly to the final answer.
A typical citation-monitoring tool queries several AI platforms using a predefined collection of prompts, collects their responses, extracts citation URLs, and calculates how frequently the monitored domain appears.
There’s nothing inherently wrong with that methodology, provided we understand what it measures.
It measures citation frequency within a particular sample of generated answers.
It doesn’t measure the proportion of all relevant searches in which your content was considered. It doesn’t establish whether your pages were retrieved but discarded. It cannot observe information absorbed during model training. And it doesn’t necessarily capture brand recommendations supported by third-party sources.
The underlying technical process is more complicated.
This is a conceptual model of web-grounded answer generation, not a claim that every LLM implements these six stages in this order.
Some responses may be generated without live web retrieval. Others may involve multiple searches, different retrieval systems, or repeated evidence-gathering steps.
Google explicitly describes how AI Overviews and AI Mode can use query fan-out: conducting multiple related searches across subtopics and data sources before constructing an answer. Its documentation also makes clear that eligibility to appear in its AI search features depends on conventional search indexing and snippet eligibility.
That creates a familiar dependency: if an important page cannot be crawled, rendered, or indexed appropriately, optimizing its content for potential AI citations misses a more fundamental problem.
But Google Search isn’t representative of every AI discovery system. That’s why we need to examine crawler behavior separately.
Not every AI bot visiting your website serves the same purpose
One of the more persistent mistakes in AI visibility discussions is treating all AI crawler activity as evidence that an assistant is discovering your content for answers.
A bot request tells you that an HTTP request occurred. You still need to establish who made it, why it was made, what the server returned, and whether the requested information was actually usable.
OpenAI’s documentation provides a useful illustration of the distinction between crawler purposes.
A successful GPTBot visit doesn’t mean your content was used for training or influenced an AI-generated answer. Similarly, an OAI-SearchBot visit doesn’t guarantee your page will appear in ChatGPT Search. Even ChatGPT-User traffic doesn’t prove your brand was mentioned or your page was cited, since users can ask ChatGPT to visit specific URLs.
Google introduces a different set of considerations. Its AI search features operate through Google’s existing search infrastructure, while Google-Extended controls certain other AI uses separately. Blocking Google-Extended is not equivalent to removing a page from Google Search.
This matters when interpreting server logs or implementing crawler-access policies.
A technical audit should distinguish at least three questions: Are the relevant systems permitted to access the content? Are they actually requesting it? And are they receiving an appropriate response?
We shouldn’t collapse those questions into a single metric called AI crawl visibility.
The crawl-to-referral gap reveals how much we’re missing
Research published by Cloudflare in 2025 makes the scale of this measurement problem difficult to ignore.
Its analysis of AI crawler activity found that approximately 79% of classified AI crawling in July 2025 was associated with training, compared with 17% for search and approximately 3% for user-initiated activities.
Cloudflare classified crawler purposes using operator disclosures and industry information. These classifications describe intended crawler functions, not verified downstream uses of every downloaded page.
Anthropic recorded approximately 38,066 crawl requests for each attributable referral in July 2025. The corresponding figures were approximately 1,091 for OpenAI and 195 for Perplexity.
We shouldn’t interpret these figures as evidence that thousands of pages were used without attribution. Training, search, and other crawler activities contribute to the totals, while referrals measure a completely different event.
This data reveals something important: the amount of information AI systems request from the web, and the measurable traffic they send back, differ dramatically in scale.
A citation dashboard observes neither the full acquisition process nor the complete downstream contribution of that information.
But server logs aren’t the missing answer either.
They tell us which verified bots requested which URLs, when those requests occurred, and what responses our infrastructure delivered. They’re diagnostic evidence of access, not proof of selection or influence. That difference matters even more at the next stage of the process.
Retrieval and citation are different events
One of the most useful studies on this subject is The Attribution Crisis in LLM Search Results, published in the Cambridge University Press journal Data & Policy in 2026.
The researchers analyzed approximately 14,000 real-world conversation logs involving search-enabled language models. Here are some of the findings:
Even in search mode, 15.6% of LLM responses involved no website visits, while 30% included no citations at all, highlighting a substantial gap between content consumption and recognition.
According to the models’ logs, 24% of GPT Search answers and 34% of Gemini answers did not visit any relevant websites when answering questions, instead relying on pretraining.
25% of GPT-4o answers and 92% of Gemini answers provided no citations at all.
Perplexity Sonar accessed approximately ten relevant pages per query but cited only three or four.
The researchers also spotted many differences in retrieval and attribution behavior between AI systems. And this is exactly what we should investigate as SEO practitioners.
In a conventional retrieval system, candidate documents may undergo several selection processes. An initial retrieval stage identifies potentially relevant material. Subsequent ranking or filtering prioritizes certain results, and the answer-generation process may incorporate only a subset of those documents.
A retrieved document can therefore disappear from the observable answer without anything being technically wrong with it. It could have been considered less relevant, supplied redundant information, or simply not been selected for the final response.
We shouldn’t assume that an uncited document influenced the answer. But equally, we cannot interpret the absence of a citation as evidence that the document was never retrieved.
Another peer-reviewed study, Correctness Is Not Faithfulness in Retrieval Augmented Generation Attributions, identified a deeper problem.
The researchers distinguished between citation correctness (whether a cited document supports a generated claim) and citation faithfulness (whether the system actually relied on that document to produce the claim).
Under their experimental conditions, they found that up to 57% of citations lacked faithfulness. The result applies to their evaluated systems and experimental setup, not to every commercial AI search product.
That challenges an assumption embedded in many AEO dashboards: that a citation reliably identifies the source responsible for an answer. Sometimes it may. Sometimes it may simply identify a document that supports information generated through a different process.
Counting citations without acknowledging that distinction creates an appearance of measurement precision that the underlying technology doesn’t necessarily support.
Are we confusing source attribution with brand exposure?
Even if citations were perfectly faithful, we’d still face another measurement problem.
A citation identifies a source. Brand visibility tells what users actually learn about brands.
Those two outcomes can occur independently.
Suppose an AI assistant responds to a question about content attribution by citing a research paper published on your company’s website. The response explains the findings accurately but never identifies your company.
Your domain receives a citation. Your brand receives no explicit mention in the generated answer.
Now suppose someone asks the same assistant to recommend tools for measuring AI visibility.
Your product appears among the recommendations. But the assistant supports its recommendation with links to review platforms, community discussions, and industry publications.
This time, your brand appears, but your website doesn’t receive a citation.
Which interaction is more valuable?
That depends on the intended outcome, the user’s intent, and whether the recommendation ultimately influences a decision. A citation may be valuable to a research publisher even without a brand mention. A product recommendation may be commercially meaningful even without an owned-domain citation.
The problem is treating the two outcomes as interchangeable.
A stronger measurement approach needs to differentiate between owned-domain citations, third-party citations, explicit brand mentions, and recommendations.
We should also consider the context of each recommendation. Being mentioned as an unsuitable option is not the same as being recommended for a relevant purchasing scenario. Citation frequency alone doesn’t capture these differences.
Fixing the measurement methodology
There’s a practical problem that receives surprisingly little attention in AEO reporting: sampling.
Most visibility platforms don’t observe every answer generated by every AI system. Instead, they monitor a selected collection of prompts, platforms, and locations.
That makes prompt selection a critical component of measurement.
A study tracking 20 informational questions will produce a completely different view of brand visibility from one tracking 20 purchase-oriented comparison prompts.
Neither is necessarily wrong. But neither represents the entire discovery market.
The problem becomes worse when teams change their prompts, monitoring frequency, or AI platforms while continuing to compare the resulting citation percentages as though the underlying measurement conditions were constant.
I’d approach AEO monitoring much like a technical SEO experiment.
Establish a fixed panel of commercially relevant prompts. Segment them by intent, industry, geography, and stage of the buying journey. Run repeated observations across the platforms that matter to your audience.
Record the complete answer rather than merely extracting its citations.
Track whether your brand appears, how it’s described, which competitors are mentioned, which third-party sources receive attribution, and whether relevant links are present.
Then preserve those conditions when comparing reporting periods.
Even this methodology has limitations. Generated answers vary, user experiences may differ, and external monitoring cannot fully reproduce every real-world interaction.
But the resulting data becomes considerably more interpretable.
What I’d actually measure: an AEO observability framework
The better approach isn’t replacing citation counts with crawler counts or creating another composite visibility score.
It’s measuring distinct stages of the system and preserving the limitations of each dataset. I’d organize an AEO measurement program around six layers.
The distinction between these layers is particularly useful when diagnosing a problem.
If citations decline, inspect the answer-level evidence before assuming that the underlying cause is technical. If relevant pages aren’t receiving search-related crawler requests, investigate accessibility and discovery. If pages remain indexed and accessible but aren’t cited, investigate retrieval and source-selection hypotheses.
And if citations increase without any corresponding brand mentions, relevant traffic, or measurable commercial activity, investigate what those citations actually contribute.
For the technical foundation, I’d start with three practical requirements.
First, verify crawler identity. User-agent strings can be spoofed. Where crawler operators publish IP ranges or verification methods, use them instead of blindly accepting every request claiming to originate from an AI platform. Google, for example, explicitly recommends verification using its published IP ranges or reverse DNS.
Second, evaluate successful content delivery, not just request volume. A crawler receiving a 403, 429, or 5xx response has not experienced the same interaction as one receiving the intended page content. Analyze response codes, redirects, canonical URLs, response times, and whether important content is available in the delivered HTML.
Third, separate observation from inference. A successful fetch is observed evidence. A subsequent citation is another observation. The hypothesis that one caused the other requires investigation, ideally with controlled changes and repeated measurements.
This is also where established SEO fundamentals remain relevant. Technical accessibility, internal linking, canonicalization, and index eligibility still matter. Google explicitly states that its AI search features build on its existing Search infrastructure, rather than requiring an entirely separate set of AI-specific technical optimizations.
The opportunity isn’t to abandon technical SEO for a new collection of AEO tactics. It’s to extend our existing diagnostic discipline to systems that expose considerably less information about how they operate.
The real problem with citation-led AEO reporting
For years, SEO practitioners have understood that rankings alone don’t explain search performance.
A ranking doesn’t guarantee an impression. An impression doesn’t guarantee a click. And a click doesn’t guarantee commercial value.
We’re now at risk of repeating that mistake with citations.
An AI crawler request isn’t necessarily evidence of discovery. Retrieval doesn’t guarantee selection. Selection doesn’t guarantee attribution. Attribution doesn’t guarantee brand exposure. And brand exposure doesn’t automatically translate into commercial impact.
Each transition introduces uncertainty.
Yet much of the emerging AEO measurement market compresses that uncertainty into citation counts and visibility percentages, often without explaining which parts of the underlying process remain unobservable.
Citations are useful. They provide evidence of how AI systems attribute information in the answers we’re able to monitor.
But they are not a comprehensive measure of AI visibility.
The more important task is understanding how information travels through an increasingly fragmented discovery ecosystem, identifying which stages we can reliably observe, and being explicit about the conclusions our data cannot support.
AEO doesn’t just have a visibility problem. It has an observability problem. And solving that problem requires more than counting links.


