Written by: André Pitì Tags: analytics, data, technical-seo, future of seo
Published: Sep 3, 2026 | Last Updated: Oct 5, 2026
Every SEO professional out there is trying to answer the same question: how do I measure AI search visibility?
In return, they find only fragmented data and grey information, together with an army of third party tools which spawned to answer that.
The truth is that the notorious "how much traffic are we getting from AI?" that marketing leaders love to ask is really multiple questions in one.
While we won't delve into the pros and cons of the so-called GEO tools here, this piece wants to bring clarity and order about these questions, and help you build sensible answers using data that you and your account likely already own.

Let's keep this in mind: AI search visibility measurement, as of today, is imperfect.
That is partially due to the non-deterministic nature of an AI system but, at a deeper level, it's mainly because:
Therefore, a sensible approach is to measure what we can where we can using first-party data.
After that, we can make inferences using both first-party and external information for the pieces of the puzzle we're not able to put together directly.
We also encourage you to reframe the whole "AI search measurement" buzz into "collecting directional signals" when reporting to leadership.
We'll dive into practical ways to do this, later in this article, but first let's flag some of the common myths on this topic.
The attribution and measurement problems in AI search helped disinformation and wobbly solution spawn.
For this, it's important to understand what to keep, what to discard, and how to set up a proper tracking of AI search signals.
Google added a native AI Assistant channel to GA4 on 13 May 2026 [1], recognising sources like ChatGPT, Gemini and Claude automatically.
That's useful, but incomplete.
Perplexity still lands in Referral, AI Overviews count as Organic Search, and a large share of AI traffic arrives with no referrer at all.
Worse, traffic from a single source such as chatgpt.com can split across the AI Assistant, Referral and Unassigned channels at once.
There are better ways to track AI-referred traffic in GA4 but, when in doubt, refer to Google's own default channel group documentation [2].
You cannot. Google bundles AI Overviews clicks into Organic Search in both GA4 and Search Console, with no separate label.
Now, several popular SEO tools have bolted-in features with AI Overviews traffic figures, which is fine as long as everyone knows it is a model.
What this means is that every tool has its own system to estimate:
And they do it using methods that include ML predictions, third parties data sets, merging of traditional SERP results, and others.
Let's take Cloudflare's Radar data [3], a quoted number in this space.
In one sample week the ratios ran from Anthropic's roughly 70,900 crawls per referral down to Mistral sending ten times more referrals than crawl requests. Those are network-wide aggregates across millions of sites.
Your client's B2B SaaS site with less than 1,000 URLs is not the average of the internet.
⚠️ Two further cautions
The ratios are tied to a specific measurement window, and Cloudflare itself notes that referrals from native apps arrive without a referrer header, so the published ratios may overstate the imbalance.
We'll get down on how to track AI bot activity further down.
That is a hypothesis, not evidence.
A reasonable heuristic is that direct traffic rising to deep technical pages alongside measured AI referrals points to dark AI traffic, but the same pattern is produced by email clients, Slack links, a podcast mention and a broken UTM, for instance.
What you can do instead:
SparkToro published an analysis [4] showing that 68.01% of US Google searches ended without a click in the first four months of 2026, up from 60.45% in 2024 and around 45% a decade ago. Clicks of any kind fell 9.51 points between 2024 and 2026, a 22.9% decline, while the share of searchers running another search instead rose 7.2 points.
So, judging AI by referral count is applying a shrinking yardstick to a growing behaviour. The half-percent that does arrive converts at roughly 23x, and the rest of the influence shows up as branded search and Direct.
In light of all of this, what should focus on instead?
Here are three core distinct factors which represent three distinct data sources (although I encourage investigating other ancillary signals to collect and interpret).
They all answer different questions:
Below is the core vocabulary, including definitions, scope and how to measure the most relevant pieces of the puzzle with tools and techniques that we should already own as SEO pros.
AI systems use bots to crawl websites, which therefore request pages from your server.
Here, we'll explore what types of bots exist, what they are for, which one you should really care about, how to measure them, and what are the possible blockers.
Important
Bots' requests means content being taken and seen from AI, not users registering a website visit. This said, certain types of bots indicate the activity of a human user, as we'll see in the tables below.
GPTBot, ClaudeBot, PerplexityBot and friends are some examples of those, and live in your server logs, or in anything else sitting in the request path: CDN and WAF dashboards, bot-management platforms, or middleware inside your own application (what it never reaches is JavaScript analytics).
Bots can be divided by several factors in different clusters: to me, one of the most important distinctions is separating them by business relevance.
Here is how I suggest to view them:
| Cluster | Business meaning | What 1 hit means | Priority | Suggested reporting cadence |
|---|---|---|---|---|
| A. Live AI conversation traffic | Real customers asking AI systems | A live AI user just got information from our site | Highest | Weekly |
| B. AI search indexing | Shelf space in AI search | Our page is eligible for citation | High | Monthly |
| C. AI training crawls | Brand presence in AI memory | Future models may know us better | Medium | Quarterly |
| D. Secondary data brokers | Indirect / hygiene | Content scraped by intermediaries | Low | Quarterly, by exception |
Here’s a comprehensive list I usually pick from when I want to track bot activity for my clients.
You also might want to skim out unnecessary bots to choose the most relevant ones for the business, and to limit your costs.
| Bot user-agent | Operator | Purpose |
|---|---|---|
| ChatGPT-User | OpenAI | In-chat fetch |
| OAI-SearchBot | OpenAI | ChatGPT Search index |
| GPTBot | OpenAI | Training |
| Claude-User | Anthropic | In-chat fetch |
| Claude-SearchBot | Anthropic | Search index |
| ClaudeBot | Anthropic | Training |
| anthropic-ai | Anthropic | Training (legacy UA) |
| PerplexityBot | Perplexity | Search index |
| Perplexity-User | Perplexity | In-chat fetch |
| MistralAI-User | Mistral | In-chat fetch |
| DuckAssistBot | DuckDuckGo | AI assistant |
| Bingbot | Microsoft | Search index (powers ChatGPT Search) |
| Googlebot | Search index (powers AI Overviews, Gemini) | |
| Google-Extended | Gemini training (robots token) | |
| GoogleOther | General-purpose crawl | |
| Applebot | Apple | Search index |
| Applebot-Extended | Apple | Apple Intelligence training (robots token) |
| Meta-ExternalAgent | Meta | Training |
| Meta-ExternalFetcher | Meta | In-chat / on-demand fetch |
| Amazonbot | Amazon | Alexa, Q, training |
| Bytespider | ByteDance | Doubao and LLM training |
| cohere-ai | Cohere | Training |
| cohere-training-data-crawler | Cohere | Training |
| CCBot | Common Crawl | Open training datasets |
| YandexAdditional | Yandex | AI uses |
| PetalBot | Huawei | AI products |
| Diffbot | Diffbot | Structured-data layer for AI |
| Timpibot | Timpi | AI search index |
| ImagesiftBot | TheHive.ai | Image AI training |
| Omgilibot | Webz.io | Data resold to AI companies |
| iaskspider | iAsk | Chinese AI search |
| YouBot | You.com | AI search |
Step zero - remove blockers
Before concluding that a bot is not visiting, check these 4 places
| Step | What to do |
|---|---|
| Step 1 Find the record |
Origin access logs first, usually at /var/log/nginx/access.log or the equivalent in your hosting panel. If you have no origin access, go to the CDN instead: Cloudflare's Logpush export or its bot analytics view is the practical substitute. |
| Step 2 Pull and filter |
Pull them on visually-clear dashboards. Never the current one, which always reads short. You can use tools like BigQuery or PostHog to organize your views — you might be surprised by the quantity of AI bots hitting your servers. Other systems like GoAccess turn the same file into a readable dashboard for free. Screaming Frog's Log File Analyser does it without a terminal. |
| Step 3 Verify before you believe |
User-agent strings are trivially spoofed. OpenAI, Anthropic and Perplexity publish IP ranges, and Google supports reverse DNS verification. Anything claiming to be a major crawler from an unlisted IP is someone else wearing the name, and counting it inflates everything downstream. |
| Step 4 Cluster bots and URLs |
The first tells you who is interested, and the second tells you if they are reaching the pages you want cited, hitting 301s and 404s, or burning requests on pagination and tag archives. |
A human reads an AI answer, clicks a link, lands on your site through a trackable link.
This lives in your analytics system, like GA4, and still constitutes one of the most direct pieces of information you can get ahold of (although allegedly still a partial data whenever tracking that visit is blocked by a legal policy).
There are two main ways of isolating AI-referred traffic in GA4:
Here is an example of a decently comprehensive list you can use in the formula:
chatgpt\.com|chat\.openai\.com|openai\.com|gemini\.google\.com|perplexity\.ai|claude\.ai|anthropic\.com|copilot\.microsoft\.com|edgeservices|deepseek\.com|grok\.com|x\.ai|meta\.ai|mistral\.ai|poe\.com|you\.com|phind\.com
To configure GA4 properly, build a custom channel group with a regex for AI sources and place the custom AI channel above Referral in the priority order , otherwise chatgpt.com sessions get classified as Referral instead. Review the regex quarterly, because new surfaces appear constantly.
Also, to see AI-referred sessions by country in GA4, open Reports > Acquisition > Traffic acquisition, select Session default channel group, filter to AI Assistant, and add Country as a secondary dimension, using Sessions as the metric.
This shows AI-referred visits by location (it cannot recover missing attribution, and Google AI Overviews and AI Mode remain under Organic Search).
AI Overviews and AI Mode clicks and impressions are folded into the Web search totals with no separate label, so the signal has to be inferred.
Some ways you can use GSC to collect AI search visibility gaps are:
1. Isolating impressions holding steady while clicks and CTR fall, at query and URL level, which is the classic footprint of an answer being served above you (not a proof)
2. Building cohorts with regex at a query level
GSC accepts regex in the query filter, so split your queries into an informational cohort
(^(how|what|why|when|which|who|where|can|does|is)\b), a branded cohort, and a transactional one.
Then track CTR per cohort over time. AI Overviews trigger disproportionately on informational queries, so the divergence between cohorts is a cleaner signal than any single query's decline.
3. Another useful analysis I love to run is about isolating queries that have more than 9 words in it, in the attempt to isolate relevant “prompt-alike” queries (usually with some impressions and no clicks), to understand which topical direction we should take for an account, and what we could further optimize for.
Now, this would need a separate article, but if you want to play with it, you can use this Regex in the Query filter: ^(\S+\s+){7,}\S+$
Bing Webmaster Tools, instead, shows how often your content is cited in AI-generated answers across Copilot, Bing's AI summaries and selected partner integrations, which URLs get cited, and how that activity moves over time.
It’s an underused, free tool, and the only first-party window into the index Copilot draws from.
One of the newest new metrics is grounding queries: the reformulated queries Copilot generates internally to retrieve content, which are not the user's prompt. That is the machine's query rather than the human's, and nobody else publishes it. For anyone trying to write content that gets retrieved rather than ranked, it is the closest thing to seeing the actual input.
In June 2026 Microsoft added Intents, Topics, Citation Share and Compare, which moved it from a citation counter toward a competitive view.
First, an important distinction.
A citation is a linked source: the answer points at your page, usually as a clickable reference.
A mention is your brand quoted in the body of the answer with no link attached, as in "teams in this space typically use X, Y or Z".
Both happen off your server, so no session exists and no analytics product will show you either one.
Also, they behave differently afterwards: a citation can produce a click you might partially measure. A mention can only produce a memory, and memories arrive later through a different door.
Logic would say that citations are more important than mentions, but given the instability that those links carry, several leaders in the SEO industry would advise against using only citations as a strong benchmark.
???? Follow the user
Here is an example of a user sequence: someone asks an assistant which tools handle a problem, your brand is named in the body with no link, and they go looking for you themselves.
| Google search | Address bar | Mobile app tap |
| Lands in Organic Search as a branded query | Lands in direct | Referrer is usually stripped and it lands in Direct, Unassigned or (not set) |
This is why rising branded search volume and rising direct are worth watching as proxies, and why neither is proof of anything on its own.
The shift worth making is from chasing a traffic figure to assembling signals from wherever they live, then being honest about which are measured, which are modelled and which are unknowable.
That is a conversation that touches four teams (marketing, developers, sales, and product), and the companies that get there first will spend the next years making better decisions on the same imperfect data everyone else has.