Cookie Consent by Free Privacy Policy Generator

Back to the Campfire Blog

How to Grow Your SEO Agency and Land Leads

Written by: André Pitì Tags: analytics, data, technical-seo, future of seo

Published: Sep 3, 2026  |  Last Updated: Oct 5, 2026

Every SEO professional out there is trying to answer the same question: how do I measure AI search visibility?

In return, they find only fragmented data and grey information, together with an army of third party tools which spawned to answer that.

The truth is that the notorious "how much traffic are we getting from AI?" that marketing leaders love to ask is really multiple questions in one.

While we won't delve into the pros and cons of the so-called GEO tools here, this piece wants to bring clarity and order about these questions, and help you build sensible answers using data that you and your account likely already own.

What's wrong with AI search visibility measurement?

Let's keep this in mind: AI search visibility measurement, as of today, is imperfect.

That is partially due to the non-deterministic nature of an AI system but, at a deeper level, it's mainly because:

  1. Not every AI system or AI engine has an accessible, complete tracking system in place
  2. During an AI-driven search journey, the touchpoints are often inconsequential and sometimes unattributable

Therefore, a sensible approach is to measure what we can where we can using first-party data.

After that, we can make inferences using both first-party and external information for the pieces of the puzzle we're not able to put together directly.

We also encourage you to reframe the whole "AI search measurement" buzz into "collecting directional signals" when reporting to leadership.

We'll dive into practical ways to do this, later in this article, but first let's flag some of the common myths on this topic.

Common myths about AI search: buzz VS reality

The attribution and measurement problems in AI search helped disinformation and wobbly solution spawn.

For this, it's important to understand what to keep, what to discard, and how to set up a proper tracking of AI search signals.

Myth 1: the GA4 AI Assistant channel shows your AI traffic

Google added a native AI Assistant channel to GA4 on 13 May 2026 [1], recognising sources like ChatGPT, Gemini and Claude automatically.

That's useful, but incomplete.

Perplexity still lands in Referral, AI Overviews count as Organic Search, and a large share of AI traffic arrives with no referrer at all.

Worse, traffic from a single source such as chatgpt.com can split across the AI Assistant, Referral and Unassigned channels at once.

There are better ways to track AI-referred traffic in GA4 but, when in doubt, refer to Google's own default channel group documentation [2].

Myth 2: you can isolate AI Overviews traffic

You cannot. Google bundles AI Overviews clicks into Organic Search in both GA4 and Search Console, with no separate label.

Now, several popular SEO tools have bolted-in features with AI Overviews traffic figures, which is fine as long as everyone knows it is a model.

What this means is that every tool has its own system to estimate:

  • If AIOs get triggered after a certain keyword
  • The volume of brand mentions in AIO (generally, and according to certain prompts)

And they do it using methods that include ML predictions, third parties data sets, merging of traditional SERP results, and others.

Myth 3: public crawl-to-refer ratios apply to your site

Let's take Cloudflare's Radar data [3], a quoted number in this space.

In one sample week the ratios ran from Anthropic's roughly 70,900 crawls per referral down to Mistral sending ten times more referrals than crawl requests. Those are network-wide aggregates across millions of sites.

Your client's B2B SaaS site with less than 1,000 URLs is not the average of the internet.

⚠️ Two further cautions

The ratios are tied to a specific measurement window, and Cloudflare itself notes that referrals from native apps arrive without a referrer header, so the published ratios may overstate the imbalance.

We'll get down on how to track AI bot activity further down.

Myth 4: a rise in direct traffic proves AI referrals

That is a hypothesis, not evidence.

A reasonable heuristic is that direct traffic rising to deep technical pages alongside measured AI referrals points to dark AI traffic, but the same pattern is produced by email clients, Slack links, a podcast mention and a broken UTM, for instance.

What you can do instead:

  • Watch the actual URLs falling in the direct channel: if they have high crawl depth, or under folders with long slugs (like /blog/measure-ai-search), they're likely to not be the result of a type-in, which leaves room for inferring AI attribution
  • Check user bot activity on specific URLs for a specific timeframe: if they overlap with the URLs coming as direct, that's another signal
  • Similarly, check in your GEO tool if pulled URLs are also overlapping with the direct ones

Myth 5: The click is no longer the unit of value

SparkToro published an analysis [4] showing that 68.01% of US Google searches ended without a click in the first four months of 2026, up from 60.45% in 2024 and around 45% a decade ago. Clicks of any kind fell 9.51 points between 2024 and 2026, a 22.9% decline, while the share of searchers running another search instead rose 7.2 points.

So, judging AI by referral count is applying a shrinking yardstick to a growing behaviour. The half-percent that does arrive converts at roughly 23x, and the rest of the influence shows up as branded search and Direct.

In light of all of this, what should focus on instead?

4 critical factors to track AI search visibility signals

Here are three core distinct factors which represent three distinct data sources (although I encourage investigating other ancillary signals to collect and interpret).

They all answer different questions:

  1. Crawler hits answer "is my content being consumed?"
  2. Referred sessions answer "is anyone arriving to my website?"
  3. Mentions and citations answer "am I in the conversation?"

Below is the core vocabulary, including definitions, scope and how to measure the most relevant pieces of the puzzle with tools and techniques that we should already own as SEO pros.

1 - Crawler bot hits

AI systems use bots to crawl websites, which therefore request pages from your server.

Here, we'll explore what types of bots exist, what they are for, which one you should really care about, how to measure them, and what are the possible blockers.

Important

Bots' requests means content being taken and seen from AI, not users registering a website visit. This said, certain types of bots indicate the activity of a human user, as we'll see in the tables below.

GPTBot, ClaudeBot, PerplexityBot and friends are some examples of those, and live in your server logs, or in anything else sitting in the request path: CDN and WAF dashboards, bot-management platforms, or middleware inside your own application (what it never reaches is JavaScript analytics).

Dividing bots by clusters and business relevance

Bots can be divided by several factors in different clusters: to me, one of the most important distinctions is separating them by business relevance.

Here is how I suggest to view them:

Cluster Business meaning What 1 hit means Priority Suggested reporting cadence
A. Live AI conversation traffic Real customers asking AI systems A live AI user just got information from our site Highest Weekly
B. AI search indexing Shelf space in AI search Our page is eligible for citation High Monthly
C. AI training crawls Brand presence in AI memory Future models may know us better Medium Quarterly
D. Secondary data brokers Indirect / hygiene Content scraped by intermediaries Low Quarterly, by exception
The complete bot list to track on your clients’ server

Here’s a comprehensive list I usually pick from when I want to track bot activity for my clients.

You also might want to skim out unnecessary bots to choose the most relevant ones for the business, and to limit your costs.

The must-have ones are typically bots related to ChatGPT, Googlebot, Claude, and Perplexity. 
Bot user-agent Operator Purpose
ChatGPT-User OpenAI In-chat fetch
OAI-SearchBot OpenAI ChatGPT Search index
GPTBot OpenAI Training
Claude-User Anthropic In-chat fetch
Claude-SearchBot Anthropic Search index
ClaudeBot Anthropic Training
anthropic-ai Anthropic Training (legacy UA)
PerplexityBot Perplexity Search index
Perplexity-User Perplexity In-chat fetch
MistralAI-User Mistral In-chat fetch
DuckAssistBot DuckDuckGo AI assistant
Bingbot Microsoft Search index (powers ChatGPT Search)
Googlebot Google Search index (powers AI Overviews, Gemini)
Google-Extended Google Gemini training (robots token)
GoogleOther Google General-purpose crawl
Applebot Apple Search index
Applebot-Extended Apple Apple Intelligence training (robots token)
Meta-ExternalAgent Meta Training
Meta-ExternalFetcher Meta In-chat / on-demand fetch
Amazonbot Amazon Alexa, Q, training
Bytespider ByteDance Doubao and LLM training
cohere-ai Cohere Training
cohere-training-data-crawler Cohere Training
CCBot Common Crawl Open training datasets
YandexAdditional Yandex AI uses
PetalBot Huawei AI products
Diffbot Diffbot Structured-data layer for AI
Timpibot Timpi AI search index
ImagesiftBot TheHive.ai Image AI training
Omgilibot Webz.io Data resold to AI companies
iaskspider iAsk Chinese AI search
YouBot You.com AI search
Setting AI bot tracking up, step by step

Step zero - remove blockers

Before concluding that a bot is not visiting, check these 4 places

  1. The edge
    CDN and WAF rules are where most AI crawler blocking now happens, increasingly by default rather than by choice.
    Because Googlebot, Applebot and Bingbot all crawl for more than one purpose, a training block can take Googlebot with it [5]. Check the zone settings before you read anything into the logs.
  2. The origin
    Server-level deny rules in .htaccess, nginx config, ModSecurity or a host firewall.
  3. The application layer
    CMS, security and SEO plugins now ship AI-bot toggles, sometimes enabled by default, and they sit outside anything an infrastructure team can see.
  4. Robots.txt (as a signal)
    Not a block, a request, but a well-behaved crawler honours it, so a stray disallow explains a missing bot just as effectively as a firewall rule.
Step What to do
Step 1
Find the record
Origin access logs first, usually at /var/log/nginx/access.log or the equivalent in your hosting panel. If you have no origin access, go to the CDN instead: Cloudflare's Logpush export or its bot analytics view is the practical substitute.
Step 2
Pull and filter
Pull them on visually-clear dashboards. Never the current one, which always reads short. You can use tools like BigQuery or PostHog to organize your views — you might be surprised by the quantity of AI bots hitting your servers. Other systems like GoAccess turn the same file into a readable dashboard for free. Screaming Frog's Log File Analyser does it without a terminal.
Step 3
Verify before you believe
User-agent strings are trivially spoofed. OpenAI, Anthropic and Perplexity publish IP ranges, and Google supports reverse DNS verification. Anything claiming to be a major crawler from an unlisted IP is someone else wearing the name, and counting it inflates everything downstream.
Step 4
Cluster bots and URLs
The first tells you who is interested, and the second tells you if they are reaching the pages you want cited, hitting 301s and 404s, or burning requests on pagination and tag archives.

2 - Referred sessions

A human reads an AI answer, clicks a link, lands on your site through a trackable link.

This lives in your analytics system, like GA4, and still constitutes one of the most direct pieces of information you can get ahold of (although allegedly still a partial data whenever tracking that visit is blocked by a legal policy).

There are two main ways of isolating AI-referred traffic in GA4:

  • The first is GA4's native AI Assistant. It runs on Google's own list of AI sources, which means you cannot add to or remove from it.
  • The second one (which I like better) is to apply a custom channel regex with a list you control, for instance, over a session source filter.

Here is an example of a decently comprehensive list you can use in the formula:

chatgpt\.com|chat\.openai\.com|openai\.com|gemini\.google\.com|perplexity\.ai|claude\.ai|anthropic\.com|copilot\.microsoft\.com|edgeservices|deepseek\.com|grok\.com|x\.ai|meta\.ai|mistral\.ai|poe\.com|you\.com|phind\.com

To configure GA4 properly, build a custom channel group with a regex for AI sources and place the custom AI channel above Referral in the priority order , otherwise chatgpt.com sessions get classified as Referral instead. Review the regex quarterly, because new surfaces appear constantly.

Also, to see AI-referred sessions by country in GA4, open Reports > Acquisition > Traffic acquisition, select Session default channel group, filter to AI Assistant, and add Country as a secondary dimension, using Sessions as the metric. 
This shows AI-referred visits by location (it cannot recover missing attribution, and Google AI Overviews and AI Mode remain under Organic Search). 

3 - Google Search Console and Bing

AI Overviews and AI Mode clicks and impressions are folded into the Web search totals with no separate label, so the signal has to be inferred.

Some ways you can use GSC to collect AI search visibility gaps are:

1. Isolating impressions holding steady while clicks and CTR fall, at query and URL level, which is the classic footprint of an answer being served above you (not a proof)

2. Building cohorts with regex at a query level

GSC accepts regex in the query filter, so split your queries into an informational cohort

(^(how|what|why|when|which|who|where|can|does|is)\b), a branded cohort, and a transactional one.

Then track CTR per cohort over time. AI Overviews trigger disproportionately on informational queries, so the divergence between cohorts is a cleaner signal than any single query's decline.

3. Another useful analysis I love to run is about isolating queries that have more than 9 words in it, in the attempt to isolate relevant “prompt-alike” queries (usually with some impressions and no clicks), to understand which topical direction we should take for an account, and what we could further optimize for.

Now, this would need a separate article, but if you want to play with it,  you can use this Regex in the Query filter: ^(\S+\s+){7,}\S+$

Bing Webmaster Tools, instead, shows how often your content is cited in AI-generated answers across Copilot, Bing's AI summaries and selected partner integrations, which URLs get cited, and how that activity moves over time.

It’s an underused, free tool, and the only first-party window into the index Copilot draws from.

One of the newest new metrics is grounding queries: the reformulated queries Copilot generates internally to retrieve content, which are not the user's prompt. That is the machine's query rather than the human's, and nobody else publishes it. For anyone trying to write content that gets retrieved rather than ranked, it is the closest thing to seeing the actual input.

In June 2026 Microsoft added Intents, Topics, Citation Share and Compare, which moved it from a citation counter toward a competitive view.

4 - Mentions and citations

First, an important distinction.

A citation is a linked source: the answer points at your page, usually as a clickable reference.

A mention is your brand quoted in the body of the answer with no link attached, as in "teams in this space typically use X, Y or Z".

Both happen off your server, so no session exists and no analytics product will show you either one.

Also, they behave differently afterwards: a citation can produce a click you might partially measure. A mention can only produce a memory, and memories arrive later through a different door.

Logic would say that citations are more important than mentions, but given the instability that those links carry, several leaders in the SEO industry would advise against using only citations as a strong benchmark.

???? Follow the user

Here is an example of a user sequence: someone asks an assistant which tools handle a problem, your brand is named in the body with no link, and they go looking for you themselves.

Google search Address bar Mobile app tap
Lands in Organic Search as a branded query Lands in direct Referrer is usually stripped and it lands in Direct, Unassigned or (not set)

This is why rising branded search volume and rising direct are worth watching as proxies, and why neither is proof of anything on its own.

The shift worth making is from chasing a traffic figure to assembling signals from wherever they live, then being honest about which are measured, which are modelled and which are unknowable.

That is a conversation that touches four teams (marketing, developers, sales, and product), and the companies that get there first will spend the next years making better decisions on the same imperfect data everyone else has. 



More articles