Written by: Adrienne Kmetz Tags: technical-seo, data-pipelines, template
Published: Apr 10, 2026 | Last Updated: Apr 16, 2026
Our First Community Article
Experts in The SEO Community have shared their insights on this topic in Slack, and we've collated the perspectives into this group article. If you have something to add, reach out via Slack.
Contributors:
Victor Pan — multiple insights on tracking limitations
Mike Sonders — prompt methodology
Martin McGarry — the AI ranking paradox
Pedro Dias — foundational research via Visively
Adrienne Kmetz — Wrote and structured the surrounding article
The frustrating reality with measuring your brands' visibility in LLMs is that with hyper personalized results, the citation data changes with every prompt. And getting the volume you need to tease a trend out of the data, costs money.
This article is a community-written guide pulling together hard-won lessons from practitioners working in AI search optimization and visibility measurement.
We'll cover how AI visibility tracking works, where it breaks down, how to DIY it, what tools exist, and what you should do regardless of the data you have.
At its core, AI visibility tracking means asking an AI model a question (or a series of questions) and recording how your brand shows up in the response. Simple enough in theory. But the moment you go deeper, things get complicated.
AI models don't return a static ranked list like a SERP. They generate an answer that is, to varying degrees, personalized to the user: their history, their location, the model version they're using, even whether they're on a free or paid plan.
People who are pushing for AI Search tracking are missing the core point. AI is meant to be a generational leap, delivering dynamic, deeply personal results unique to the user's immediate chat/search & context. Yet, ranking demands stability and repeatability to some extent.
— Martin McGarryThere's a real philosophical tension here. The more "intelligent" and personalized an AI becomes, the less its outputs look like something you can track at scale.
That said, the community consensus is that tracking is still worth doing, you just have to do it carefully, be open about your assumptions, and understand what you're measuring and where it falls short.
This assumes that AI search optimization is about optimizing for very specific prompts. But if you're doing it right, you're just helping the AI connect entities (like your brand) to the desired concepts.
— Mike SondersAppearances that aren't backed by strong retrieval signals are inconsistent and unreliable. LLMs are probabilistic generators, not ranked indices.
— Pedro Dias, VisivelyRead Pedro's full article on how AI visibility works, how the bots work, and the underlying tech that you'll need to understand before moving forward on an AI data plan.
In other words: Your goal is to be so clearly associated with a topic or concept that the AI reaches for your brand naturally in case the user types a short keyword or has a long, context-rich conversation.
Before you invest heavily in AI visibility tracking, you need to understand the constraints. There are a lot of them.
Your visibility scores are only as good as the prompts you're tracking. But prompts are different for everyone based on what they're looking for, and they tend to be longer, not shorter, than question-based keywords. Therefore tracking one prompt, or the wrong prompts, doesn't give you any data that you can make a decision on.
Visibility scores are currently dependent on the prompts you track. Garbage prompts in, garbage breakdown of visibility out.
— Victor PanAI answers aren't deterministic. The same prompt can yield different results on consecutive runs, which means a single measurement gives you very little signal.
You don't get a truly representative result if you just submit a prompt once — just because your brand appears in the response this time, it doesn't mean it'll appear next time. I think you need to submit each prompt (or a close variant of each prompt with the same search intent) at least 3 times to see if your brand appears every time, sometimes, a bit of the time, or never.
— Mike SondersMost tracking tools use free accounts to query AI platforms but free and paid users don't get the same experience. This is a significant gap.
Paid accounts use a different model, pull results faster (lower latency), might even use a different web search set for grounding, and generally has a higher willingness-to-ground prompts to web browsing. Currently nobody (AFAIK) is doing paid account visibility tracking.
— Victor PanMany AI platforms rate-limit the number of "deep thinking" queries free users can run. If your tracking tool depends on those, it may be systematically undercounting how often enhanced reasoning features are used in real answers.
Free accounts have token limits for the amount of times they can 'think deeply' so you need a lot of them otherwise you're undercounting how often Thinking is used.
— Victor PanReal users accumulate memory in their AI accounts. The AI learns their preferences, context, history, and surfaces answers accordingly. Tracking tools start fresh every time.
Real users have memory in their free accounts, which bias the direction of the answer. Prompt tracking does not (and should not).
— Victor PanIf you need country- or language-specific data, you need accounts tied to those locations. That's an additional operational and logistical burden that most tools haven't fully solved. AI visibility data is useful directional signal, not ground truth. Treat it the way you'd treat early organic search data in the late 2000s, that was worth tracking, worth acting on, but requiring a healthy dose of skepticism.
All of these limitations point to something deeper. Martin McGarry put it best:
The paradox of AI rank tracking: for the ranking system to even function, we would have to admit the system is not acting as a generational AI; but it is merely repeating stable data. Which is not the AI we are being sold and immediately undermines the entire promise of AI.
— Martin McGarryWhich is why the goal was never to find a perfect measurement system, it was to understand what you're measuring, and act accordingly.
Victor put it well. When evaluating new AI tools, look at:
Not ready to invest in a dedicated tool? You can get meaningful signal by doing this manually or with a lightweight setup. Here's how.
Start with the questions your ideal customer is asking when they're close to a buying decision. These should map to your ICPs (Ideal Customer Profiles), not general industry queries. Think: "best [category] for [specific use case]" instead of "what is [category]."
The closest we can get to simulating a real user right now is adding a role to the prompt, and Victor Pan explains why that matters:
The system prompts reveal the extent OpenAI's willing to store a bio of the user. So the approximation on personalization has been "role" in prompt. That's hotly debated but our best approximation for now.
— Victor PanIn practice: if your ICP is a marketing director at a mid-size SaaS company, say so in the prompt. It won't perfectly replicate a real user's context, but it's the best signal we have right now.
A good starting prompt set has 20–50 queries spread across awareness, consideration, and decision-stage intent.
At minimum: ChatGPT, Perplexity, and Gemini. Add Claude and Copilot if your audience skews enterprise or technical. Don't try to track everything at once it's better that you find depth on fewer platforms.
Following the guidance from the community: run each prompt at least 3 times, ideally on different days and at different times. Record whether your brand appears (yes/no), where in the answer, and how it's framed (positive, neutral, negative).
Some would argue you'd need to get to 150 runs in order to have a dataset that is reliable enough to make decisions on. Once you get past this part, you'll need to use a tool to get that kind of volume.
A simple spreadsheet works. Columns: prompt, platform, date, brand mentioned (Y/N), mention type (top rec / list item / passing mention), sentiment. Review monthly for trends.
Read the mentions. Is the AI describing your brand accurately? Is it citing the right use cases? Is it mentioning outdated pricing or discontinued features? This qualitative layer is often more actionable than the raw visibility number.
A growing category of AI visibility platforms has emerged to automate what we described above. The space is still maturing rapidly, capabilities and pricing change often, so verify current features directly with vendors.
When evaluating any AI tool, ask these questions:
The TLDR — it's a technical and long problem to troubleshoot when 'AEO visibility' is down. Every limitation becomes scrutinized. The user behavior is changing. The tools are desperately trying to catch up.
— Victor PanNo tool currently solves all the limitations described in the previous section. Use them as directional dashboards, not authoritative measurement systems.
So I tested ~20 tools over the last few months. Paid for most of them myself just to understand the differences.
Quick thoughts:
Enterprise stuff (Profound, Conductor, Semrush AI, AthenaHQ, etc.)
→ Very polished. Very expensive. Feels built for big marketing teams.
Mid-tier tools (AIclicks.ai, Peec, Writesonic, Surfer AI Tracker, Rankscale, Scrunch, Omnia, LLMClicks.ai, etc.)
→ More practical. Prompt tracking, brand mentions, share of voice across models.
Some give actual action steps. Some are just dashboards.
Indie tools (Otterly, Rank Prompt, Waikay, Passionfruit Labs, etc.)
→ More focused. Some track citations. Some check if AI is getting your facts wrong. Budget-friendly entry points.
What I learned:
No tool is “Search Console for AI” yet.
Results change depending on how you phrase the prompt.
API results and actual chat UI don’t always match.
Most tools are 70% similar.
The real value isn’t just “are we mentioned?”
It’s:
Why are we mentioned?
Which sources triggered it?
What does AI think our brand actually is?
Also… tools don’t fix weak positioning. Clear messaging + strong entity signals still matter more than dashboards.
The best things you can do for AI visibility don't require perfect measurement. They're about making your brand easy for AI systems to understand, trust, and recommend.
AI models work with entities, named things with defined attributes. Make sure your brand, products, and key people are clearly described and consistently represented across your website, PR, and third-party profiles.
Wikipedia entries, Wikidata, and authoritative directory listings all help. Before any of this works, AI systems need to be able to crawl your content in the first place. The Community has extensive thoughts about how AI bots behave with noindex and robots.txt.
Generic content doesn't give AI models anything to latch onto. The more specific and opinionated your content is about your use cases, your ideal customers, and your differentiation, the easier it is for AI to associate you with specific concepts.
AI models are grounded in web content and they favor sources they've learned to trust: major publications, established industry sites, academic papers. Getting mentioned and cited in those contexts matters more than publishing on your own domain.
As Mike Sonders noted: the goal is entity-concept association, not keyword ranking. Ask yourself: "When an AI thinks about [my category], do the signals it's learned clearly point to us?"
Build content and PR strategies that reinforce that connection. Think about what a knowledgeable friend would say about your brand in conversation. That's what you want AI to say. Write, publish, and earn coverage that makes that happen.
Your AI visibility strategy should include a review process for what the AI actually says. Wrong pricing, outdated product names, misattributed claims. These are reputational risks that don't show up if you're only counting brand mentions.
Use tracking tools to find inaccuracies to fix, not positions to climb. Entity-level monitoring is more stable than prompt-level rankings.
— Pedro Dias, VisivelyAI visibility is not a channel you can optimize in a sprint. It reflects the cumulative signal across your entire digital presence. Expect a minimum 3–6 month lag between content and PR efforts and measurable changes in how AI models talk about you.
The tools and methodologies are still catching up to the complexity of the problem. But that's not a reason to wait, it's a reason to build good habits now, before the space matures and your competitors figure it out.
Measure what you can. Act on what's directionally clear. And stay close to communities where practitioners are sharing what works.