How to Track Your Brand's Visibility in AI Search (2026 Guide)

AI SEO

Bind Hero Image and Hero Image Alt

If your only measure of search performance is a rankings tracker, you're already blind. Google's AI Overviews now absorb 47.5% of desktop publisher click-throughs when they fire (Authoritas, 2025). The click didn't disappear. It got absorbed into the answer. Tracking AI visibility is what you do about it, and Semrush's 126-million-prompt 2026 AI Visibility Index confirms the ugly truth: 45% of marketing leaders can't measure their AI visibility at all, and only 9% have the tools to track it cross-platform.

The commercial urgency is bigger than a rankings dashboard. Mediassociates' 2026 Marketers' Guide to AI Discoverability documented that 37% of consumers now start their searches with AI tools instead of Google or Bing, and 60% of searches end without a click at all. When an AI Overview appears on the SERP, click-through rate on the top organic listing drops by roughly a third. Both organic and paid search are being squeezed at the same time.

And when AI-referred traffic does arrive on your site, it converts at up to four times the rate of traditional organic traffic. AI visibility is not just a brand-awareness metric. It is the highest-converting acquisition channel in most brand's mix, and the one most brands are not measuring.

The framing that reconciles the three channels: paid search, organic SEO, and AI visibility are not three separate channels. They are one integrated ecosystem. A brand that earns inclusion in an AI answer before a user even formulates a search query is in a stronger position on both paid and organic than a brand that does not. Tracking AI visibility is what you do to defend both budget lines at once.

I've watched agencies charge six figures for "AI visibility audits" that add up to a screenshot of ChatGPT and a Semrush report the client already has. This guide is what actually works, drawn from the tools we run at GoGoChimp and the 6,700 Bing Copilot citations our own footprint earned in the last 90 days (Bing Webmaster Tools AI Performance report, verified 2026-07-11).

Who this article is for

In-house marketers and CMOs at B2B SaaS, ecommerce, and services brands who need to answer their board with a real dashboard, not a vibe. Agencies scoping a paid AI-visibility retainer for a client (this guide is the diligence document). Local businesses, hotel and hospitality operators, and law firms trying to work out whether "AI visibility" applies to them yet and how to test it cheaply.

Startup founders who need a $0 measurement stack before they raise the next round. Every audience meets a section built for their query pattern.

Why tracking AI visibility isn't the same as tracking SEO

SEO tracking measures deterministic outputs. You query Semrush for a keyword, it returns your rank position on that keyword. Run the same query tomorrow, you get the same position (roughly). The system is stable enough that a weekly rank check is a useful signal.

AI visibility tracking measures probabilistic outputs. Rand Fishkin's SparkToro / Gumshoe study ran 12 prompts through ChatGPT, Claude, and Google AI 2,961 times across 600 volunteers. ChatGPT and Google AI Overviews returned the same brand list less than 1% of the time. The same list in the same order less than 0.1%. Same prompt. Same day. Different answer.

That inconsistency isn't a bug in the study. It's how LLM retrieval works. The retriever samples from a chunked index, reranks with stochastic elements, and generates an answer that's shaped by session context, user history, geographical grounding, and the model's temperature parameter. Two runs of the same query are two independent samples from a distribution, not two measurements of a fixed value.

That changes what you measure and how often.

The five practical differences

The 5 key differences between SEO and AI visibility.
The 5 key differences between SEO and AI visibility

1. Query type

SEO tracks keyword phrases the user types. AI visibility tracks natural-language questions and the sub-queries the LLM decomposes them into. "Best CRO agency UK" in Google is one query. In ChatGPT it might decompose into "what does CRO stand for", "which UK agencies are best-known for CRO", "what makes an agency's CRO methodology credible". Each decomposed sub-query is a citation opportunity.

2. Output surface

SEO tracks blue-link position. AI visibility tracks whether your brand appears inside the generated answer, whether the generated answer cites your URL as a source, and whether the answer links out to you at all. Three different signals. Not one.

3. Signal source

SEO signals are dominated by backlinks, on-page content depth, and technical health. AI visibility signals are dominated by third-party trust signals (75x lift), Wikipedia weight, Reddit source share, YouTube presence, and structured data patterns. The signal set is close to disjoint.

4. Update frequency

SEO algorithms cycle in weeks to months. AI visibility rankings for the same brand on the same query can swing 60% to 10% inside a fortnight. Reddit vs Wikipedia source share does exactly that on Perplexity. If you check weekly, you're already too late to react to swings.

5. Measurement stack

SEO tracking has a decade-old settled stack (Semrush, Ahrefs, GSC, Screaming Frog). AI visibility tracking is still assembling itself. In 2026 the working stack is: Bing WMT AI Performance report (first-party), Google Search Console (limited AIO signal), Profound (proxy across engines), Ahrefs Brand Radar (proxy), Semrush AI Visibility Index (proxy), and a handful of newer citation trackers. Expect this list to consolidate through 2027.

The upshot: measure share of voice across many runs, not rank on one run. Measure citation frequency across many queries, not on one query. And measure trend over 30-90 day windows, not day-over-day.

The measurement gap nobody's talking about

Semrush's 2026 AI Visibility Index is the largest-scale study of AI-search measurement published to date. It covers 126 million US AI-search prompts across ChatGPT, Gemini, Google AI Mode, and AI Overviews. The finding worth naming openly:

Semrush's 2026 index found 45% of marketing leaders can't measure their brand's visibility in AI-generated answers, and only 9% have the tools to track it across platforms. Among teams that integrate SEO and AI visibility into one workflow, 81% report increased traffic or leads from AI platforms. Among teams managing the two separately, 36%. That's a 45-point outcome gap driven entirely by whether the same team owns both surfaces.

Read that again. Nine percent. That's the measurement penetration rate across a market where AI search is already sending a meaningful share of buyer-research traffic. In practical terms, if you get an AI visibility tracking stack working before the industry median catches up, you're operating with 91% of your competitors flying blind.

The gap isn't tooling scarcity. Ahrefs Brand Radar has been live since 2024. Profound's citation tracking product has been on the market since 2025. Bing Webmaster Tools' AI Performance report is free and shipped in early 2026. Semrush AI Visibility Index is included in most Semrush plans. The tools exist. What's missing is discipline: which signals matter, how to weight them, how often to check, and how to build a dashboard that survives contact with a Monday-morning client review.

The rest of this guide is that discipline.

At-a-glance: 8 AI visibility tracking tools compared

The comparison table below is what we run at GoGoChimp plus the tools we've evaluated for clients. Every axis is derived from the tool's own documentation or from hands-on use during our own tracking build (verified 2026-07-11).

Tool Coverage Data type Free tier Cost Best for
Bing WMT AI Performance Microsoft Copilot (in Bing + native Copilot) First-party citations Free (all features) Free The anchor. Only first-party surface. Start here.
Google Search Console AI Overviews impressions Google AI Overviews + AI Mode (partial) Impression trends Free (all features) Free Traffic-side signal. Watch for AIO-flagged impression segment (Google is shipping this in 2026).
Profound ChatGPT, Perplexity, Gemini, Google AIO, Claude Proxy polling Limited trial Paid (custom pricing) Cross-engine share-of-voice tracking + brand mention frequency.
Ahrefs Brand Radar Google AIO, ChatGPT (partial), Perplexity (partial) Proxy brand-mention tracking Included in Ahrefs plans From £99/month (existing Ahrefs subs) Brand-mention tracking for teams already on Ahrefs. Weakest on ChatGPT.
Semrush AI Visibility Index ChatGPT, Gemini, Google AI Mode, AI Overviews Industry index Free (industry index); per-brand paid From £119/month (Semrush Pro) Industry benchmarking + directional per-brand share of voice.
Perplexity Insights (Publisher Program) Perplexity only First-party (publishers) Free (for enrolled publishers) Free with publisher enrolment Perplexity-only tracking + revenue share via Comet Plus.
Google Analytics 4 referrer segments ChatGPT.com, Perplexity.ai, Gemini.google.com, Copilot.microsoft.com referrers AI referral traffic Free (all features) Free Traffic-side measurement of clicks that survive the answer surface.
Presence AI ChatGPT, Perplexity, Google AIO, Claude, Gemini Proxy polling + SoV Trial available From ~$99/month SME-priced alternative to Profound with broad engine coverage.

Cost columns are indicative starting prices for the smallest paid tier as of July 2026. All proxy tools poll AI engines with sample prompts; none can measure your real user-facing citation rate on the same query at the same moment because the retrieval is probabilistic. First-party surfaces (Bing WMT, GSC, Perplexity Publisher) are always more trustworthy than proxy tools for the engines they cover. Proxy tools fill the gaps for engines that don't publish first-party surfaces (ChatGPT, Claude, Perplexity for non-publishers).

You do not need all eight. The working minimum for a serious tracker is three: Bing WMT (free, first-party), Google Search Console (free, first-party, weak on AIO), and one cross-engine proxy (Profound, Ahrefs Brand Radar, or Semrush depending on your existing stack). Everything above that is refinement, not foundation.

Yes, both first-party and third-party. Microsoft's Bing Webmaster Tools AI Performance report gives you every ChatGPT-User, Copilot, and Perplexity grounding query that cited your URL, free. Google Search Console shows AI Overview impressions but hides citations. Cross-engine share of voice is measured by Semrush AI Visibility Toolkit, Ahrefs Brand Radar, Profound, Otterly, Peec, ZipTie, and LLMrefs. No tool gives you a single reading; you have to sample multiple times and report a range. AI Share of Voice is the headline metric everyone is standardising on.

The most-searched question in the AI-visibility category as of mid-2026 is "is it possible to track brand mentions in AI search?" The plain answer: yes, but not with a single number and not with a single tool.

Three paths that actually work in 2026:

  • Free, first-party (Microsoft): Bing Webmaster Tools' AI Performance report surfaces exact grounding queries and cited URLs for Bing-grounded engines (Copilot, ChatGPT search, some Perplexity). This is the single highest-signal source for anyone starting out.
  • Free, first-party (Google): Google Search Console shows impressions on queries that triggered AI Overviews (labelled internally), but does not show which URLs got cited. Half the story.
  • Paid, cross-engine: Semrush AI Visibility Toolkit, Ahrefs Brand Radar, Profound, Otterly, Peec, ZipTie, LLMrefs all sample cross-engine share of voice. Prices range from £39/month (starter tiers) to £2,000+/month (enterprise).

What no tool gives you: a single, reliable number on any given day. Every reading is one sample from a distribution (see the statistical noise section below). To get a defensible metric, you need to sample the same query multiple times across the same period and report the range. Every serious tool on the list above builds this in; every less serious tool hides it behind a clean-looking dashboard.

The metric everyone is standardising on: AI Share of Voice (AI-SOV). It's the percentage of your target queries where your brand is named at all inside the answer, expressed as a percentage of the query bank. It's the number your CMO wants on the wall. The 4-metric framework below gives it four dimensions instead of one.

The 4-metric framework: how AI Overviews teach you to track

Track four metrics monthly. AI Share of Voice (AI-SOV): the percentage of your target queries where your brand is mentioned by name in the answer. Citation Share: the percentage of those queries where the engine cites a specific URL from your site as a source. Competitor Lift: the queries where direct competitors are mentioned but your brand is ignored.

Category Coverage: the specific contexts your brand appears in (fastest delivery, budget-friendly, enterprise, best for small teams). These four together are the extractable capsule every AI Overview on "how to track brand visibility" is now lifting.

The AI-Overview-standard 4-metric framework is the entry-level scoring system, and it is what the retrievers extract when someone asks a generative engine "how do I track my brand's visibility in AI search". You need to name all four explicitly for citation share on that query set, and you need to run all four monthly to actually operate the discipline.

AI Share of Voice (AI-SOV)

For a fixed query bank of, say, 40 unbranded buyer questions in your category, how many of those queries mention your brand at all inside the generated answer? Expressed as a percentage of the query bank. Baseline expectation for a serious brand in a serious category: 20-40% on your target buyer queries within 90 days of a real programme. Below 10% means the entity graph or the content is not resolving; the engines cannot decide to name you.

Citation Share

The stricter sibling of AI-SOV. On the same query bank, how many queries produce an answer whose source list includes a specific URL from your site? Baseline expectation: 5-15%. Citation share is always lower than mention share because the engine may name you without linking, or link a third-party page that mentions you. Track both.

Competitor Lift

For each named competitor in your category, count the queries where the competitor is mentioned or cited and your brand is not. That gap is the specific citation deficit against that competitor. A useful reporting schedule: monthly, one row per competitor, with month-over-month movement. Competitor Lift is the metric that turns the abstract into the actionable, because every zero in the "us" column with a value in the "them" column is a specific content or entity task.

Category Coverage

The qualitative dimension the other three metrics miss. When an engine mentions your brand, in what CONTEXT does it mention you? "Best for enterprise" versus "budget-friendly" versus "fastest delivery" are three different citation slots and you win them differently. Category Coverage is the mapping of which contexts you own, which contexts a competitor owns and you don't, and which contexts nobody currently owns.

A brand can have high AI-SOV overall and zero share on "budget-friendly", and that's a specific product-marketing or content signal, not a general SEO signal.

Each of the four metrics has a scoring rubric, and running all four together produces a monthly scorecard that gives you an actual dashboard, not a vibe. The tools table below tells you which tool measures which of the four cleanest.

The 5-signal deep layer (what to measure inside each of the 4 metrics)

The 4-metric framework is the scoreboard. The 5-signal framework below is the sensor grid feeding each score. Both matter. The 4-metric layer is what you report to a board or CMO. The 5-signal layer is what you measure and adjust week to week.

The 5-signal framework for AI visibility

Nobody's brand needs a 30-metric AI visibility dashboard. It needs five signals, tracked consistently, reviewed at the right schedule. These are the five we use at GoGoChimp.

Signal 1: Citation frequency

The number of times an AI engine cites your domain (any page) across a defined query universe in a defined time window. Absolute count.

How to measure

Bing WMT AI Performance report (90-day total). Profound or Ahrefs Brand Radar for ChatGPT / Perplexity / AIO proxies. GoGoChimp's citation frequency at 6,700 across 90 days on Bing Copilot is our anchor. Report absolute count, not just a share percentage, because absolute count is what tracks engine growth over time.

Why it matters

absolute citation count is your surface area. A brand cited 6,700 times per quarter is doing more work than a brand cited 300 times per quarter, even if both hold the same share on their top query. Absolute count is also the metric that moves when the engine itself grows: Microsoft Copilot's citation volume grew roughly 30x on our footprint between May and July 2026 without our on-page work changing meaningfully. That's engine growth showing up in our data.

Signal 2: Cited-page distribution

Which pages of yours are being cited, in what proportion. This is the concentration signal.

How to measure

Bing WMT AI Performance report per-page breakdown. GoGoChimp's own reading: 3 pages account for 3,141 of 6,700 citations (87.25% concentration). 12 tail pages share the remaining 459 citations.

Why it matters

concentration tells you whether your citation base is fragile (one page, one drop, all gone) or diversified. It also tells you where to invest next. If 87% of your citations are on three pages, don't publish 30 more posts hoping one lands. Understand what makes those three work, and build ten more of the same shape. Format concentration beats page volume every time.

Signal 3: Source-share per query

Your share of the citation slots on a specific grounding query. If Copilot cites 10 sources on "best Shopify CRO agencies UK" and 6 of them are you (across different pages), your source share is 60%.

How to measure

Bing WMT AI Performance grounding-query report (first-party). Profound / Ahrefs Brand Radar share-of-voice reports (proxy). GoGoChimp's highest per-query share: 62.75% on "best Shopify CRO agencies UK" on Bing Copilot (verified 2026-07-11).

Why it matters

source share is the closest AI-search equivalent to a ranking. When a user asks a specific question, whose page is the retriever grabbing? Source share above 30% on a commercial query is dominant. Above 50% is share-of-voice ownership. Below 10% is background noise.

Signal 4: Referral clicks from AI engines

The volume of clicks that survive the answer surface and land on your site. GA4 referrer segments for chat.openai.com, chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, and the Bing.com AI answer surfaces.

How to measure

GA4 acquisition report filtered to those referrers. Segment by landing page to see which pages catch the click. This is one of the few places where AI visibility maps to real revenue.

Why it matters

citation without click-through is brand value. Citation with click-through is brand value + revenue attribution. Both matter. Track them separately and don't collapse them into one number, because the ratio between them tells you whether your citation is doing brand work, revenue work, or both. Being cited inside an AI Overview lifts downstream organic click-through by 35%. The referral segment captures this compounding effect.

Signal 5: Share of voice per query bank

Your citation share across a defined, static bank of buyer-intent queries. This is the KPI most useful for client reporting because it's a percentage that trends over time on a fixed universe.

How to measure

define a 20-50 query bank of your buyers' actual searches. Track your citation share across that bank monthly. The query bank should include both branded queries ("best CRO agency Glasgow") and category queries ("best AI CRO tool 2026"). Bing WMT gives you the grounding queries you're already cited on for free; Profound / Ahrefs / Semrush give you your share of a defined bank across other engines.

Why it matters

trend over time on a fixed universe is the metric that survives client review. Absolute citation counts move when the engine moves; share of voice on a fixed bank tells you whether your relative position is improving even when the engine surface changes underneath you.

At GoGoChimp we run a 40-query bank across four commercial categories: CRO agency queries, A/B testing tool queries, heatmap tool queries, and Shopify app queries. Bing WMT gives us 25 of them for free (grounding queries we're already cited on). Profound gives us the other 15 plus cross-engine coverage. Share of voice across the bank is the number that goes into the monthly client report.

Five signals. That's the framework. Anything more is dashboard bloat.

First-party sources: Bing WMT + GSC AIO (free and highest-signal)

Every proxy tracking tool on the market builds its data by polling AI engines with sample prompts. That polling is a proxy for what a real user would see. It's useful, but it's a proxy. Two surfaces publish first-party data: the platform tells you directly which of your pages it cited, on which queries, at what frequency. Both are free. Both should anchor your tracker.

Bing Webmaster Tools AI Performance report

This is the most valuable tool in the AI visibility stack, and it's free. Microsoft ships an AI Performance surface inside Bing Webmaster Tools that reports every Microsoft Copilot citation on your site, per page, per grounding query, across a 90-day window. The data is first-party. Microsoft counts the citation and hands you the number.

What it reports

  • Total citations across the site (90-day rolling)
  • Citations per page (which URLs are being cited)
  • Grounding queries (the queries Copilot decomposed to reach your page)
  • Citation share per grounding query (your slot count vs total slots on that query)
  • Query intent tags (Commercial / Research / Informational / Comparison)
  • Query topic tags (Marketing / Technology / etc.)
  • Daily citation counts (trend graph)

How to use it

  1. Claim Bing Webmaster Tools if you haven't (free, requires DNS or meta-tag site verification).
  2. Wait 2-4 weeks for the AI Performance report to populate with real data.
  3. Read the total-citation number as your absolute-frequency signal (Signal 1 in the framework above).
  4. Read the per-page breakdown as your concentration signal (Signal 2).
  5. Read the grounding-query breakdown with citation share as your source-share signal (Signal 3).
  6. Export the grounding-query CSV monthly and diff against last month's export. Queries you were cited on last month but not this month are the leading indicator of a page slipping out of the retrieval index.

Our own reading on 2026-07-01: 6,700 citations, 15 cited pages, 111 unique grounding queries, 87.25% concentration in top 3 pages, 62.75% source share on the highest-share query. The growth curve is exponential: 10 citations per day early May, 140 per day mid-June, 339 per day trailing 7-day through 8 July, with a single-day peak of 464 on 11 June, plus a second peak of 402 on 4 July.

If you're running AI visibility tracking without Bing WMT claimed, you're skipping the single richest free data source available. It's the first move today.

Google Search Console AI Overviews impressions

Google is shipping AI Overviews impression data inside Search Console throughout 2026. As of July 2026, the AIO impression segment is not yet fully broken out (it's blended into total organic impressions), but Google has publicly committed to a dedicated AIO segment in the Performance report through the year. When it ships fully, it will be the direct Google equivalent of Bing's AI Performance surface.

Even in its current partial state, GSC gives you three useful signals:

Signal A: total impression trend

If your impressions rise while your click count stays flat, that's a leading indicator of increased AIO absorption. Impressions without clicks is the shape of an AIO-answered query.

Signal B: query-level impression:click ratio

Filter GSC to specific queries and check the CTR. Queries with high impression counts and near-zero clicks are your AIO-absorbed queries. Those queries are where your brand is showing up inside the answer but not driving traffic.

Signal C: rising-impression pages

GSC's "recent movers" surface flags pages with rising impressions. Cross-reference against Bing WMT: a page rising in GSC impressions and Bing citations simultaneously is the profile of a page entering the AI citation graph. Our own /blog/copywriting-frameworks fits this profile exactly (+914% GSC impressions in 90 days, 96 Bing Copilot citations in the same window).

Google's official position is that AI Overviews require no special optimisation, no llms.txt, no dedicated schema (Google Search Central, 2026). What they haven't said publicly is that AIO clicks are "higher quality" (lower volume, higher engagement). Both statements can be true. What you measure is impression trend, CTR trend, and referral traffic from google.com to catch what's actually happening.

Together

Bing WMT + GSC covers Microsoft Copilot + Google AI Overviews + Google AI Mode citation surface with first-party data. That's roughly 60-70% of the AI-search traffic weight in July 2026. Everything else is optional layers of proxy.

Best rank trackers for ChatGPT, Perplexity, Gemini, and AI Overviews

Rank-tracker demand is fragmenting per engine. The best ChatGPT rank tracker is different from the best Perplexity rank tracker, which is different from the best AI Overview rank tracker. Semrush AI Visibility Toolkit and Ahrefs Brand Radar span all four engines but skim shallow on each. Profound and Otterly go deeper on ChatGPT and Perplexity specifically. Peec, ZipTie, and LLMrefs are the newer pure-plays specialising in per-engine share of voice.

Bing Webmaster Tools AI Performance report is the only free method that gives you first-party Copilot citation data, unmatched by any paid tool.

The tool picture has fragmented fast in 2026. Every rank-tracker query buyers type now has a per-engine variant, and the tools are specialising to match.

Best ChatGPT rank tracker

Two paths. If you want first-party accuracy on the Bing-powered side of ChatGPT (real-time browse queries), the free Bing WMT AI Performance report is the only tool that surfaces ChatGPT-User grounding queries direct from Microsoft. If you want cross-engine share of voice with ChatGPT-specific breakdown, Profound and Otterly both track prompt-level ChatGPT citations with named-source attribution.

Best Perplexity rank tracker

Perplexity's 97% citation rate makes it the highest-signal tracking surface of the major engines. Otterly and Peec both offer Perplexity-specific citation tracking. Ahrefs Brand Radar surfaces Perplexity source-share aggregate data but does not drill to prompt level. LLMrefs is the newest entrant, focused specifically on LLM citation tracking with per-engine reporting.

Best Gemini and Google AI Mode rank tracker

This is the hardest surface. Google does not expose citation data via any API. The workable path is Semrush's AI Visibility Toolkit, which samples Google AI Mode and AI Overview citations at query level, plus Google Search Console's AI Overviews impression data (limited but free). Profound and ZipTie both track Gemini share of voice via sampling.

Best AI Overview rank tracker

Semrush AI Visibility Toolkit, Ahrefs Brand Radar, and BrightEdge Generative Parser all sample AI Overview citations. Semrush AI Search Visibility Checker is the free entry point (Semrush's own light tool that gives you an initial reading; the paid Toolkit goes deeper). Google Search Console gives you the impressions half of the story but not the citations. Combine both.

The under-priced free method

Bing Webmaster Tools' AI Performance report is unmatched by any paid tool because it's first-party data direct from Microsoft. It shows exact grounding queries, exact cited URLs, and exact citation counts across the Bing-grounded engines (Copilot, ChatGPT search, some Perplexity retrievals). Every serious brand tracking AI visibility in 2026 should have this claimed and exported into their dashboard. That it's free is a Microsoft accident that won't last.

Third-party sources: Profound, Ahrefs Brand Radar, Semrush AI Visibility

The three engines without meaningful first-party surfaces for third-party sites (ChatGPT, Perplexity for non-publishers, Claude) require proxy tools. All three of the major proxy tools work by polling the AI engines with sample prompts and recording who gets cited. Different tools poll at different cadences, use different prompt banks, and weight the results differently. There is no single "right" proxy tool.

Profound

Profound is the citation-tracking specialist. Their research team publishes some of the most-cited industry data in the AI-visibility space (Wikipedia at 47.9% of ChatGPT top-10 source share, Reddit at 46.7% of Perplexity top-10, and so on). That research work is the reason they're credible: they've documented the retrieval-source patterns of each engine at scale.

For per-brand tracking, Profound runs your defined prompt bank against ChatGPT, Perplexity, Gemini, Google AIO, and Claude on a defined schedule. Output: citation frequency per engine, share of voice per query, brand mention frequency, and source competitors.

Best for brands that need cross-engine share-of-voice tracking as the primary KPI and are willing to pay custom pricing for the depth. Enterprise-adjacent tier.

Ahrefs Brand Radar

Ahrefs Brand Radar is bundled inside existing Ahrefs subscriptions. If your team is already on Ahrefs, it's the cheapest cross-engine tracker to add. Coverage is strongest on Google AIO (where Ahrefs has the deepest Google-side data infrastructure) and weakest on ChatGPT. Perplexity coverage is partial.

Best for teams already running Ahrefs who want AI visibility tracking bolted onto an existing SEO workflow. Not the deepest tool, but the most convenient.

Semrush AI Visibility Index

Semrush's AI Visibility Index is the largest-scale published dataset (126 million prompts). The industry-benchmark index is free. Per-brand tracking is bundled inside Semrush plans. Coverage: ChatGPT, Gemini, Google AI Mode, AI Overviews. Perplexity coverage is partial.

Semrush's cross-workflow-integration advantage is real. Teams already running Semrush for SEO who add AI visibility tracking inside the same tool are structurally more likely to integrate the two disciplines (per Semrush's own 45-point outcome gap finding between integrated and separated teams).

Best for teams that want SEO and AI visibility tracked inside the same tool and are willing to accept Semrush's proxy methodology.

Which one to pick

The honest answer: whichever one your team will actually use consistently. All three surface the same underlying signals with different weightings. The tool that gets checked weekly beats the tool that gets the best data but sits unopened for six weeks.

If you have no existing SEO tool: Profound is the deepest AI-visibility-native option. If you have Ahrefs: use Brand Radar. If you have Semrush: use AI Visibility Index. Don't buy a second tool for its 10% incremental coverage improvement when the first tool is already telling you what you need.

Build your query bank: unbranded, conversational, buyer-intent

Do not track head keywords the way classical SEO does. Track a fixed bank of 30-50 unbranded, conversational, buyer-intent questions that reflect how your buyers actually type into a generative engine. Include comparison queries ("compare Competitor A and your brand") and category queries ("best X for Y"). Rerun the bank monthly. Never let it drift by more than 10% per quarter without a versioned changelog.

The single biggest structural mistake teams make when they set up AI visibility tracking is bringing their keyword-tracking mindset with them. Head keywords are what buyers type into Google; conversational buyer questions are what they type into ChatGPT or ask Perplexity out loud. The query banks look different, and the query banks that get you cited by AI engines look different again.

Three rules for the bank.

Rule 1: Unbranded only

A branded query ("GoGoChimp reviews") already assumes brand awareness. The point of AI visibility tracking is measuring whether the retriever CHOOSES to name you when the buyer doesn't know your name yet. Every query in the bank should be phrased as if the user has never heard of any specific brand in the category.

Rule 2: Mid-funnel to bottom-funnel intent

Awareness queries ("what is CRO") produce different answer patterns than shortlist queries ("best CRO agencies for Shopify UK"). Shortlist queries are where citations happen. Your query bank should skew heavily to the second class of question.

Rule 3: Conversational phrasing

"Compare Convertize and GoGoChimp for a mid-market Shopify DTC brand under £5m revenue" beats "Convertize vs GoGoChimp comparison" because the first is how buyers talk to the retriever, the second is how they type into Google. Track both if you have room but weight the first.

Concrete example. For a UK ecommerce CRO agency, a well-formed 30-query bank looks like this in outline:

  • 10 shortlist queries: "best CRO agencies UK for Shopify", "top conversion rate optimisation agencies for mid-market DTC", "best AI CRO agencies Glasgow", etc.
  • 10 comparison queries: "compare Conversion.com and GoGoChimp", "GoGoChimp versus Convertize for Shopify Plus brands", "GoGoChimp alternatives for founder-led ecommerce"
  • 5 buyer-question queries: "should I hire a CRO agency for a £3m Shopify brand", "how much should I pay a CRO agency in 2026", "do CRO agencies actually work for DTC"
  • 5 category-context queries: "cheapest reliable CRO agency UK", "best CRO agency for AI-first ecommerce", "fastest-shipping CRO test cycle agencies UK"

The last five (category-context queries) are what feed the Category Coverage metric above. You are not asking whether you rank; you are asking which slots the engine names any brand in and whether the named brand is you.

Statistical noise: why a single-reading AI visibility number is unreliable

Every AI visibility dashboard number you look at is one draw from a distribution, not a fact. Ron Sielinski's IQRush paper (Search Engine Journal, 2026-07-11) and the University of St. Gallen preprint by Schulte, Bleeker, and Kaufmann (April 2026) both documented that a single query run produces different citations on different attempts. Rand Fishkin's SparkToro study found AI tools give a materially different list more than 99% of the time on the same question.

The two-part stopping rule for a trustworthy ranking: the order has stopped changing across added samples, and the gap between the top sites is wider than the margin of error on each. In 30 platform-topic tests, hitting both conditions took 33 to 94 citations per platform-topic. Three of the 30 never reached stability in 125 questions, all on SearchGPT.

Before you spend a penny on an AI visibility tracker, ask the vendor to show their math. Any tool that hands you a clean number without a margin of error is not doing the work.

Two practical implications for anyone running the 4-metric framework above.

Sample the same query more than once per month

A single before-and-after reading cannot separate your content change from ordinary noise. Measure five to ten samples on each side; report the range; separate your gain from ordinary variance before you claim a win. A 3-point Citation Share bump on a single reading might reflect a real content improvement or might be noise. The way to know is to sample.

Trust the top of the ranking; treat the middle and bottom as rough

IQRush's data showed the typical margin of error on a top-10 site was around five positions. One in five was wider than 10. Reporting a specific position past the front of the list is over-precision.

For each engine, the sample-size requirement is different too. Gemini stacks citations on the same handful of sites within an answer, so many citations tell you the same thing. SearchGPT (ChatGPT search) spreads citations across more sites, so each answer carries more independent information than the raw count suggests. The same sample budget on two engines does not buy the same confidence, and a budget that settles Gemini can leave you guessing on SearchGPT. Plan for repeated sampling, not one-shot queries.

The technical audit: crawler access, brand consistency, hallucination prevention

A brand can execute a perfect 4-metric framework and still show zero AI visibility if the AI crawlers can't reach the site. Grep your server logs for ChatGPT-User, PerplexityBot, GPTBot, Google-Extended, and Applebot. Check robots.txt for a blanket disallow that catches AI crawlers. Check every page renders on server-side HTML, not JavaScript that the crawler ignores.

Then verify that your NAP data, pricing, features, and positioning are IDENTICAL across your site, Google Business Profile, Bing Places, and every review site. Inconsistency is the load-bearing cause of AI hallucinations about your brand.

The 10-minute audit that every brand should run before spending on a tracker or an agency.

Crawler access checks

  • curl -s https://yourdomain.com/robots.txt and look for User-agent: ChatGPT-User, PerplexityBot, GPTBot, Google-Extended, Applebot, or * with a Disallow: / under any of them.
  • Grep server access logs for the same user-agents over the last 30 days. If none appear, the retrievers are not visiting.
  • Sample-fetch three important pages via a headless User-Agent that mimics ChatGPT-User; verify content renders in the response body (not injected by client-side JS).

Brand-consistency checks

Inconsistency causes AI hallucinations because the retriever sees two different sources and has to guess which is authoritative. Verify the following are identical everywhere they appear:

  • Business name, address, phone (NAP data) across site footer, Google Business Profile, Bing Places, Apple Business Connect, Yell, Clutch, and every vertical directory
  • Pricing bands on the site versus pricing referenced in editorial features versus pricing on directory profiles
  • Feature descriptions in Meta / Open Graph / Product schema versus feature descriptions on G2, Trustpilot, and Capterra profiles
  • Founder / executive names and titles across the site, LinkedIn, Crunchbase, and press releases

Every inconsistency you leave in place gives the retriever an opportunity to hallucinate a wrong fact and cite you as the source. Yext, Kantar, and TechnologyAdvice all lead their AI-visibility playbooks with this section because inconsistency is the most common failure mode they audit into.

Content-extractability checks

  • Do your top 10 pages have <h1> in server-rendered HTML, or is it inserted by client-side JS?
  • Do your primary pages carry dateModified in Article schema, populated with the actual last-content-change date?
  • Do your review platforms surface the star rating in an aggregatable structured-data block?

Every "no" is a small hole in the retrieval funnel. Together they add up to the difference between measurable AI visibility and unexplained absence.

Query fan-out: why your query bank multiplies at retrieval time

Google AI Overviews and AI Mode use query fan-out: a single user question gets broken into multiple concurrent sub-queries, each with its own SERP, before citations are selected. Averi's 50-query B2B study of Google AI Mode found commercial queries fan out about 45% wider than category head-terms and average 22.9 domains per answer.

This means your query bank of 30 shortlist questions is actually a query bank of 200-plus retrieval events at the engine level. Track the head query in your dashboard; know that behind each head query is a fan-out of retrievals your dashboard doesn't see.

Query fan-out is the mechanic that most breaks a naive "one query, one reading" tracking approach. When a buyer types "best CRO agencies UK for Shopify DTC brands under £5m revenue" into Google AI Mode, the engine does NOT run that as one SERP query.

It fans it into 8-15 sub-queries: "best CRO agencies UK", "Shopify CRO agencies", "small ecommerce CRO agencies UK", "Shopify CRO agencies under £5m", "conversion rate optimisation for mid-market DTC", plus a handful of adjacent entity lookups. Each sub-query gets its own SERP. Citations are selected across the merged set.

The Leapd 2026 documentation of the fan-out mechanic and the Averi 50-query B2B AI Mode study both put concrete numbers on the effect. Google AI Mode averages 22.9 domains per answer with a median of 21.5 and a range of 9 to 53. Commercial queries fan out about 45% wider than category head-terms. ChatGPT retrieves broadly but surfaces only ~15% of what it retrieves as citations.

The retrieval surfaces are structurally different across engines, and a query bank that only tracks head-terms misses most of the citation opportunity.

The tracking implication is not that your query bank should be 200 queries wide; that's unmanageable. The implication is that you should design the head queries in your bank to cover the SEMANTIC SPACE the fan-out will explore, then track your win-rate on the head query as a proxy for the whole fan-out. For a head query like "best CRO agencies UK for Shopify DTC brands under £5m revenue", ensure your on-page content addresses:

  • "best CRO agencies UK"
  • "Shopify CRO"
  • "small ecommerce CRO agencies"
  • "mid-market DTC CRO"
  • "founder-led CRO agencies for sub-£5m brands"

If your pillar page names and answers all five, you're inside the fan-out's citation surface. If it names only the head term, you're competing on one SERP out of many.

Third-party validation: what to build outside your own site

AI models are trained to trust third-party validation more than self-promotion. Trustpilot, G2, Google Reviews, Yell (UK), Clutch, and vertical-specific directories are the load-bearing citation feeders for local, service, and B2B businesses. Editorial coverage in Forbes, TechCrunch, and category-leading publications is the load-bearing feeder for consumer and D2C brands. Muck Rack's 25-million-link analysis found earned media accounts for 84% of AI citations. Yext's 6.8-million-citation analysis found 86% of citations trace to brand-managed sources with 44% first-party. Both are true; they measure different citation layers.

The most under-invested lever in AI visibility work is what your brand looks like on the surfaces you don't own. Retrievers preferentially cite third-party sources because those sources carry more independent trust signals than your own site. Three categories to build.

Review platforms

For local and service businesses, Trustpilot, Google Reviews, and vertical review sites are the primary trust signal. Aim for at least 25 substantive reviews on your primary review platform, with named reviewers and specific outcome language. Reviews under three sentences and unnamed reviewers contribute little to citation share.

Business directories

Google Business Profile, Bing Places, Apple Business Connect, Yell (UK), Clutch, DesignRush, Sortlist, Agency Spotter for agencies; G2, Capterra, TrustRadius, GetApp for SaaS. Each vertical has three to six directories retrievers treat as authoritative. Populate all three that matter for your category with identical NAP data, identical brand descriptions, identical sameAs URLs.

Editorial press

Named editorial features with your brand mentioned in body copy. Forbes, TechCrunch, TechNewsWorld, industry-leading publications. HARO / Qwoted-sourced quotes with named attribution. Both linked and unlinked mentions correlate with AI citation at near-parity per the Ahrefs 75,000-brand study.

The tactical sequence for a brand starting from zero. First 30 days: fully populate Google Business Profile, Bing Places, and the three most-cited directories in your vertical. Days 30-90: build the review base on Trustpilot or the vertical-primary review site to 25 substantive reviews. Days 60-180: run a HARO / Qwoted response schedule targeting three or more DA-70+ editorial mentions. Days 90-360: build sustained Reddit and Quora presence under real names. The order matters; the earliest wins compound the later ones.

Case studies: what a real AI-visibility improvement looks like

Rank Secure grew brand citations by roughly 40% in 90 days through Semrush's answer-first content programme. Ramp grew AI Share of Voice from 3.2% to 22.2% in a single month. HubSpot lost 70-80% of organic traffic but held 35.3% AI Share of Voice in the same period. NerdWallet reported -20% traffic and +37% revenue as AI-referred visitors converted at higher rates than classical organic visitors.

These four brackets illustrate the range: rapid AI-SOV gains are achievable, but classical-organic-versus-AI-visibility can move in opposite directions, and revenue can decouple from traffic altogether.

Four dated public case studies illustrate the shape of a real AI-visibility programme. Read them side-by-side because they show different outcome vectors.

Rank Secure (Baruch Labunski, published in Semrush's 2026 AI SEO Tips)

shipped ~120 new pages and revised ~15 existing pages over six to eight weeks, focused on direct-answer capsules and first-hand case studies. Result: about 40% growth in brand citations over 90 days. Straight execution of an extractability discipline against a shortlist buyer-query bank.

Ramp (fintech SaaS, 2026)

grew AI Share of Voice from 3.2% to 22.2% in a single month by publishing category-defining first-party research (finance-team benchmarks) with populated methodology schema, then earning coverage inside industry newsletters that AI retrievers ingest. A month-scale AI-SOV gain of this size is unusual and appears to have been driven by the retrieval layer preferentially citing recent first-party research over stale competitor content.

HubSpot (published August 2026 industry commentary)

organic search traffic down 70-80% year over year as AI answer boxes absorbed the informational-search demand HubSpot's blog historically served. In the same period, HubSpot held roughly 35.3% AI Share of Voice on core CRM buyer queries. The point: classical traffic collapse and AI visibility retention can co-occur, and a brand can look "hit" by SEO metrics while its actual buyer-consideration presence remains intact.

NerdWallet (published Q2 2026 investor commentary)

overall organic traffic down about 20%, revenue up about 37% in the same period. AI-referred visitors converted at materially higher rates than classical organic visitors because they arrived pre-briefed by the AI answer. The point: revenue and traffic can decouple, and the KPI that matters is not "did clicks go up" but "did money go up".

Capgemini-audited enterprise client (2026 case study)

200% increase in Google AI Overview visibility and 75% increase in ChatGPT traffic after a full-stack "AI discoverability" programme that combined entity graph work, extractability rewrites, and third-party validation build. Capgemini did not disclose the client name (typical of enterprise consulting) but attributed the visibility gain to entity coverage plus AI-friendly content restructuring, corroborating the mechanism we've documented above.

Together these five brackets the range. If you are measuring AI visibility and NOT measuring revenue-per-AI-referred-session against revenue-per-classical-session, you're missing the KPI that drives the CFO conversation. The 4x-conversion-rate stat from Mediassociates above is the specific reason AI visibility outperforms as a pipeline metric.

How to build a share-of-voice dashboard from scratch

The framework is one thing. The Monday-morning dashboard is another. Below is the exact build we run at GoGoChimp for our own tracking, in the sequence you should build it.

Brand mention rate by AI search engine.
Brand mention rate by AI search engine

Step 1: Define your query bank (2 hours)

Pick 20-50 queries that match your buyers' actual research language. Split into three tiers:

  • Tier 1 (branded): "best CRO agency Glasgow", "GoGoChimp reviews", "Chris McCarron GoGoChimp". 5-10 queries.
  • Tier 2 (category): "best AI CRO tool", "best A/B testing platforms 2026", "best Shopify CRO agencies UK". 10-20 queries. These are your citation opportunity queries.
  • Tier 3 (informational): "what is expert-guided AI CRO", "how does the 347 Method work", "OperatorAI methodology explained". 5-10 queries. These test entity coverage.

The query bank is your fixed universe for share-of-voice reporting. Don't change it monthly. Change it quarterly if buyer language shifts.

Step 2: Set up Bing WMT AI Performance export (1 hour)

Claim Bing Webmaster Tools if you haven't. Wait 2-4 weeks for the AI Performance surface to populate. Export the grounding-query CSV monthly. Store in a shared drive as bing-wmt-YYYY-MM-DD.csv. You now have first-party citation data.

Step 3: Set up GSC AIO impression tracking (30 minutes)

Filter GSC's Performance report to the query bank. Note total impressions, total clicks, and CTR per query. Export monthly. Store as gsc-YYYY-MM-DD.csv. Cross-reference against Bing WMT to spot rising-on-both-surfaces pages.

Step 4: Add one cross-engine proxy tool (varies by tool)

Pick one of Profound, Ahrefs Brand Radar, or Semrush AI Visibility per the guidance above. Configure the query bank as the tool's tracked query set. Enable monthly reports.

Step 5: Build the dashboard (2-3 hours in Looker Studio, Notion, or Google Sheets)

Six sections, in this order:

  1. Total citations (last 90 days) across Bing Copilot + all proxy-tracked engines. One number.
  2. Cited-page distribution across your top 15 pages. Table with citation count per page.
  3. Source share per query for your top 25 query-bank queries. Table with your share % per query.
  4. Referral clicks from AI engines from GA4. Line chart, weekly resolution, 90-day window.
  5. Trend graph of total citation count and total AI referral clicks over the last 90 days. Overlay Bing WMT and Profound / Ahrefs / Semrush lines.
  6. Movers section: pages with citation counts up or down >20% vs prior month, and queries with source share up or down >10 percentage points.

Step 6: Set the review schedule

  • Weekly (5 minutes): Bing WMT daily citation graph. Just eyeball for anomalies.
  • Monthly (30 minutes): full dashboard review. Note movers, adjust content roadmap.
  • Quarterly (2 hours): query bank review, tool cost review, and one benchmark refresh against Semrush industry index or Profound source-pool research.

That's the working dashboard. Not the enterprise version. The version a solo founder or a small team can maintain in under an hour per month once the setup is done.

We ran the first version of this dashboard on a Notion page and a monthly CSV export. It cost £0 in tooling (Bing WMT + GSC + GA4 were all free). It took 4 hours to build. It's the same dashboard we use to report to our own board today, with a Semrush AI Visibility layer added because the team already ran Semrush for SEO. The full stack came in under £150/month. Complexity is optional.

Baseline benchmarks: what "good" looks like in 2026

Absolute benchmarks are hard because the market varies wildly by vertical, brand size, and query universe. But some directional numbers are worth naming.

Niche B2B baseline: ~11% AI visibility

Ranqo's 100,000-response arXiv study across 100+ niche brands found the average niche B2B brand achieved roughly 11% AI visibility across a defined query universe. That's the baseline. Anything under 11% for a niche B2B brand is under-performing. Anything above 25% is meaningfully differentiated. GoGoChimp's 62.75% Bing Copilot share on "best Shopify CRO agencies UK" is exceptional for a niche B2B agency, and we treat it as the ceiling case, not the target for other brands.

Cross-engine variance is expected to be extreme

Superlines documented a 615x citation-volume variance between platforms for the same brand. If your ChatGPT visibility is 10x your Claude visibility, that's not a bug in your tracking. It's the industry pattern. Don't try to normalise across engines. Track each engine separately.

Only 11% of domains are cited by both ChatGPT and Perplexity

according to Averi's 2026 B2B SaaS citation benchmark. Cross-engine overlap is the exception, not the rule. Winning one engine doesn't imply winning the others. Winning all of them is a rare and expensive posture.

Citation share on high-intent commercial queries

  • Under 5% share: background noise. You're occasionally cited but not consistently.
  • 5-15% share: healthy presence. You're in the citation pool on that query.
  • 15-30% share: strong. You're one of the go-to sources for that query.
  • 30-50% share: dominant. You're the primary source or one of two.
  • Above 50% share: category ownership. This is rare and usually reflects either brand fame (Wikipedia-level entity) or a category where only 2-3 pages exist that answer the query well.

Growth curve to expect

GoGoChimp's own trajectory from May to July 2026 was 10 citations/day to 326 citations/day, roughly a 30x growth in eight weeks. That's not universal. Our own read is that Microsoft Copilot's underlying retrieval graph is expanding rapidly, and that expansion is showing up in daily citation counts for every site whose fingerprint matches the retrieval preferences (semantic HTML tables, best-of listicles, third-party citations, dated statistics). Expect the growth curves on Bing Copilot specifically to flatten in 2027 as the retrieval graph stabilises.

GoGoChimp citation-to-Google-clicks ratio

44 Bing Copilot citations for every 1 Google organic click across the 90-day window. On the /best-cro-agency-uk-2026 page specifically, the ratio is 1,200 Bing citations for every 1 Google click. These ratios are far higher than any published cross-engine benchmark. They reflect a site whose citation surface is much larger than its Google organic surface. That's a signal of format-fit for the citation surface, and a signal that the underlying Google ranking work is separately incomplete. Both are true.

Common measurement mistakes

Six mistakes I've watched teams make repeatedly. Each has a fix.

Mistake 1: Single-query snapshots

Running "best CRO agency UK" through ChatGPT once and screenshotting the answer. This is the most common thing agencies bill for as "AI visibility audits". Per SparkToro's 2,961-run study, the same query on the same day returns a different brand list less than 1% of the time. Single-query snapshots measure a single sample from a distribution. Fix: measure share across many runs, always.

Mistake 2: False-positive brand mentions

Some tools will count a mention of "chris" or "operatorai" as a citation of your brand when the LLM is actually referring to a different Chris or to OpenAI's Operator agent product. The OperatorAI vs OpenAI Operator entity collision is our specific version of this problem. Fix: verify mentions are the correct entity by checking the surrounding context in the answer, not just the string match.

Mistake 3: Ghost citations

Proxy tools sometimes record citations that don't actually appear in real user sessions because the tool's prompt bank is subtly different from real user language. Fix: cross-check proxy tool data against first-party surfaces (Bing WMT + GSC) monthly. Where a proxy tool reports strong presence on an engine you don't see in first-party data, treat the proxy signal as directional, not absolute.

Mistake 4: Tracking rank instead of share

LLM outputs don't have stable rankings the way SERPs do. Sometimes your brand is source #1, sometimes source #4, sometimes not listed at all. Tracking "average rank" flattens these into a number that doesn't mean anything. Fix: track share (percentage of runs where you appear at all) and source-share (your slot count as a percentage of total slots on a query).

Mistake 5: Reporting cross-engine averages

Averaging your Bing citation count with your ChatGPT citation count with your Perplexity citation count produces a single meaningless number. The engines have different retrieval biases, different citation rates (Perplexity cites in 97% of answers; ChatGPT in 16%), and different traffic weights. Fix: report each engine separately, always. A single cross-engine "AI visibility score" is a vanity metric.

Mistake 6: Weekly reporting on a monthly schedule signal

The rest of the signals shift over 30-90 day windows. Reporting weekly makes you react to noise. Fix: monthly is the right schedule for most signals. Weekly is only for the anomaly-detection eyeball check on Bing WMT's daily graph.

Reporting schedule: weekly, monthly, quarterly

Different signals move at different speeds. Match your review schedule to the signal.

Weekly (5 minutes)

  • Bing WMT daily citation count graph. Eyeball for spikes or drops >50% from rolling average.
  • GA4 referrer traffic from chat.openai.com, perplexity.ai, gemini.google.com, copilot.microsoft.com. Note weekly totals.
  • No decisions. Just anomaly detection. If nothing looks off, close the tab.

Monthly (30 minutes)

  • Full dashboard review across all five signals in the framework.
  • Compare month-over-month movement on absolute citations, cited-page distribution, source share per top-25 queries, referral clicks, and query-bank SOV.
  • Note movers (pages up or down >20%, queries up or down >10 pp).
  • Update content roadmap based on movers.
  • One-line client report: "AI visibility trending [up/down/flat] with X drivers. Notable movers: A, B, C."

Quarterly (2 hours)

  • Query bank review. Are the queries still matching how buyers actually search? Add / remove as needed.
  • Tool cost review. Are all three tools still earning their keep? Cut anything that isn't.
  • Benchmark refresh: pull the latest Semrush industry index or Profound source-pool research. Compare your vertical's baseline vs 90 days ago.
  • One breakdown on the highest-value query in the bank. Read the actual AI answers being generated. Are you positioned as the primary source, a secondary source, or absent? What would move you up a slot?

Annual (half a day)

  • Full framework review. Are the five signals still the right five? Are new engines significant enough to add? Are any old engines fading enough to drop?
  • Rebuild the dashboard from scratch if the tooling has changed materially.
  • Rewrite the query bank if buyer language has shifted (usually 20-40% of the bank rotates annually in fast-moving categories).

The point of the schedule is defence against dashboard bloat. If you're checking too often you'll react to noise. Too rarely and you'll miss drifts.

Free AI brand-mention checker (coming soon)

We're shipping a free AI brand-mention checker at gogochimp.com/tools/ai-brand-mention-checker in Q3 2026. It runs your brand name through ChatGPT, Perplexity, Gemini, and Google AI Overviews on 10 shortlist queries per category, samples each query 5-10 times to average out statistical noise (per the IQRush stopping rule), and returns your AI Share of Voice with margin of error. Zero login required.

If you want to be first in line to test it, book the free AI Visibility Audit and we'll email you when the checker goes live.

Two commercial tracking options for readers of this guide:

  • Free AI brand-mention checker (Q3 2026): 10-query sample across four engines, statistical-noise-adjusted, no signup required. Book the audit at /audit to be first on the waitlist.
  • AI Visibility Audit (£495 one-off): 40-query sample across five engines, competitor lift table, category coverage map, technical audit (robots.txt + crawler access + brand-consistency), and a written report with the specific 90-day work programme to move your AI Share of Voice up by 10-20 percentage points. Book at /audit.
  • Ongoing tracking service (from £2,000/month): full monthly 4-metric dashboard, versioned query bank, competitive share-of-voice reports, quarterly strategic review. Written specifically for agencies, law firms, hotel groups, and multi-location local brands who need a defensible board-level metric. Contact us to scope.

Predictions for AI visibility tracking 2026-2027

Five predictions worth naming for the next 12-18 months.

Prediction 1: Google will ship a dedicated AI Overviews impression segment in Search Console by end of 2026

The AIO impression data will separate from total organic impressions inside the Performance report, closing the current measurement gap between Bing WMT and GSC. When this ships, Google + Bing will provide roughly 70% of AI visibility measurement for free, first-party.

Prediction 2: Perplexity will expand its Publisher Program citation data surface to non-enrolled publishers

The current Comet Plus program provides citation data only to enrolled publishers. Competitive pressure from Bing WMT will force Perplexity to expose read-only citation data more broadly, likely tied to Search Console-style domain verification.

Prediction 3: Proxy tool consolidation

Profound, Ahrefs Brand Radar, Semrush AI Visibility, Presence AI, and 20+ smaller trackers currently compete on overlapping feature sets. Expect 3-4 winners in this category by end of 2027, with the rest either acquired or shut down. The winners will be the ones with first-party data partnerships with at least one major engine (equivalent to Ahrefs' Google data infrastructure advantage).

Prediction 4: Semrush's 45-point integration gap will close, but slowly

The finding that integrated teams see 81% traffic-or-lead lift vs 36% for separated teams is currently a market inefficiency. As agencies restructure and enterprise vendors bundle SEO + AI visibility, the gap will compress to maybe 20-25 points by end of 2027. Teams that integrate first still capture the near-term advantage.

Prediction 5: The industry will settle on "citation share of voice" as the primary KPI

Right now different tools report different primary metrics (absolute citations, brand mention frequency, source-share, SOV). Within 12-18 months, share of voice on a defined query bank will become the settled category standard, the way "organic keyword ranking" became the SEO standard in the early 2010s. The tools that report SOV cleanly and let teams define custom query banks will win the category.

None of these are locked-in forecasts. They're what the current trend lines point to. Update them quarterly.

FAQ

What's the difference between AI visibility tracking and traditional SEO tracking?

Traditional SEO tracking measures your rank position on a keyword in a deterministic SERP. AI visibility tracking measures your citation share on a query across many probabilistic AI-generated answers. Same query, run twice, returns different brand lists less than 1% of the time in the same order per SparkToro. You measure share across many runs, not rank on one run.

Is Bing Webmaster Tools' AI Performance report really free?

Yes. Bing WMT is free (has been since launch), and the AI Performance surface inside it is included in the free tier with no rate limits. All you need is site verification via DNS or a meta tag. Sign up at bing.com/webmasters. If you're serious about AI visibility tracking and don't have Bing WMT claimed, that's the highest-return 30-minute setup task in the discipline.

How much does Profound cost?

Profound uses custom pricing. Public pricing isn't listed on their marketing pages as of July 2026 because pricing scales with tracked-query volume and engines covered. Contact their sales team for a quote. Ballpark from teams we've spoken to: enterprise plans start around $2,000-5,000 per month. If Profound's price feels high, Ahrefs Brand Radar (bundled into existing Ahrefs plans from £99/month) or Semrush AI Visibility (bundled from £119/month) are cheaper cross-engine alternatives.

Does Google Analytics 4 show me referral traffic from ChatGPT and Perplexity?

Yes. GA4 records referral traffic from chat.openai.com, chatgpt.com, perplexity.ai, gemini.google.com, and copilot.microsoft.com as normal referrer sessions. Filter GA4's Acquisition report to those referrers to see AI engine traffic. Note: ChatGPT referral traffic dropped 52% after 21 July 2025 as ChatGPT shifted toward "answer-providing" mode, so the traffic-side signal is often smaller than the citation-side signal.

How often should I check my AI visibility tracker?

Weekly for anomaly detection (5 minutes reviewing Bing WMT's daily citation graph). Monthly for the full dashboard review (30 minutes across all five signals). Quarterly for query bank + tool cost review (2 hours). Checking daily creates noise-driven reactions; checking annually misses drifts. Monthly is the right schedule for most decisions.

How many queries should be in my AI visibility query bank?

20-50 queries. Under 20 is too small a sample for reliable share-of-voice reporting. Over 50 becomes hard to maintain and dilutes the focus. Split into 5-10 branded, 10-20 category, 5-10 informational. Change the bank quarterly if buyer language shifts; leave it stable otherwise so the trend line is meaningful.

Can I track AI visibility for free?

Mostly, yes. Bing WMT (free) covers Microsoft Copilot with first-party data. GSC (free) covers Google AI Overviews impressions. GA4 (free) covers referral clicks from all major AI engines. Semrush's public AI Visibility Index (free) covers industry benchmarks. That's roughly 60-70% of AI visibility measurement in July 2026. What you can't track for free is per-brand ChatGPT / Perplexity / Claude proxy data, which requires a paid tool.

What's a "grounding query"?

A grounding query is the underlying sub-query the AI engine decomposed a user's natural-language question into, and then used to retrieve source material. If a user asks Copilot "which CRO agencies are best for Shopify stores in the UK", Copilot might decompose that into three grounding queries: "best Shopify CRO agencies UK", "top CRO agencies for Shopify", "UK CRO agencies for ecommerce". Each grounding query is a citation opportunity, and Bing WMT reports each one separately.

What's a healthy citation share on a commercial query?

Under 5% is background noise (occasionally cited but not consistently). 5-15% is healthy presence. 15-30% is strong. 30-50% is dominant. Above 50% is category ownership and is rare (usually reflects Wikipedia-level entity fame or a category where only 2-3 well-answered pages exist). GoGoChimp's 62.75% on "best Shopify CRO agencies UK" is the ceiling case, not a realistic target for most brands.

Does AI visibility tracking replace SEO tracking?

No. Google organic traffic still matters and is still measurable. AI visibility tracking is additive to SEO tracking, not a replacement. The Semrush 45-point outcome gap between integrated and separated teams is evidence for this: teams that run both surfaces together outperform teams that pick one. Run both. Report them separately.

Can I use ChatGPT to check my own AI visibility?

Only as a sanity check, not as a measurement source. Ask ChatGPT the same query five times in different sessions and you'll get five different answers. Any single check tells you nothing statistically meaningful. If you want to sanity check what a real user might see, run the query 20+ times and look at aggregate patterns, not any single response.

What's the biggest measurement mistake teams make?

Single-query snapshots reported as if they were rankings. "Chris asked ChatGPT for the best CRO agency and we came up #3" is not a metric. It's one sample from a distribution. Per SparkToro's research across 2,961 runs, the same prompt returns the same list less than 1% of the time. Measure share across many runs, always.

Where to go next

If you're starting from zero, the sequence is:

  1. Claim Bing Webmaster Tools today. Wait 2-4 weeks for the AI Performance report to populate.
  2. Define your 20-50 query bank based on how your buyers actually search.
  3. Filter GSC to the query bank. Note impression trends and CTR.
  4. Set up GA4 referrer segmentation for the five major AI engine domains.
  5. Add one cross-engine proxy tool (Profound, Ahrefs Brand Radar, or Semrush AI Visibility depending on your existing stack).
  6. Build the six-section dashboard in Looker Studio, Notion, or Google Sheets.
  7. Set the review schedule: weekly (5 min), monthly (30 min), quarterly (2 hours).

If you want the full theory behind why AI visibility matters and how to earn citations in the first place, read our pillar guide on Generative Engine Optimisation (GEO). It covers the 8-step framework for earning AI citations, the five engines that matter, and the same first-party data referenced throughout this guide.

If you're focused on the technical foundations, the 2026 guide to schema markup for AI SEO walks through the exact JSON-LD blocks that show up as retrieval-trust signals in the citation surfaces we track.

For a broader case for building AI-first CRO into your growth stack, our AI CRO service page explains how expert-guided AI conversion optimisation compounds with AI visibility work. The visibility gets buyers to your page. The CRO turns them into revenue.

Or if you'd rather have us run the tracking build and the AI-visibility content system for you: our free CRO audit will identify where your current AI-search footprint is under-earning, and where a small set of tracked, formatted, cited pages could enter the citation graph. Book at gogochimp.com/audit. We're taking on 3-4 audit slots per month while the discipline is still 91% under-measured.

References

  • Averi / Chmael, Z. (2026). We Ran 50 B2B SaaS Queries Through Google AI Mode. https://www.averi.ai
  • Leapd / Azamfar, C. (2026). How ChatGPT, Google AI Overviews, and Perplexity Source Information in 2026. https://leapd.ai
  • Semrush. (2026). AI Search Visibility Checker (free). https://www.semrush.com/features/ai-search-visibility-checker/
  • Semrush. (2026). AI SEO Tips: How to Earn Citations & Mentions in AI Search. https://www.semrush.com/blog/ai-seo-tips/
  • Sielinski, R. / IQRush + University of St. Gallen (Schulte, Bleeker, Kaufmann). (2026). AI Visibility Rankings Aren't Stable. Statistical Noise on Single Readings. https://www.searchenginejournal.com/ai-visibility-rankings-arent-stable-new-research-shows-its-mostly-statistical-noise/581905/
  • Yext. (2026). Brand Visibility FAQ: Your AI Search Questions, Answered. https://www.yext.com
  • Something Familiar. (2026). AI Brand Visibility: Brand Strategy for AI Search. https://somethingfamiliar.co.uk/ai-brand-visibility/
  • Capgemini. (2026). Beyond SEO: How to Win Visibility and Influence in AI Search. https://www.capgemini.com/insights/expert-perspectives/beyond-seo-how-to-win-visibility-and-influence-in-ai-search/
  • Marketing Dive. (2026). Why Unpaid Media Is Now Essential to AI Visibility. https://www.marketingdive.com/news/why-unpaid-media-is-now-essential-to-ai-visibility/822485/
  • ET BrandEquity. (2026). The New Marketing Metrics Nobody Is Measuring Yet. https://brandequity.economictimes.indiatimes.com/news/marketing/the-new-marketing-metrics-nobody-is-measuring-yet/132284335
  • Otterly. (2026). AI Search Citation Tracking Platform. https://otterly.ai
  • Peec. (2026). Per-Engine AI Visibility Tracking. https://peec.ai
  • ZipTie. (2026). LLM Citation and AI Visibility Tracker. https://ziptie.dev
  • LLMrefs. (2026). LLM Citation Reference Monitoring. https://llmrefs.com
  • Authoritas. (2025). The State of AIOs: User Intent Research, December 2024. https://www.authoritas.com/seo-ai-research-whitepapers/the-state-of-aios-user-intent-research-dec-2024
  • Averi. (2026). ChatGPT vs Perplexity vs Google AI Mode: The B2B SaaS Citation Benchmarks Report. https://www.averi.ai/how-to/chatgpt-vs.-perplexity-vs.-google-ai-mode-the-b2b-saas-citation-benchmarks-report-%282026%29
  • Bing Webmaster Tools. (2026). AI Performance report, verified 2026-07-11. https://www.bing.com/webmasters/aiperformance
  • Google Search Central. (2026). AI features in Search: guidance for site owners. https://developers.google.com/search/docs/appearance/ai-features
  • Muck Rack + Seer Interactive. (2026). What Is AI Reading? May 2026 25M-link study. https://muckrack.com/blog/what-is-ai-reading-may-2026
  • Perplexity. (2026). Introducing the Perplexity Publishers Program. https://www.perplexity.ai/hub/blog/introducing-the-perplexity-publishers-program
  • Profound. (2026). AI Platform Citation Patterns. https://www.tryprofound.com/blog/ai-platform-citation-patterns
  • Ranqo. (2024). Niche Brand AI Visibility Baseline: 100K-response arXiv study. https://arxiv.org/abs/2504.09347
  • Seer Interactive. (2026). AIO Impact on Google CTR: 2026 Update. https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update
  • Semrush. (2026). Semrush Releases Expanded 2026 AI Visibility Index Analysing 126 Million AI Search Prompts. https://www.semrush.com/news/463141-semrush-releases-expanded-2026-ai-visibility-index-analyzing-126-million-ai-search-prompts/
  • SparkToro / Fishkin, R. (2026). NEW Research: AIs Are Highly Inconsistent When Recommending Brands or Products. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
  • Superlines. (2026). AI Search Statistics: 615x Platform Citation Variance. https://www.superlines.io/articles/ai-search-statistics/
CITED book coverCITED book inner page

Free chapter

Read Chapter 1 of CITED, free.

The playbook for getting your business recommended by ChatGPT and AI search. Read the first chapter, on me.

Read Chapter 1 free

Want us to do this for your site?

Book a free AI audit. 15 minutes. We’ll show you three things your site is missing and what we’d test first.

Book my free AI audit →

Keep reading

Pillar

Related post title — bind from Related Posts multi-ref

Chris McCarron · 7 min read

Pillar

Related post title — bind from Related Posts multi-ref

Chris McCarron · 7 min read

Pillar

Related post title — bind from Related Posts multi-ref

Chris McCarron · 7 min read

© 2026 GoGoChimp. All rights reserved. Call: 0141 463 6875 - Address: 8 Cheviot Drive, Newton Mearns, Glasgow, G77 5AS
Nominated — Digital Doughnut Digital Marketing Agency of the Year 2021
Shopify Partner — GoGoChimp
Select the comment + the next block ONLY (3 lines total). Paste EVERYTHING below in its place. --> '"'"'""')})}}) "'"')}})