
GoGoChimp Research · White paper · October 2026
On 6 October 2026 we crawled 484 agency websites and scored 415 of them for how likely AI engines are to quote them. This is what a sharp prospect would find.
415
agency sites scored
70
median Rubric score
52
median Trusted score, the weakest pillar
14 of 646
service pages signed and sourced
If you sell AI search work, a sharp prospect will check your own website first. On 6 October 2026 we crawled 484 agency sites and scored 415 of them, to see what that prospect would find. The access checks passed almost everywhere, but the trust signals didn’t. On Rubric’s Trusted pillar, 168 of the 415 scored under 50. Only 14 of the 646 service pages we crawled had both a named author and two external sources. So did just 14 of the 570 case studies.
Ten findings
Contents
How to read this paper. Agency figures come from GoGoChimp’s Rubric crawl of 6 October 2026. Any other source is named where it’s used. Rubric scores pages from 0 to 100, and 70 is the line at which a page counts as quotable. Rubric predicts citability, it doesn’t count citations. Top performers are named, positively. No other agency is named.
Your clients are going to ask whether ChatGPT, Claude, Perplexity or Google’s AI answers can quote them. Before they believe your answer, they can run the same test on your own site. It’s the cheapest due diligence a buyer can do.
Agency sites ought to be the easiest test case on the web. They’re written by people who sell search for a living.
The plumbing’s fine. What’s missing is the layer that tells an AI engine what a page is and who stands behind it.
That gap costs you twice: with no author markup, you’ve got a weaker site and a weaker pitch. The client can see you haven’t done for yourself what you’re selling.
We crawled 533 sites on one engine, at one page cap, on one day. That’s 484 agencies and 49 SEO-tool vendors, all on Rubric scoring version 2026-10-01 at up to 40 pages each. The crawl ran on 6 October 2026 between 14:35 and 20:12. We’ve kept the tool vendors separate. Sites we couldn’t crawl properly are counted, but they’re never scored.
We crawled every site on the five lists behind our earlier crawls, apart from three Prolific North domains that now redirect to another brand.
| List | Who’s on it | Agencies attempted | Agencies scored |
|---|---|---|---|
| 12-13 Aug seven-country list | SEO agencies in the USA, UK, Canada, Australia, Germany, France and Spain, plus the 49 SEO-tool vendors | 243 | 200 |
| 23 Aug English-speaking list | Agencies in the US, UK, Canada, Australia, New Zealand, Ireland and South Africa | 100 | 89 |
| 23 Aug international list | International and mid-tier agencies (Netherlands, Ireland, India, the Nordics, Italy, Poland, Portugal, Singapore, UAE, South Africa, Brazil and a few US, UK and Australian sites) | 99 | 89 |
| Prolific North Top 50 Digital Agencies 2025 | Digital agencies in the North of England | 47 | 42 |
| 10 UK SEO agencies | A cross-check list | 10 | 10 |
Fourteen agencies sit on two or three lists, so the rows add up to more than 484 and 415. We ran every crawl ourselves rather than through Rubric’s hosted service, with the broken-link check and the Common Crawl lookup switched off.
Rubric crawls a site up to a page cap and runs every page through the same set of checks. They roll up into three pillars and a score out of 100.
Rubric predicts citability, it doesn’t count citations. A high score means a page is easy to quote. The score says nothing about an agency’s client work. The per-engine fits in Appendix A are modelled weightings of the same checks.
Every crawl that clears six failure rules gets a score. We checked each site against them in this order, and the first rule that applied set its status.
In all, 415 of 484 agencies (85.7%) and 47 of 49 tool vendors were scored.
| Why an agency wasn’t scored | Sites |
|---|---|
| Firewall challenge pages | 25 |
| Unreachable (DNS 5, certificate 3, timeout 2, browser error 2) | 12 |
| Homepage only | 10 |
| Engine block flag | 9 |
| Degraded | 8 |
| Failed | 5 |
| Total | 69 |
Yes, 24 of them. We gave 14 a second pass, two sites at a time, mostly because their homepage had timed out or thrown a browser error. We tried 20 on their www address after the bare domain failed, and some sites got both. Six of the 24 are scored now, two of them from the www address.
We sorted pages by URL first, so /services/ marks a service page and /case-studies/ or /work/ a case study. Otherwise we used the crawler’s own label, and anything left went into “other page”. The crawler read 15,658 pages at the 415 scored agencies: 26.5% blog posts or guides, 4.1% service pages, 3.6% case studies and 45.4% other pages.
The engine also labels the root of any subdomain it reaches as a homepage, and some of those roots were error pages. So the page-type tables count 537 homepages at 415 sites. Every site-level homepage comparison in this paper uses only the homepage on the address we crawled.
Yes, by 5 points at the median. Sites where blog posts made up under a quarter of the crawl had a median of 68 (n=265). Sites where they made up half or more had 73 (n=103). Blog posts score above most inner pages, so a crawl that reaches mostly posts lifts the site’s score. Every site had the same 40-page cap, but the crawler chose which pages it read.
Scores barely move on the same engine. We’d crawled the 42 Prolific North sites on this engine on 5-6 October. A day later, the median absolute move was 0 and no site moved 10 points or more. The list’s median (65.5) and Trusted median (48) came out identical. The 10 UK SEO agencies, crawled at 25 pages on 5-6 October, moved by a median of 1.5.
Trusted scores fell under the current engine, which is stricter on named authors and on content-type schema. Our August crawls ran on an older, unversioned engine, and the August figures below are the only ones in this paper. Don’t set them beside anything else here.
We scored 203 sites from the 12-13 August list both times, tool vendors included. Their Trusted median fell from 59 in August to 54 on 6 October. Of those 203, 97 moved 10 points or more on Trusted, while the overall median moved by 1 point.
The August block flag also over-fired. Of the sites it capped as blocked in August, 44 were scored normally on 6 October.
Across the 415 agencies, the Trusted median was 52, against 80 for Known and 71 for Findable. Only 61 agencies (14.7%) reached 70 on Trusted, and 168 scored under 50.
Median pillar score, 415 agencies
Trusted is the only pillar with a median below 70, the quotable line. 168 of the 415 agencies scored under 50 on it. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
| Pillar | Median | Sites at 70+ | Sites under 50 |
|---|---|---|---|
| Known | 80 | 339 (81.7%) | 3 |
| Findable | 71 | 251 (60.5%) | 1 |
| Trusted | 52 | 61 (14.7%) | 168 |
Trusted came last on every list in the crawl, from 46.5 for the 10 UK SEO agencies to 54 for the English-speaking list. Appendix A has each list.
The Trusted pillar checks for evidence: a named author, at least two external source links and statistics that carry a source. It also wants at least 1.5 statistics per 100 words and a date inside the last 12 months. I’d put every one of those in a client audit. On agency sites, they’re the signals that go missing.
My read is that agency sites are built as brochures. Most service pages carry nobody’s name and link to nothing outside the site. That’s fine for a buyer who’s already booked a call. But it gives an AI engine nobody to credit and nothing to check the claim against.
Trusted was the weakest of Rubric’s three pillars across the 415 agency sites GoGoChimp crawled on 6 October 2026, with a median of 52. Of the 415, 168 scored under 50 on it.
Trust goes missing on the pages away from the homepage and the blog. The median homepage scored 67 on Trusted and the median blog post 66. The median service page scored 46, the median case study 45 and the median about, team or contact page 48. They’re page-level medians, which is why they differ from the site-level 52.
Just under half of agency blog posts are more than a year old. Of 4,151 blog posts, 3,944 (95.0%) carried a date, and the median dated post was 357.5 days old. Of the dated posts, 1,943 (49.3%) were more than a year old and 1,339 (34.0%) more than two.
Rubric’s freshness check failed at only 89 of 367 sites (24.3%), so the dates tell you more than the check does. Because the crawl stops at 40 pages, a site’s newest post isn’t always among them.
The median service page scored 70 and the median case study 68, against 75 for a blog post and 80 for a homepage. About, team and contact pages scored lower still, at 66. The crawl reached service pages at 175 of 415 sites and case studies at 157.
Median page score by page type
Service pages and case studies, the pages that sell, sit at or below the line. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
| Page type | Pages | Sites | Median page score | Known | Findable | Trusted | Median words | Median external links |
|---|---|---|---|---|---|---|---|---|
| Homepage | 537 | 415 | 80 | 86 | 80 | 67 | 895 | 1 |
| Blog post or guide | 4,151 | 292 | 75 | 90 | 71 | 66 | 995 | 1 |
| Service page | 646 | 175 | 70 | 77 | 74 | 46 | 884 | 0 |
| Case study | 570 | 157 | 68 | 77 | 71 | 45 | 513 | 0 |
| About, team, contact | 1,131 | 347 | 66 | 77 | 69 | 48 | 429 | 1 |
| Other page | 7,101 | 388 | 68 | 75 | 71 | 48 | 798 | 0 |
The homepage row includes subdomain roots, which is why it’s 80 here and 82 in the homepage section. Appendix C has every page type.
Named author fails most, on about eight pages in ten. Each cell is the share of applicable pages that came back bad.
| Check | Blog post or guide | Service page | Case study |
|---|---|---|---|
| Named author | 9% | 82% | 78% |
| Content-type schema | not reported | 61% | 60% |
| Two or more external source links | 44% | 60% | 61% |
| Answer-first opener | 52% | 52% | 46% |
| Question or claim-shaped headings | 43% | 50% | 68% |
| Statistics carry a source | 28% | 27% | 49% |
We haven’t reported the blog-post schema figure. The crawler labels pages with Article schema as articles, so it would be partly circular.
Hardly any: 14 of 646 service pages (2.2%) and 14 of 570 case studies (2.5%). We call a page signed and sourced when it passes both the named-author check and the two-or-more external links check.
The byline is the bigger gap. Only 28 of 600 applicable service pages (4.7%) passed the named-author check, and 379 of 646 (58.7%) had no external link at all.
Most agency case studies are unsigned and unsourced, and one in four runs under 300 words. Of 570 case studies, 347 (60.9%) had no external link and 140 (24.6%) were that short. Only 77 of 554 applicable case studies (13.9%) passed the named-author check, and 83 (14.6%) declared an Article-family schema type.
A buyer’s specific questions land on these pages. “Can this agency handle technical SEO for a Shopify store?” is a service-page question. “Have they done it for anyone like us?” is a case-study question. Most of those pages are unsigned, and most link to nothing a buyer could check.
In GoGoChimp’s 6 October 2026 crawl of 415 agency sites, 14 of 646 service pages and 14 of 570 case studies had both a named author and two external source links.
The homepage outscored the median inner page at 373 of 415 sites (89.9%). The median homepage scored 82 and the median inner page 70. Site by site, the median gap was 12 points.
Median gap below the homepage, in points
Median points each page type scored below the same site’s homepage. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
The homepage is the agency in its interview suit. The inner pages are the same agency at home on a Sunday, in a dressing gown, eating cereal out of the box.
The service page a buyer lands on is the one in the dressing gown.
Every kind of inner page scores below the homepage at most agencies. For service pages and case studies, it’s nine sites in ten.
| Page type | Scored below the homepage at | Median gap (points) |
|---|---|---|
| About, team, contact | 322 of 347 sites (92.8%) | 16 |
| Service page | 158 of 175 sites (90.3%) | 13 |
| Case study | 143 of 157 sites (91.1%) | 13 |
| Blog post or guide | 223 of 292 sites (76.4%) | 7 |
The fewer pages a crawl reaches, the more the homepage carries the score. One agency’s crawl reached 2 pages and scored 88, which would top our table if we ranked it. We don’t rank any crawl under 10 pages, and we never score a one-page crawl. Thirteen scored crawls reached fewer than 10 pages, and dropping them leaves the median at 70.
Test a service page and a recent post, and do the same when you audit a prospect. A homepage score on its own flatters the site.
Across 415 agency sites crawled by GoGoChimp on 6 October 2026, the median homepage scored 82 on Rubric and the median inner page 70.
The gap comes from content and attribution checks. Question or claim-shaped headings failed at 85 of the 104 lowest-scoring agencies, against 18 of 102 in the top quarter. Entity clarity, content-type schema, external sources, answer-first openers and named author split them the same way. The HTTP, robots.txt and canonical checks didn’t separate them at all, because no agency in the crawl failed them.
Share of sites failing the check, top quarter against bottom quarter
Question or claim-shaped headings
Content-type schema
Two or more external sources
Answer-first opener
Named author
Entity clarity
Statistic density
Counts are sites failing out of 104 unless shown. The gap sits on attribution and markup, not on access. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
The top quarter is the 104 agencies scoring 75 to 88, and the bottom quarter the 104 scoring 46 to 65. Their medians were 78 and 62 overall, and 67.5 and 43 on Trusted.
| Check | Top quarter failing | Bottom quarter failing |
|---|---|---|
| Question or claim-shaped headings | 18 of 102 | 85 of 104 |
| Entity clarity (Organization or Person plus sameAs) | 6 of 104 | 57 of 104 |
| Content-type schema | 22 of 104 | 76 of 104 |
| Two or more external source links | 26 of 104 | 77 of 104 |
| Answer-first opener | 29 of 102 | 72 of 104 |
| Named author | 35 of 102 | 74 of 104 |
| Statistic density (1.5 or more per 100 words) | 3 of 104 | 32 of 104 |
Counts out of 102 are for checks that didn’t apply at every site. Appendix E has five more checks. Every one of these checks feeds the score that sorts agencies into quarters, so part of each gap is built in. None of it proves what an engine rewards.
The gap is mostly on the inner pages. Top-quarter homepages had a median of 89 and bottom-quarter homepages 76, 13 points apart. Their median inner pages were 78 and 61, 17 points apart.
The best agencies sign and source most of their inner pages. At conquerradigital.com.au (85, Trusted 92), 34 of 38 applicable inner pages had a named author and 36 of 38 had two or more external source links. All but one had sourced statistics. At firestarterseo.com, the top scorer at 87, 33 of 38 had a named author and 33 of 38 external sources. All 38 carried content-type schema.
In GoGoChimp’s 6 October 2026 crawl, 85 of the 104 lowest-scoring agencies failed Rubric’s question-heading check. In the top quarter, 18 of 102 did.
Named author failed at the most sites, 249 of 413 (60.3%). External source links (224), content-type schema (208), question or claim-shaped headings (205) and answer-first openers (200) came next. No agency failed the HTTP, robots.txt, canonical, internal-link, reachability or word-count checks.
Share of agencies failing the check on most pages
Named author: 249 of 413. External sources: 224 of 415. Schema: 208 of 415. Headings: 205 of 413. Answer-first: 200 of 413. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
A site fails a check here when it comes back bad on more than half the pages it applies to.
| Check | Pillar | Sites failing | Median share of pages failing |
|---|---|---|---|
| Named author | Trusted | 249 of 413 (60.3%) | 70% |
| Two or more external source links | Trusted | 224 of 415 (54.0%) | 55% |
| Content-type schema (Article, Service and so on) | Known | 208 of 415 (50.1%) | 51% |
| Question or claim-shaped H2 and H3 headings | Findable | 205 of 413 (49.6%) | 50% |
| Answer-first opener (40-70 words) | Findable | 200 of 413 (48.4%) | 50% |
Without content-type schema, an engine has to work out from the copy what kind of page it’s reading. On the median agency site, 51% of applicable pages failed the check.
A heading like “Our approach” doesn’t tell an engine what question the section answers. An opener that spends its first 70 words on brand copy gives it nothing to lift. Both are writing problems, and they’re quick to fix.
A named author gives an engine a person to credit with the expertise, and external links give it something to cross-check. Without either, the page is an unsigned claim. On the median agency site, 70% of applicable pages failed the named-author check.
Signed, sourced posts score about 14 points more than posts with neither. Of 4,151 agency blog posts, 1,691 (40.7%) had a named author and two or more external source links, and their median page score was 81. The 195 posts (4.7%) with neither scored a median of 67.
Part of that gap is mechanical. Both checks feed the page score, so a page that passes them scores higher by construction. The gap shows what the two fixes are worth inside Rubric’s scoring. This crawl can’t tell you whether an engine rewards them.
In GoGoChimp’s 6 October 2026 crawl, 249 of 413 agency sites failed Rubric’s named-author check on more than half their pages. None failed the HTTP, robots.txt or canonical checks.
Service schema is rare, even on service pages. Only 139 of 415 agencies (33.5%) used Service as its own type on any crawled page, and 130 of 646 service pages (20.1%) declared it. More agencies carried FAQPage (194) than Service.
The table counts sites with each type on at least one crawled page, and Appendix D has every type.
| Schema type | Agencies (n=415) |
|---|---|
| Organization (or a subtype) | 366 (88.2%) |
| WebSite | 333 (80.2%) |
| BreadcrumbList | 310 (74.7%) |
| Person | 285 (68.7%) |
| Article family | 261 (62.9%) |
| FAQPage | 194 (46.7%) |
| Service as its own type | 139 (33.5%) |
| Offer | 114 (27.5%) |
| Review or AggregateRating | 106 (25.5%) |
Agency sites mostly carry generic schema. Organization and WebSite top the table, and those two tell an engine whose site it is. Neither says what the agency sells.
My guess is it’s the CMS. Organization, WebSite and BreadcrumbList are the three commonest types. In my experience, that’s the graph an SEO plugin emits out of the box. Rubric didn’t detect any plugin, and I haven’t checked site by site.
Review markup is missing at most agencies, and a few carry no schema at all. Only 106 agencies (25.5%) declared review or rating markup, and 21 (5.1%) had no schema on any crawled page. HowTo turned up at 15.
In GoGoChimp’s 6 October 2026 crawl, 130 of 646 agency service pages declared Service schema. More of the 415 agencies carried FAQPage (194) than Service (139).
Close to a quarter of agencies (92 of 415, 22.2%) declared no sameAs links, and only 6 linked to Wikidata. The median agency declared 4. JavaScript-only schema turned up at 44 sites (10.6%), and on more than half the crawled pages at 23.
| Signal | Agencies (n=415) |
|---|---|
| No sameAs links | 92 (22.2%) |
| Median sameAs count | 4 |
| Wikidata link | 6 |
| Entity clarity check failing | 106 (25.5%) |
| JS-only schema on 1+ page | 44 (10.6%) |
| JS-only schema on more than half the pages | 23 |
| llms.txt present | 185 (44.6%) |
Rubric doesn’t score llms.txt.
The sameAs property ties a brand’s schema to its profiles elsewhere: LinkedIn, Companies House, review sites, Wikidata. It’s how an engine confirms the agency on the page is the entity it’s read about somewhere else.
LinkedIn, then X, then YouTube, and very little else.
| Profile | Agencies listing it in sameAs |
|---|---|
| 279 (67.2%) | |
| X | 207 (49.9%) |
| YouTube | 149 (35.9%) |
| Crunchbase | 18 (4.3%) |
| GitHub | 7 (1.7%) |
| Wikipedia | 5 (1.2%) |
| G2 | 5 (1.2%) |
| 4 (1.0%) | |
| Trustpilot | 4 (1.0%) |
| Capterra | 1 (0.2%) |
Rubric only counts a fixed list of platforms. Clutch, Google Business Profile and Companies House aren’t on it, so these figures say nothing about them.
A crawler that doesn’t run JavaScript never sees schema that exists only in the rendered page. In Vercel’s 2024 analysis, “none of the major AI crawlers currently render JavaScript”, with OpenAI’s three bots, ClaudeBot and PerplexityBot among them (Vercel and MERJ, 2024). It’s a December 2024 study, and serving schema in the HTML costs nothing either way.
Rubric catches JS-only schema by comparing each page’s raw HTML with a rendered copy.
In GoGoChimp’s 6 October 2026 crawl, 92 of 415 agency sites declared no sameAs links in their schema, and only 6 linked to Wikidata.
Rubric found a comparison page at 192 of 415 agency sites (46.3%) and a how-to-choose page at 124 (29.9%). Only 43 (10.4%) had a machine-readable price on any crawled page.
Rubric sorts a buyer’s follow-up questions into five kinds and looks for a page that answers each, by URL, title and question headings.
| Kind of question | What Rubric looks for | Agencies with a page (of 415) |
|---|---|---|
| Compare | vs, alternatives, best X for Y | 192 (46.3%) |
| Constrain | how to choose, buyer’s guide, checklist, pricing guide, use cases | 124 (29.9%) |
| Clarify | what is, guide to, overview, FAQ, glossary | 293 (70.6%) |
| Validate | case study, results, testimonials, reviews | 308 (74.2%) |
| Act | pricing, get a quote, book a call, contact | 328 (79.0%) |
These counts are lower bounds. A page the crawler didn’t reach in its 40 doesn’t count, so some sites will have pages we’ve missed.
Buyers ask specific, comparative questions that a homepage can’t answer. These four are illustrative and don’t come from our data.
Agencies have plenty of pages for the last kind and far fewer for the first two.
A buyer will ask what it costs, and an engine can only quote a number someone’s published. You don’t have to publish your rates. A starting price or a range on one page gives an engine something of yours to quote.
GoGoChimp’s 6 October 2026 crawl found a comparison page at 192 of 415 agency sites, and a machine-readable price at 43.
Two agencies paired a robots.txt block on an AI bot with a live 403 to that same bot, and both were blocking ClaudeBot. That pairing is what we count as a deliberate block. We probed 460 agencies, and we’ve counted the two rather than named them.
Thirteen do. Twelve of them block training crawlers only, and one also blocks a search or user-triggered bot. All 13 block CCBot. GPTBot is blocked at 8, Google-Extended at 6, ClaudeBot at 5 and PerplexityBot at 1.
A firewall can turn away any bot it can’t verify, whatever the site’s AI policy. A user-agent probe sends a request labelled as GPTBot or ClaudeBot and records the reply. Ours came from our own IP address, which isn’t on any vendor’s network. Of the 460 agencies we probed, 70 (15.2%) returned a 403 to at least one AI agent.
Of those 70, 26 refused every bot we sent, Googlebot and Bingbot included. Our Googlebot probe got a 403 at 35 agencies, more often than our OAI-SearchBot probe (30). That’s what a firewall turning away unverified bots looks like.
If a client’s tool says their site blocks AI, check the firewall before the robots.txt. If your own audit says it, re-probe before it goes in a deck. That includes ours.
Of 460 agency sites GoGoChimp probed on 6 October 2026, 70 returned a 403 to at least one AI user agent. Only 2 paired it with a robots.txt block on the same bot.
Allow all the search crawlers and answer-time fetchers, and make training a separate decision. The vendors’ own documentation splits their bots that way.
| Kind | Bot | What the vendor says |
|---|---|---|
| Search | OAI-SearchBot | “OAI-SearchBot is for search” (OpenAI) |
| Search | Claude-SearchBot | Blocking it “prevents our system from indexing your content for search optimization” (Anthropic) |
| Search | PerplexityBot | “designed to surface and link websites in search results on Perplexity” (Perplexity) |
| Search | Googlebot | Its rules “affect Google Search (including Discover and all Google Search features)” (Google) |
| Answer-time fetch | ChatGPT-User | “Because these actions are initiated by a user, robots.txt rules may not apply” (OpenAI) |
| Answer-time fetch | Claude-User | Used “When individuals ask questions to Claude” (Anthropic) |
| Answer-time fetch | Perplexity-User | “this fetcher generally ignores robots.txt rules” (Perplexity) |
| Training | GPTBot | Disallowing it “indicates a site’s content should not be used in training generative AI foundation models” (OpenAI) |
| Training | ClaudeBot | Restricting it “signals that the site’s future materials should be excluded from our AI model training datasets” (Anthropic) |
| Control token | Google-Extended | Covers Gemini training and grounding, and “does not impact a site’s inclusion in Google Search” (Google) |
Google-Extended isn’t a crawler of its own. Google calls it a “standalone product token” that’s used “in a control capacity”, and Google’s usual crawlers do the fetching (Google).
Most of the refusals we saw weren’t aimed at AI search. Besides the 26 agencies that refused every bot, 33 of the 70 refused only ClaudeBot, which is Anthropic’s training crawler. That’s 59 of the 70.
On an agency site, I’d let the four search crawlers and three fetchers through both robots.txt and the firewall. The fetchers may ignore robots.txt anyway, by OpenAI’s and Perplexity’s own account, so the firewall decides what they get.
A firewall challenge page met our crawler on most URLs at 25 of the 484 agency sites we attempted (5.2%). We didn’t score those sites.
A challenge page has nothing on it to quote, whoever it’s meant to stop.
We can’t see from outside whether verified AI crawlers get the same page. Put that question to your host or firewall provider: what do GPTBot, ClaudeBot and PerplexityBot receive on their first request? If it’s a challenge page, that’s what they read.
Firewall defaults have shifted under agencies, too. In July 2025, Cloudflare said it was “changing the default to block AI crawlers unless they pay creators for their content” (Cloudflare, 2025). In July 2026 it said “non-verified bots are still default blocked” (Cloudflare, July 2026). On 15 September 2026, it said its Block settings “now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot”, so they affect search as well as training (Cloudflare, September 2026). We can’t tell which provider or setting any of our challenged sites used.
On 6 October 2026, 25 of the 484 agency sites GoGoChimp tried to crawl served a firewall challenge page on most URLs, about 1 in 20.
Rubric’s model rated Gemini the weakest engine fit at 152 of 415 agency sites, and AI Overviews at 111. ChatGPT was weakest at 63, Perplexity at 51, Claude at 20 and Copilot at 18. These fits are modelled, and we didn’t query any engine.
In Rubric’s current engine, Gemini’s heaviest checks are content-type schema, entity clarity and indexability. For AI Overviews, they’re answer-first openers, question headings and indexability. Three of those are among the five checks agencies failed most often, so the two fits that lean on them sit lowest. Appendix A has the median fit for each engine.
I wouldn’t chase one engine’s fit. The checks that pull Gemini and AI Overviews down are the same ones the fixes below deal with.
SEO-tool vendors scored a median of 67 (n=47) and agencies 70 (n=415), on the same engine, page cap and day. The gap sits in Known, at 73 against 80. Findable went a point the other way (72 against 71), and Trusted was close (51 against 52).
| SEO-tool vendors (n=47) | Agencies (n=415) | |
|---|---|---|
| Median | 67 | 70 |
| Mean | 67.5 | 69.7 |
| Known | 73 | 80 |
| Findable | 72 | 71 |
| Trusted | 51 | 52 |
| Scoring 70+ | 13 (27.7%) | 222 (53.5%) |
I didn’t expect that.
I suspect product sites lean on app pages, docs and pricing tables. Those carry less of the content-type and entity markup Rubric checks under Known, though I haven’t tested it. The best tool sites still score well: keyword.com at 86, diib.com at 82 and yoast.com at 81.
In GoGoChimp’s 6 October 2026 crawl, SEO software vendors scored a median of 67 on Rubric (n=47). SEO agencies in the same crawl scored 70 (n=415).
Agencies scored a median of 70, against 66 for 244 other websites we crawled on 7 October, a day later, with the same engine and the same 40-page cap. Those sites are Rubric’s benchmark corpus: online shops, blogs and publishers, SaaS companies and other businesses. Because both crawls ran under the same rules, the scores compare directly.
Median overall and Trusted score
Other websites: Rubric’s benchmark corpus, re-crawled on 7 October 2026 on the same engine and 40-page cap. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
| Agencies (n=415) | Other websites (n=244) | |
|---|---|---|
| Median | 70 | 66 |
| Scoring 70+ | 222 (53.5%) | 55 (22.5%) |
| Known | 80 | 71 |
| Findable | 71 | 68 |
| Trusted | 52 | 49 |
| Homepage ahead of inner pages | 373 (89.9%) | 225 (92.2%) |
More than half the agencies cleared the 70 line, against fewer than a quarter of the other sites. The biggest gap is on Known, at nine points. On Trusted, the agency lead shrinks to three.
Agencies sign, structure and mark up their pages more often than other sites do. They don’t cite outside sources much more often.
| Check failing on most pages | Agencies | Other websites |
|---|---|---|
| Named author | 249 of 413 (60.3%) | 198 of 242 (81.8%) |
| Answer-first opener | 200 of 413 (48.4%) | 197 of 242 (81.4%) |
| Question or claim-shaped headings | 205 of 413 (49.6%) | 181 of 242 (74.8%) |
| Content-type schema | 208 of 415 (50.1%) | 171 of 243 (70.4%) |
| Two or more external source links | 224 of 415 (54.0%) | 138 of 244 (56.6%) |
So agencies practise more of what they preach than the average site, except on sources. Sources are one of the Trusted signals, and missing them helps keep both groups under the line on trust.
In GoGoChimp’s October 2026 crawls, 415 agency websites scored a median of 70 on Rubric, against 66 for 244 other websites crawled a day later on the same engine.
Agency medians ran from 73 in Australia (n=36) to 63.5 in South Africa (n=10), and the UK sat at 69 (n=89). With 10 to 92 scored sites per group, read the table as description and ignore a 1 or 2 point gap between neighbours.
| Country | Sites | Median | Trusted median |
|---|---|---|---|
| Australia | 36 | 73 | 57 |
| USA | 43 | 71 | 51 |
| Germany and Austria | 22 | 71 | 58 |
| Canada | 34 | 70.5 | 54.5 |
| France | 18 | 69.5 | 50.5 |
| Nordics | 10 | 69.5 | 50 |
| UK | 89 | 69 | 51 |
| Spain | 17 | 69 | 54 |
| South Africa | 10 | 63.5 | 57 |
| Generic domain, no country | 92 | 71 | 51.5 |
We took the country from the 12-13 August list where it gave one, and otherwise from the domain ending. The table shows groups with 10 or more scored agencies, and the other 44 sit in smaller ones. Appendix B adds Known and Findable.
The middle half of the 415 agencies spans 10 points, from 65 to 75. No agency reached 90, and 28 scored in the 80s. Another 194 sat in the 70s and 161 in the 60s.
Agencies by score band (n=415)
355 of the 415 agencies scored between 60 and 79. Source: GoGoChimp Rubric crawl of 6 October 2026, scoring version 2026-10-01, up to 40 pages per site.
For a mid-table agency, the top 10 was 12 points away, with tenth place on 82 against a median of 70.
We only rank crawls of 10 pages or more. Four agencies tied for tenth, so the table has 13.
| Rank | Site | Score |
|---|---|---|
| 1 | firestarterseo.com | 87 |
| 2 | aztekweb.com | 86 |
| 3 | conquerradigital.com.au | 85 |
| 4= | netzbekannt.de | 84 |
| 4= | seokratie.de | 84 |
| 4= | soupagency.com.au | 84 |
| 4= | thatware.co | 84 |
| 8= | kiwop.com | 83 |
| 8= | rodanet.com | 83 |
| 10= | azurodigital.com | 82 |
| 10= | blennd.com | 82 |
| 10= | mediaforce.ca | 82 |
| 10= | titanblue.com.au | 82 |
Agencies would reach a median of 95 on Rubric’s model if every flagged fix were made. That’s a median gain of 24 points, with 380 of 415 projected at 90 or more. The projection is a ceiling, not a forecast, because it assumes every fix lands on every page.
No agency in GoGoChimp’s 6 October 2026 crawl of 415 agency sites reached 90 on Rubric. The middle half scored between 65 and 75.
We have published a fuller write-up of this subset: a regional deep-dive of this benchmark.
The 42 scored northern sites had a median of 65.5, against 70 for all 415 agencies, and a Trusted median of 48, against 52. They’re a subset of the same 6 October crawl, so the two compare directly.
The middle half of the northern agencies scored between 64 and 70, and 12 of the 42 (28.6%) reached 70 or above. Only 2 reached 70 on Trusted. The homepage beat the median inner page at 37 of the 42, by a median of 12.5 points.
The list is the Prolific North Top 50 Digital Agencies 2025, and it ranks digital agencies rather than SEO agencies. Part of the gap to the full benchmark may be the kind of agency, and this data can’t tell you how much.
Three sites tied for tenth on 70, so the table has 12.
| Rank | Site | Score |
|---|---|---|
| 1= | embryo.com | 80 |
| 1= | seoworks.co.uk | 80 |
| 3= | havasmarket.co.uk | 75 |
| 3= | idhlagency.com | 75 |
| 5 | visualsoft.co.uk | 74 |
| 6= | reasondigital.com | 73 |
| 6= | riseatseven.com | 73 |
| 6= | unrvld.com | 73 |
| 9 | velstar.agency | 71 |
| 10= | addpeople.co.uk | 70 |
| 10= | connective3.com | 70 |
| 10= | housedigital.co.uk | 70 |
The joint top site, seoworks.co.uk, had a Known score of 91. The two northern sites that reached 70 on Trusted are both in the top 10: riseatseven.com at 76 and embryo.com at 72.
The top northern sites source or sign most of their inner pages. At riseatseven.com, all 38 applicable inner pages had two or more external source links and sourced statistics. At embryo.com, all 18 applicable inner pages had both. At reasondigital.com, 31 of 37 applicable inner pages carried a named author and 32 of 37 carried content-type schema. At seoworks.co.uk, 23 of 30 had a named author and 24 of 30 had sourced statistics.
GoGoChimp crawled the Prolific North Top 50 Digital Agencies 2025 on 6 October 2026. The 42 scored sites had a median Rubric score of 65.5 and a Trusted median of 48, against 70 and 52 for all 415 agencies.
Focus on the service pages and case studies, because that’s where trust goes missing. Put a named author on them first, since it’s the most common failure and Trusted is the weakest pillar. Then add sources and content-type schema. All nine fixes are markup or writing jobs. The wider playbook for client sites is in our AI search optimisation hub.
| Fix | Rubric check | Agencies (n=415) |
|---|---|---|
| 1. Name the author | Named author | 249 of 413 sites failing |
| 2. Mark up the page type | Content-type schema | 208 sites failing |
| 3. Cite two or more sources | External source links | 224 sites failing |
| 4. Ask the question in the heading | Question or claim-shaped H2-H3s | 205 of 413 sites failing |
| 5. Answer first | Answer-first opener | 200 of 413 sites failing |
| 6. Add sameAs links | Entity clarity | 106 sites failing |
| 7. Serve schema in the HTML | JS-only schema (raw HTML against rendered) | JS-only schema on more than half the pages at 23 sites |
| 8. Sign, date and source case studies | Named author plus external source links, on case-study pages | 14 of 570 case studies passed both |
| 9. Publish a comparison page and a how-to-choose page | Follow-up coverage (crawled pages only) | Comparison page found at 192 sites, how-to-choose page at 124 |
Rows 1 to 6 count sites that fail the check on more than half their applicable pages. Rows 7 to 9 count sites or pages, as each cell says.
Add a visible byline and Person schema to every post and guide, and to any page that makes an expert claim. Give the Person sameAs links to the author’s profiles.
Give each author a page marked up as a Person, and point each post’s Article schema at it. Bylines go missing most on service pages, so sign each one with the person who leads it. To check, search a post’s source for “author” and match it to the byline.
Use Article schema on posts, guides and case studies, and Service schema on service pages.
Set it once in each template so new pages inherit it, and give each Service a provider that points to your Organization. Then view-source on one page per template and read the @type values.
Link to the study you’re quoting, the client’s public result or the standard you’re citing. I understand the instinct to keep visitors on the site. But to an AI engine, a page that links nowhere has nothing to cross-check.
Put the links in the body copy, next to the claim, and start with the service pages, where most linked nowhere. To check, count the outbound links in the main content against Rubric’s floor of two.
Swap label headings for the question the section answers. “Our process” becomes “How long does a technical SEO audit take?” Our guide to mapping the follow-up questions AI asks about a topic shows how to find the questions worth asking.
Take the questions from your sales calls, in the buyer’s words. Half the service pages we crawled came back bad on question or claim-shaped headings, and case studies did worse, at 68%. To check, read the headings alone and ask whether a buyer could tell what each section answers.
Answer each heading in the first paragraph under it, in 40 to 70 words. Brand copy can follow.
Rubric’s check uses that window. Read the first 70 words under each heading and ask whether they’d make sense quoted on their own.
Link to LinkedIn, Companies House, your review profiles and Wikidata if you qualify for an entry.
Put one Organization block on the homepage with an @id, and point every other block at it. To check, search the homepage source for “sameAs”.
If a tag manager or a JavaScript framework injects your schema, move it server-side, into the page template. Test with view-source rather than the browser inspector, which shows the page after JavaScript has run. If you audit client shops, the same rule applies to price and stock. Our guide to why AI won’t recommend a product page covers that side.
From a terminal, curl https://www.example.com/services/ should return the JSON-LD in its output.
Sign each case study with the person who led the work, and date the page and the period the work covered. Link at least two sources, one of them something the client could confirm, such as their own announcement or their live site. Then add Article schema with the author pattern below.
To check, view-source for Article and author and count the body links. If a client won’t allow a link, name them and the person who signed off the result.
Write one page comparing your approach with the options a buyer weighs, such as an in-house hire or a freelancer. Write another on how to choose an agency in your speciality, including when you’re the wrong fit.
Ask ChatGPT and Claude the comparison question before and after you publish, and note which pages they cite.
Three short JSON-LD blocks cover it: Person plus Article on a post, Service on a service page, and Organization with sameAs on the homepage. Every name, URL and profile in them is made up, so swap in your own.
A post, with the author as a Person:
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "Person",
"@id": "https://www.example.com/team/jamie-example#person",
"name": "Jamie Example",
"jobTitle": "Head of Technical SEO",
"url": "https://www.example.com/team/jamie-example",
"worksFor": { "@id": "https://www.example.com/#organization" },
"sameAs": ["https://www.linkedin.com/in/jamie-example"]
},
{
"@type": "Article",
"headline": "How long does a technical SEO audit take?",
"datePublished": "2026-10-06",
"author": { "@id": "https://www.example.com/team/jamie-example#person" },
"publisher": { "@id": "https://www.example.com/#organization" },
"mainEntityOfPage": "https://www.example.com/blog/technical-seo-audit-time"
}
]
}A service page:
{
"@context": "https://schema.org",
"@type": "Service",
"name": "Technical SEO audit",
"description": "A 10-day audit of crawling, indexing, site speed and structured data, with a ranked fix list.",
"provider": { "@id": "https://www.example.com/#organization" },
"url": "https://www.example.com/services/technical-seo-audit",
"offers": { "@type": "Offer", "price": "3000", "priceCurrency": "GBP" }
}The homepage, with sameAs:
{
"@context": "https://schema.org",
"@type": "Organization",
"@id": "https://www.example.com/#organization",
"name": "Example Agency",
"url": "https://www.example.com/",
"sameAs": [
"https://www.linkedin.com/company/example-agency",
"https://x.com/exampleagency",
"https://www.youtube.com/@exampleagency",
"https://find-and-update.company-information.service.gov.uk/company/00000000"
]
}Every block points at the same Organization @id, which lets an engine join them into one entity. Use the same Article pattern on case studies. The Offer is optional, and its £3,000 is as made up as everything else. It only shows where a price goes in the markup, so don’t read it as a market rate.
A good one names the company, the person and the scope in its first 50-odd words. The before-and-after below uses the same made-up Example Agency and Jamie Example.
| Before | After | |
|---|---|---|
| Heading | Why choose us | What does a technical SEO audit from Example Agency include? |
| Opener | At Example Agency, we’re a team of passionate search specialists who put clients first. We believe great SEO starts with great relationships, and we’ve been helping brands grow since 2012. | A technical SEO audit from Example Agency takes 10 working days. Jamie Example, our head of technical SEO, checks crawling, indexing, site speed and structured data on up to 5,000 URLs. You get a ranked list of fixes, each tied to the pages it affects, and a 60-minute call to go through it. |
The after version answers a buyer’s question in 53 words with three figures. It’s built on four habits any service page can copy:
Fix the templates first, then the pages a buyer reads, using only the fixes above.
| Week | Work (fix number) |
|---|---|
| 1 | Move schema into the server HTML (7). Add Organization schema with sameAs to the homepage (6). Build author pages with Person schema (1). |
| 2 | Add Article schema to the post and case-study templates and Service schema to the service template (2). Add bylines to posts and guides (1). |
| 3 | Give your two busiest service pages question headings, answer-first openers, two sources and a named lead (1, 3, 4, 5). Rewrite the busier one with the four habits above. |
| 4 | Bring your five newest case studies up to standard (8). Publish a comparison page and a how-to-choose page (9). Re-crawl and read the inner pages first. |
I haven’t put a score gain against any week, because it depends on how many pages each fix reaches.
Check one service page and one case study, at about two minutes a step.
curl -A "GPTBot" https://www.example.com/services/, and see whether you get the page or a challenge. A 403 isn’t proof the real bot is blocked, since your request doesn’t come from its network. A challenge page is still worth raising.To see your own inner pages against these checks, run a Rubric audit and read the page-level results before the headline score.
This benchmark can’t tell you how often any of these sites is quoted, or why a site blocks what it blocks. It’s first-party crawl data from convenience lists, scored on one engine at one page cap on one day.
Every column comes from the same 6 October 2026 crawl, on one engine at one page cap, so the columns can be compared. Fourteen agencies sit on more than one list, so the list columns overlap. Engine fits are Rubric’s modelled weightings.
| All agencies | 12-13 Aug list | 23 Aug English-speaking | 23 Aug international | Prolific North Top 50 | 10 UK SEO agencies | SEO-tool vendors | |
|---|---|---|---|---|---|---|---|
| Attempted | 484 | 243 | 100 | 99 | 47 | 10 | 49 |
| Scored | 415 (85.7%) | 200 | 89 | 89 | 42 | 10 | 47 |
| Median | 70 | 71 | 71 | 69 | 65.5 | 68 | 67 |
| Mean | 69.7 | 70.5 | 70.8 | 68.6 | 66.6 | 69.9 | 67.5 |
| IQR | 65-75 | 66-75 | 66.5-77 | 65-73 | 64-70 | 65.5-75.75 | 63-70 |
| Min / max | 46 / 88 | 47 / 88 | 46 / 87 | 55 / 84 | 54 / 80 | 64 / 80 | 53 / 86 |
| Scoring 70+ | 222 (53.5%) | 116 (58.0%) | 58 (65.2%) | 40 (44.9%) | 12 (28.6%) | 3 (30.0%) | 13 (27.7%) |
| Known median (share at 70+) | 80 (81.7%) | 80 (84.0%) | 83 (84.3%) | 80 (82.0%) | 74.5 (66.7%) | 80.5 (80.0%) | 73 (68.1%) |
| Findable median (share at 70+) | 71 (60.5%) | 71 (61.0%) | 72 (64.0%) | 70 (56.2%) | 70.5 (61.9%) | 71.5 (60.0%) | 72 (66.0%) |
| Trusted median (share at 70+) | 52 (14.7%) | 53.5 (20.5%) | 54 (11.2%) | 51 (11.2%) | 48 (4.8%) | 46.5 (10.0%) | 51 (10.6%) |
| Trusted under 50 | 168 | 77 | 30 | 39 | 22 | 7 | 21 |
| ChatGPT fit | 67 | 69 | 69 | 65 | 65 | 64 | 65 |
| Perplexity fit | 68 | 69 | 69 | 68 | 66.5 | 66.5 | 69 |
| AI Overviews fit | 66 | 66 | 68 | 65 | 61 | 62.5 | 60 |
| Gemini fit | 66 | 67 | 70 | 66 | 60 | 63 | 55 |
| Copilot fit | 68 | 68 | 69 | 68 | 65 | 65 | 63 |
| Claude fit | 70 | 71 | 70 | 69 | 66.5 | 68 | 66 |
| Bands: 90-100 / 80-89 / 70-79 / 60-69 / 50-59 / under 50 | 0/28/194/161/29/3 | 0/17/99/69/14/1 | 0/9/49/23/6/2 | 0/2/38/43/6/0 | 0/2/10/27/3/0 | 0/1/2/7/0/0 | 0/3/10/31/3/0 |
Country comes from the 12-13 August list’s country label where it had one, and otherwise from the domain ending. Groups are shown where 10 or more agencies were scored, and the other 44 scored agencies sit in smaller groups. Read it as description only.
| Country | Sites | Median | Known | Findable | Trusted |
|---|---|---|---|---|---|
| Australia | 36 | 73 | 84.5 | 73 | 57 |
| USA | 43 | 71 | 81 | 70 | 51 |
| Germany and Austria | 22 | 71 | 75.5 | 71 | 58 |
| Canada | 34 | 70.5 | 80.5 | 70 | 54.5 |
| France | 18 | 69.5 | 76 | 71 | 50.5 |
| Nordics | 10 | 69.5 | 85 | 71 | 50 |
| UK | 89 | 69 | 77 | 71 | 51 |
| Spain | 17 | 69 | 83 | 71 | 54 |
| South Africa | 10 | 63.5 | 78 | 66 | 57 |
| Generic domain, no country | 92 | 71 | 80 | 72 | 51.5 |
Page-level medians, scored agencies only (n=415, 15,658 pages). Page types come from the URL first, then the crawler’s own label, and “other page” is a mixed bucket. The homepage row includes subdomain roots, so its median differs from the site-level homepage median of 82.
| Page type | Pages | Share of pages | Sites with 1+ | Median page score | Known | Findable | Trusted | Median words | Median external links |
|---|---|---|---|---|---|---|---|---|---|
| Other page | 7,101 | 45.4% | 388 | 68 | 75 | 71 | 48 | 798 | 0 |
| Blog post or guide | 4,151 | 26.5% | 292 | 75 | 90 | 71 | 66 | 995 | 1 |
| About, team, contact | 1,131 | 7.2% | 347 | 66 | 77 | 69 | 48 | 429 | 1 |
| Listing or index | 665 | 4.2% | 329 | 78 | 89 | 74 | 0 | 398 | 0 |
| Service page | 646 | 4.1% | 175 | 70 | 77 | 74 | 46 | 884 | 0 |
| Legal and utility | 592 | 3.8% | 272 | 65 | 70 | 70 | 48 | 1,124.5 | 0 |
| Case study | 570 | 3.6% | 157 | 68 | 77 | 71 | 45 | 513 | 0 |
| Homepage | 537 | 3.4% | 415 | 80 | 86 | 80 | 67 | 895 | 1 |
| Product page | 265 | 1.7% | 33 | 81 | 82 | 79 | 89 | 2,405 | 2 |
Share of applicable pages marked bad, by page type
| Check | Blog post or guide | Service page | Case study | About, team, contact | Other page |
|---|---|---|---|---|---|
| Named author | 9% | 82% | 78% | 81% | 78% |
| Content-type schema | not reported | 61% | 60% | 65% | 69% |
| Two or more external source links | 44% | 60% | 61% | 48% | 53% |
| Answer-first opener | 52% | 52% | 46% | 56% | 51% |
| Question or claim-shaped headings | 43% | 50% | 68% | 68% | 50% |
| Statistics carry a source | 28% | 27% | 49% | 31% | 33% |
| Statistic density | 26% | 43% | 16% | 25% | 30% |
| Entity clarity | 4% | 29% | 35% | 32% | 35% |
Signed, sourced, dated and sized: blog posts, service pages and case studies
| Measure | Blog post or guide | Service page | Case study |
|---|---|---|---|
| Pages | 4,151 | 646 | 570 |
| Named author and 2+ external links | 1,691 (40.7%), median 81 | 14 (2.2%), median 75.5 | 14 (2.5%), median 77.5 |
| Neither | 195 (4.7%), median 67 | 313 (48.5%), median 67 | 261 (45.8%), median 64 |
| Passed named author (applicable pages) | 3,489 of 4,147 (84.1%) | 28 of 600 (4.7%) | 77 of 554 (13.9%) |
| No external link | 1,808 (43.6%) | 379 (58.7%) | 347 (60.9%) |
| Under 300 words | 1,040 (25.1%) | 72 (11.1%) | 140 (24.6%) |
Part of each median gap is mechanical, because both checks feed the page score. Blog-post schema figures are left out because they’d be partly circular.
Sites with the type on at least one crawled page.
| Type | Agencies (n=415) |
|---|---|
| Organization (or a subtype) | 366 (88.2%) |
| WebSite | 333 (80.2%) |
| BreadcrumbList | 310 (74.7%) |
| Person | 285 (68.7%) |
| Article family | 261 (62.9%) |
| FAQPage | 194 (46.7%) |
| Service or ProfessionalService | 162 (39.0%) |
| Service as its own type | 139 (33.5%) |
| Offer | 114 (27.5%) |
| Review or AggregateRating | 106 (25.5%) |
| VideoObject | 65 (15.7%) |
| HowTo | 15 (3.6%) |
| No schema on any crawled page | 21 (5.1%) |
| Page-level measure | Agencies |
|---|---|
| Service pages declaring Service | 130 of 646 (20.1%) |
| Case studies declaring an Article-family type | 83 of 570 (14.6%) |
We don’t report how many blog posts declare Article schema, because the crawler labels pages with Article schema as articles.
The top quarter is the 104 agencies scoring 75 to 88 and the bottom quarter the 104 scoring 46 to 65.
| Median | Top quarter (n=104) | Bottom quarter (n=104) | All agencies (n=415) |
|---|---|---|---|
| Overall | 78 | 62 | 70 |
| Known | 88 | 68 | 80 |
| Findable | 77 | 64 | 71 |
| Trusted | 67.5 | 43 | 52 |
| Homepage | 89 | 76 | 82 |
| Inner pages | 78 | 61 | 70 |
Sites failing each check (bad on more than half the applicable pages). These are the 12 checks with the widest gap between the two quarters. HTTP status, robots.txt, canonical, internal links, reachability and word count failed at no agency in the crawl. The gaps are correlational, and partly mechanical, because each check feeds the score.
| Check | Top quarter failing | Bottom quarter failing |
|---|---|---|
| Question or claim-shaped H2-H3s | 18 of 102 | 85 of 104 |
| Content-type schema | 22 of 104 | 76 of 104 |
| Entity clarity (Organization or Person plus sameAs) | 6 of 104 | 57 of 104 |
| Two or more external source links | 26 of 104 | 77 of 104 |
| Answer-first opener (40-70 words) | 29 of 102 | 72 of 104 |
| Named author | 35 of 102 | 74 of 104 |
| Fresh (updated in the last 12 months) | 11 of 97 | 34 of 84 |
| Statistic density (1.5 or more per 100 words) | 3 of 104 | 32 of 104 |
| Distinct content (no near-duplicates) | 1 of 103 | 25 of 103 |
| Exactly one H1 | 4 of 104 | 27 of 104 |
| Statistics carry a source | 16 of 102 | 34 of 104 |
| Meta description (50-160 characters) | 3 of 104 | 20 of 104 |
Chris McCarron founded GoGoChimp in Glasgow in 2013 and has 13 years of conversion rate optimisation experience. He was previously Head of Growth at Roadtrippers and FOMO, and he wrote CITED, a book on getting your business recommended by AI search.
Rubric is GoGoChimp’s AI-citability auditor. It crawls a site the way AI crawlers read it and runs 44 checks across three pillars, Known, Findable and Trusted, weighted for six AI engines: ChatGPT, Perplexity, Google AI Overviews, Gemini, Microsoft Copilot and Claude.
Rubric predicts citability, it doesn’t count citations. 70 is the quotable line.
Run Rubric on your own site and compare your pillar scores and failing checks with the figures in this study.
Rubric
Free AI SEO audit
The Rubric audit is free and needs no sign-up.
Your website
Run a free Rubric auditRubric
yourstore.com · 26 pages
Citability score
67
3 points below quotable
70 is the quotable line.
By engine
Perplexity
76
Claude
74
ChatGPT
65
Copilot
63
AI Overviews
61
Gemini
57
Do these first
Add Organization schema with sameAs
Name an author on every article