| KEY FINDINGS AT A GLANCE |
| Only 22% of audited sites have both clean crawler access and accurate AI citation. Most fail on one or the other.Structured data triples citation accuracy but only when combined with open crawl access. Either alone is weak.The gap between top-quartile and bottom-quartile AI visibility is wider than the equivalent gap in organic search.B2B SaaS sites are the most underoptimised category high domain authority, terrible AI signals.Sites updated in the last 90 days receive 47% more AI citations than stale sites on the same topic. |
About this dataset
The Observatory is our internal research programme a rolling audit of domains across industry verticals, measuring AI citation frequency, brand representation accuracy, crawler access health, and structured data quality. We started it quietly in April 2025. This is our first public report.
The 1,200 domains in this cohort were selected to represent a spread of sizes, verticals, and technical sophistication not a convenience sample of our paying customers. We re-audited each domain three times across the year to capture drift.
| Metric | Value |
| Domains audited | 1,200 |
| Re-audit cycles | 3 (Apr, Sep 2025 · Apr 2026) |
| AI engines tested | ChatGPT Search, Perplexity, Gemini, Bing Copilot |
| Queries per domain | 40–60 (brand + category) |
| Industries covered | B2B SaaS, media, e-commerce, local services, fintech, health |
| Total data points | ~3.2 million |
| Methodology note: AI citation accuracy was scored by human reviewers using a 5-point rubric: correct identification, accurate product description, correct pricing tier, no hallucinated features, and no brand confusion with a competitor. A site scores 100 only if all five criteria are met across 80% of tested queries. |
The 8 findings
Finding 01
Most sites fail the “both” test
We expected to find sites that were either well-optimised or poorly-optimised. What we found instead was a distribution with a hollow middle: sites that scored well on technical access but poorly on citation accuracy, and vice versa. Only 22% of audited sites passed both dimensions.
| 61% Open crawler access, no major blocks | 38% AI describes brand correctly |
| PASS RATE BOTH ACCESS AND ACCURACY 22% The real benchmark. Access is necessary but not sufficient most sites that fix their robots.txt still have accuracy problems. |
This is the most important pattern in the dataset. Access and accuracy are correlated but not the same thing. You can unblock every AI crawler and still have AI systems describe your product incorrectly, confuse you with a competitor, or cite an outdated pricing page. The pipeline matters at every stage.
Finding 02
Structured data is a force multiplier, not a standalone fix
Average AI citation accuracy score by technical configuration
serp.fyi Observatory data 1,200 domains. Score = % of queries with accurate brand representation across all 5 criteria.
| Configuration | Score |
| Open access + schema + llms.txt | 91 / 100 |
| Open access + schema | 74 / 100 |
| Open access only | 52 / 100 |
| Schema, but crawlers blocked | 31 / 100 |
| No optimisation | 18 / 100 |
Schema markup alone with AI crawlers blocked scores 31/100. Barely better than nothing. But paired with open access, schema jumps from being a marginal signal to a primary driver of accuracy. The combination of open access, schema, and an llms.txt file scores 91/100 the highest configuration in our dataset.
Finding 03
Freshness is the most underrated signal
We expected domain authority and backlink profile to be the dominant factors in AI citation frequency. They matter but not as much as freshness. Sites with a substantive content update in the last 90 days received 47% more citations than comparable sites that hadn’t been updated. This gap was consistent across all four AI engines.
| 47% more citations for sites updated in last 90 days | 3.1× citation lift from full optimisation vs nothing | 61% of citation drops tied to content staleness | 22% of sites pass both access and accuracy |
| Why freshness matters to AI: Retrieval-augmented systems weight recency explicitly. A page last crawled 14 months ago carries an implicit staleness penalty even if the content hasn’t technically changed. Regular publishing signals that a source is actively maintained, which correlates with reliability in these systems’ training. |
Finding 04
The visibility gap is wider than in organic search
In traditional organic search, the gap between the top quartile and bottom quartile site for a given keyword set is typically 3–5 positions meaningful but bounded. In AI visibility, the gap is binary in many cases: you’re cited or you’re not. Among sites competing for the same category queries, the top quartile received an average of 8.3 citations per 10 queries. The bottom quartile received 0.4.
Citation frequency by visibility quartile same-category competing sites
Average citations per 10 queries. Sites grouped by performance quartile within the same industry/category.
| Configuration | Score |
| Top quartile (Q4) | 8.3 / 10 |
| Upper-mid (Q3) | 5.4 / 10 |
| Lower-mid (Q2) | 2.6 / 10 |
| Bottom quartile (Q1) | 0.4 / 10 |
Finding 05
B2B SaaS is the most underoptimised vertical
SaaS companies tend to have strong domain authority, good technical infrastructure, and active content programmes. Yet they scored lowest on AI citation accuracy among all verticals we tracked an average of 41/100 versus 68/100 for media sites and 59/100 for e-commerce.
The pattern is consistent: SaaS sites have feature-dense, jargon-heavy content written for existing users and keyword targeting, not for an AI system trying to quickly understand what the product does and who it’s for. Their robots.txt files also tend to have aggressive blocks added by legal or security teams and never revisited.
Average AI citation accuracy score by industry vertical
serp.fyi Observatory data, n=1,200. Score out of 100.
| Configuration | Score |
| Media / publishing | 68 / 100 |
| E-commerce | 59 / 100 |
| Local services | 54 / 100 |
| Fintech | 48 / 100 |
| Health | 44 / 100 |
| B2B SaaS | 41 / 100 |
Finding 06
Brand confusion is the most common accuracy failure
When AI systems got things wrong, the most common failure mode wasn’t hallucination of non-existent features it was brand confusion. In 34% of inaccurate citations, the AI correctly described a product or service but attributed it to the wrong company, or conflated two similar brands in the same niche.
This is particularly acute for companies with generic names, companies in crowded SaaS categories, and any brand that shares significant keyword overlap with a larger competitor. The fix is explicit disambiguation in your llms.txt and Organisation schema a “we are not X” statement is as valuable as a “we are Y” statement.
Finding 07
Perplexity is the most citation-generous engine; Gemini the most conservative
Not all AI engines behave the same way. Perplexity cited sources in 74% of its answers in our test set. Gemini cited sources in 41%. ChatGPT Search and Bing Copilot fell between at 62% and 58% respectively. If you’re tracking your AI visibility, your Perplexity score and your Gemini score will look very different and the optimisation strategies aren’t identical.
Finding 08
Fixing access takes a day. Fixing accuracy takes three months.
This is perhaps the most practically useful finding. A robots.txt fix deploys in minutes and AI crawlers respect it within 1–2 weeks. Accuracy improvements from new structured data, a fresh llms.txt, and updated content take 8–12 weeks to fully propagate through the re-crawl and re-indexing cycles of the major AI engines. Plan accordingly.
| The takeaway: Start with access fixes today. Start the accuracy work this week. Measure both separately on a 90-day cycle. Don’t conflate them. |
What surprised us
We expected domain authority to be the dominant predictor of AI citation frequency. It isn’t it’s fourth, behind freshness, structured data completeness, and crawl access. Brands with DA 40 and excellent AI signals regularly outperformed DA 80 sites with poor ones.
We also didn’t expect the B2B SaaS finding to be as stark as it is. SaaS companies spend more on content marketing than almost any other vertical. The problem isn’t effort it’s that the content is optimised for a different system entirely. The same budget that produces 40 keyword-targeted blog posts could produce 8 deeply accurate, AI-legible resource pages and see better returns in this emerging channel.
FAQ
Will you publish the full dataset?
We’re publishing anonymised aggregate data in this report and will release industry-level breakdowns quarterly. Individual domain data stays private. Researchers can request access to the anonymised microdata via our research programme.
How were domains selected? Is this representative?
Domains were stratified by industry vertical and size (traffic tier), then randomly sampled within strata. It’s not a perfectly representative sample of the entire web we over-indexed on English-language, professionally managed sites but it’s the largest systematic AI visibility dataset we’re aware of at this stage.
How do I know where my site sits in this distribution?
Run an audit. The Observatory benchmarks are the reference set we use for comparative scoring your audit report shows your score against the vertical median, the top quartile, and the overall average. You’ll see exactly which dimension is pulling your score down.
When is the next Observatory report?
We publish full cohort reports twice a year April and October. Industry-specific briefings (SaaS, media, health) go out quarterly to subscribers. Sign up at serp.fyi to be notified.
| See where your domain sits in the Observatory Citation accuracy, crawl access score, structured data health, and vertical comparison in one audit report. → Run your own audit at serp.fyi/register |






