{"id":56,"date":"2026-04-21T16:06:40","date_gmt":"2026-04-21T16:06:40","guid":{"rendered":"https:\/\/serp.fyi\/blog\/?p=56"},"modified":"2026-04-21T16:06:40","modified_gmt":"2026-04-21T16:06:40","slug":"observatory-notes-what-we-learned-from-running-1200-audits","status":"publish","type":"post","link":"https:\/\/serp.fyi\/blog\/observatory-notes-what-we-learned-from-running-1200-audits\/","title":{"rendered":"Observatory Notes: What We Learned From Running 1,200 Audits"},"content":{"rendered":"\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>KEY FINDINGS AT A GLANCE<\/strong><\/td><\/tr><tr><td>Only 22% of audited sites have both clean crawler access and accurate AI citation. Most fail on one or the other.Structured data triples citation accuracy &nbsp;but only when combined with open crawl access. Either alone is weak.The gap between top-quartile and bottom-quartile AI visibility is wider than the equivalent gap in organic search.B2B SaaS sites are the most underoptimised category &nbsp;high domain authority, terrible AI signals.Sites updated in the last 90 days receive 47% more AI citations than stale sites on the same topic.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">About this dataset<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The Observatory is our internal research programme &nbsp;a rolling audit of domains across industry verticals, measuring AI citation frequency, brand representation accuracy, crawler access health, and structured data quality. We started it quietly in April 2025. This is our first public report.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The 1,200 domains in this cohort were selected to represent a spread of sizes, verticals, and technical sophistication &nbsp;not a convenience sample of our paying customers. We re-audited each domain three times across the year to capture drift.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><td><strong>Metric<\/strong><\/td><td><strong>Value<\/strong><\/td><\/tr><\/thead><tbody><tr><td>Domains audited<\/td><td><strong>1,200<\/strong><\/td><\/tr><tr><td>Re-audit cycles<\/td><td><strong>3 (Apr, Sep 2025 \u00b7 Apr 2026)<\/strong><\/td><\/tr><tr><td>AI engines tested<\/td><td><strong>ChatGPT Search, Perplexity, Gemini, Bing Copilot<\/strong><\/td><\/tr><tr><td>Queries per domain<\/td><td><strong>40\u201360 (brand + category)<\/strong><\/td><\/tr><tr><td>Industries covered<\/td><td><strong>B2B SaaS, media, e-commerce, local services, fintech, health<\/strong><\/td><\/tr><tr><td>Total data points<\/td><td><strong>~3.2 million<\/strong><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Methodology note: <\/strong>AI citation accuracy was scored by human reviewers using a 5-point rubric: correct identification, accurate product description, correct pricing tier, no hallucinated features, and no brand confusion with a competitor. A site scores 100 only if all five criteria are met across 80% of tested queries.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">The 8 findings<\/h2>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 01<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Most sites fail the \u201cboth\u201d test<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We expected to find sites that were either well-optimised or poorly-optimised. What we found instead was a distribution with a hollow middle: sites that scored well on technical access but poorly on citation accuracy, and vice versa. Only 22% of audited sites passed both dimensions.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>61%<\/strong> Open crawler access, no major blocks<\/td><td><strong>38%<\/strong> AI describes brand correctly<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>PASS RATE &nbsp;BOTH ACCESS AND ACCURACY<\/strong> <strong>22%<\/strong> The real benchmark. Access is necessary but not sufficient &nbsp;most sites that fix their robots.txt still have accuracy problems.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This is the most important pattern in the dataset. Access and accuracy are correlated but not the same thing. You can unblock every AI crawler and still have AI systems describe your product incorrectly, confuse you with a competitor, or cite an outdated pricing page. The pipeline matters at every stage.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 02<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Structured data is a force multiplier, not a standalone fix<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Average AI citation accuracy score by technical configuration<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>serp.fyi Observatory data &nbsp;1,200 domains. Score = % of queries with accurate brand representation across all 5 criteria.<\/em><\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><td><strong>Configuration<\/strong><\/td><td><strong>Score<\/strong><\/td><\/tr><\/thead><tbody><tr><td>Open access + schema + llms.txt<\/td><td><strong>91 \/ 100<\/strong><\/td><\/tr><tr><td>Open access + schema<\/td><td><strong>74 \/ 100<\/strong><\/td><\/tr><tr><td>Open access only<\/td><td><strong>52 \/ 100<\/strong><\/td><\/tr><tr><td>Schema, but crawlers blocked<\/td><td><strong>31 \/ 100<\/strong><\/td><\/tr><tr><td>No optimisation<\/td><td><strong>18 \/ 100<\/strong><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Schema markup alone &nbsp;with AI crawlers blocked &nbsp;scores 31\/100. Barely better than nothing. But paired with open access, schema jumps from being a marginal signal to a primary driver of accuracy. The combination of open access, schema, and an llms.txt file scores 91\/100 &nbsp;the highest configuration in our dataset.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 03<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Freshness is the most underrated signal<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We expected domain authority and backlink profile to be the dominant factors in AI citation frequency. They matter &nbsp;but not as much as freshness. Sites with a substantive content update in the last 90 days received 47% more citations than comparable sites that hadn\u2019t been updated. This gap was consistent across all four AI engines.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>47%<\/strong> more citations for sites updated in last 90 days<\/td><td><strong>3.1\u00d7<\/strong> citation lift from full optimisation vs nothing<\/td><td><strong>61%<\/strong> of citation drops tied to content staleness<\/td><td><strong>22%<\/strong> of sites pass both access and accuracy<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Why freshness matters to AI: <\/strong>Retrieval-augmented systems weight recency explicitly. A page last crawled 14 months ago carries an implicit staleness penalty &nbsp;even if the content hasn\u2019t technically changed. Regular publishing signals that a source is actively maintained, which correlates with reliability in these systems\u2019 training.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 04<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">The visibility gap is wider than in organic search<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In traditional organic search, the gap between the top quartile and bottom quartile site for a given keyword set is typically 3\u20135 positions &nbsp;meaningful but bounded. In AI visibility, the gap is binary in many cases: you\u2019re cited or you\u2019re not. Among sites competing for the same category queries, the top quartile received an average of 8.3 citations per 10 queries. The bottom quartile received 0.4.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Citation frequency by visibility quartile &nbsp;same-category competing sites<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Average citations per 10 queries. Sites grouped by performance quartile within the same industry\/category.<\/em><\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><td><strong>Configuration<\/strong><\/td><td><strong>Score<\/strong><\/td><\/tr><\/thead><tbody><tr><td>Top quartile (Q4)<\/td><td><strong>8.3 \/ 10<\/strong><\/td><\/tr><tr><td>Upper-mid (Q3)<\/td><td><strong>5.4 \/ 10<\/strong><\/td><\/tr><tr><td>Lower-mid (Q2)<\/td><td><strong>2.6 \/ 10<\/strong><\/td><\/tr><tr><td>Bottom quartile (Q1)<\/td><td><strong>0.4 \/ 10<\/strong><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 05<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">B2B SaaS is the most underoptimised vertical<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">SaaS companies tend to have strong domain authority, good technical infrastructure, and active content programmes. Yet they scored lowest on AI citation accuracy among all verticals we tracked &nbsp;an average of 41\/100 versus 68\/100 for media sites and 59\/100 for e-commerce.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The pattern is consistent: SaaS sites have feature-dense, jargon-heavy content written for existing users and keyword targeting, not for an AI system trying to quickly understand what the product does and who it\u2019s for. Their robots.txt files also tend to have aggressive blocks added by legal or security teams and never revisited.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Average AI citation accuracy score by industry vertical<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>serp.fyi Observatory data, n=1,200. Score out of 100.<\/em><\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><td><strong>Configuration<\/strong><\/td><td><strong>Score<\/strong><\/td><\/tr><\/thead><tbody><tr><td>Media \/ publishing<\/td><td><strong>68 \/ 100<\/strong><\/td><\/tr><tr><td>E-commerce<\/td><td><strong>59 \/ 100<\/strong><\/td><\/tr><tr><td>Local services<\/td><td><strong>54 \/ 100<\/strong><\/td><\/tr><tr><td>Fintech<\/td><td><strong>48 \/ 100<\/strong><\/td><\/tr><tr><td>Health<\/td><td><strong>44 \/ 100<\/strong><\/td><\/tr><tr><td>B2B SaaS<\/td><td><strong>41 \/ 100<\/strong><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 06<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Brand confusion is the most common accuracy failure<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">When AI systems got things wrong, the most common failure mode wasn\u2019t hallucination of non-existent features &nbsp;it was brand confusion. In 34% of inaccurate citations, the AI correctly described a product or service but attributed it to the wrong company, or conflated two similar brands in the same niche.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is particularly acute for companies with generic names, companies in crowded SaaS categories, and any brand that shares significant keyword overlap with a larger competitor. The fix is explicit disambiguation in your llms.txt and Organisation schema &nbsp;a \u201cwe are not X\u201d statement is as valuable as a \u201cwe are Y\u201d statement.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 07<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Perplexity is the most citation-generous engine; Gemini the most conservative<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Not all AI engines behave the same way. Perplexity cited sources in 74% of its answers in our test set. Gemini cited sources in 41%. ChatGPT Search and Bing Copilot fell between at 62% and 58% respectively. If you\u2019re tracking your AI visibility, your Perplexity score and your Gemini score will look very different &nbsp;and the optimisation strategies aren\u2019t identical.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><em>Finding 08<\/em><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Fixing access takes a day. Fixing accuracy takes three months.<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This is perhaps the most practically useful finding. A robots.txt fix deploys in minutes and AI crawlers respect it within 1\u20132 weeks. Accuracy improvements &nbsp;from new structured data, a fresh llms.txt, and updated content &nbsp;take 8\u201312 weeks to fully propagate through the re-crawl and re-indexing cycles of the major AI engines. Plan accordingly.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>The takeaway: <\/strong>Start with access fixes today. Start the accuracy work this week. Measure both separately on a 90-day cycle. Don\u2019t conflate them.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">What surprised us<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We expected domain authority to be the dominant predictor of AI citation frequency. It isn\u2019t &nbsp;it\u2019s fourth, behind freshness, structured data completeness, and crawl access. Brands with DA 40 and excellent AI signals regularly outperformed DA 80 sites with poor ones.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We also didn\u2019t expect the B2B SaaS finding to be as stark as it is. SaaS companies spend more on content marketing than almost any other vertical. The problem isn\u2019t effort &nbsp;it\u2019s that the content is optimised for a different system entirely. The same budget that produces 40 keyword-targeted blog posts could produce 8 deeply accurate, AI-legible resource pages and see better returns in this emerging channel.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">FAQ<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Will you publish the full dataset?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We\u2019re publishing anonymised aggregate data in this report and will release industry-level breakdowns quarterly. Individual domain data stays private. Researchers can request access to the anonymised microdata via our research programme.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How were domains selected? Is this representative?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Domains were stratified by industry vertical and size (traffic tier), then randomly sampled within strata. It\u2019s not a perfectly representative sample of the entire web &nbsp;we over-indexed on English-language, professionally managed sites &nbsp;but it\u2019s the largest systematic AI visibility dataset we\u2019re aware of at this stage.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How do I know where my site sits in this distribution?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Run an audit. The Observatory benchmarks are the reference set we use for comparative scoring &nbsp;your audit report shows your score against the vertical median, the top quartile, and the overall average. You\u2019ll see exactly which dimension is pulling your score down.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>When is the next Observatory report?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We publish full cohort reports twice a year April and October. Industry-specific briefings (SaaS, media, health) go out quarterly to subscribers. Sign up at serp.fyi to be notified.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-cyan-bluish-gray-background-color has-background has-fixed-layout\"><tbody><tr><td class=\"has-text-align-center\" data-align=\"center\"><strong>See where your domain sits in the Observatory<\/strong> <br>Citation accuracy, crawl access score, structured data health, and vertical comparison \u00a0in one audit report. <br><strong>\u2192 Run your own audit at <a href=\"http:\/\/serp.fyi\">serp.fyi\/register<\/a><\/strong><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>KEY FINDINGS AT A GLANCE Only 22% of audited sites have both clean crawler access and accurate AI citation. Most fail on one or the other.Structured data triples citation accuracy &nbsp;but only when combined with open crawl access. Either alone is weak.The gap between top-quartile and bottom-quartile AI visibility is wider than the equivalent gap<\/p>\n","protected":false},"author":1,"featured_media":57,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[98,93,96,92,95,97,94],"class_list":["post-56","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tool-guides-tutorials","tag-content-audit-findings","tag-data-driven-seo","tag-seo-audit-insights","tag-seo-case-study","tag-seo-patterns-trends","tag-technical-seo-insights","tag-website-audit-analysis"],"_links":{"self":[{"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/posts\/56","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/comments?post=56"}],"version-history":[{"count":1,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/posts\/56\/revisions"}],"predecessor-version":[{"id":58,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/posts\/56\/revisions\/58"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/media\/57"}],"wp:attachment":[{"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/media?parent=56"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/categories?post=56"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/serp.fyi\/blog\/wp-json\/wp\/v2\/tags?post=56"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}