Research · Machine Gatekeeper

How the gatekeeper actually decides.

The Machine Gatekeeper framework says AI systems now sit between a brand and its audience, deciding who gets named. This page is the evidence for the mechanism: what a 129,000-domain study found actually correlates with getting cited, what Google itself says about optimizing for it, and why a large share of sites accidentally lock the gatekeeper out entirely.

Last verified 20 August 2026

In one paragraph

Independent research on 129,000 domains found citation frequency tracks closely with referring domains, domain trust, content freshness, and depth, ordinary authority signals, not hidden AI tricks. Google's own documentation says the same thing in its own words: there is no separate optimization layer for AI features, the core search quality systems are what generative answers draw on. And Cloudflare's own 2025 traffic data shows AI crawlers are growing fast while robots.txt files disallow them more than any other bot category, meaning a meaningful share of sites are shutting the gatekeeper out without realizing it.

What The Gatekeeper Weighs

Ordinary authority signals, not tricks.

SE Ranking analyzed 216,524 pages across 129,000 domains and 20 industry niches to see what actually separated frequently-cited sites from rarely-cited ones in ChatGPT. Four factors showed the clearest gap between the weakest and strongest tier, shown here as average citations per page.

Domain Trust score1.6 vs 8.4 avg. citations
Below 43
1.6
97–100
8.4
Referring domains1.7 vs 8.4 avg. citations
Up to 2,500
1.7
350,000+
8.4
Content freshness3.6 vs 6.0 avg. citations
Outdated
3.6
< 3 months
6.0
Content depth3.2 vs 5.1 avg. citations
< 800 words
3.2
2,900+ words
5.1

Domain Trust and referring domains showed the strongest relationship of anything tested, both proxies for Earned Authority: what independent sites already link to and vouch for. Page speed also correlated (pages loading in under 0.4 seconds averaged 6.7 citations), a Clarity-adjacent signal, the gatekeeper can only weigh what it can render quickly. Source: SE Ranking, 129,000 domains, published 26 November 2025.

In Google's Own Words

There is no separate game to learn.

Every AI-visibility vendor sells a version of "AI optimization" as something distinct from SEO. Google's own developer documentation says otherwise, directly.

The best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems.

Google, AI features and your website, developers.google.com · Read the source

This is consistent with the SE Ranking data above: the strongest citation signals, Domain Trust and referring domains, are exactly the same signals traditional SEO and PR have always tracked. The gatekeeper is not running a hidden second scoring system. It is applying the same quality bar it always has, to a different output format.

What Stays Hidden, Even From Google

You can see that you were shown. Not why.

In June 2026, Google added a Generative AI tab to Search Console, the first official window into AI Overviews and AI Mode performance. What it shows, and deliberately withholds, says a lot about how opaque the gatekeeper's logic still is, even to the people building it.

What it shows

Impressions: every time a URL from your domain is rendered inside an AI-generated answer block, filterable by URL, region, device, and date range.

What it withholds

Click-through data, the actual prompts or queries that triggered the citation, and position within the answer. Google cites user privacy and interface volatility as the reasons.

In short: you can now confirm the gatekeeper noticed you. You still cannot see which question earned that notice, or why one page was chosen over another. Source: Google Search Console Generative AI reporting, launched June 2026.

The Self-Inflicted Wound

The gatekeeper cannot recognize what it cannot reach.

Cloudflare's own network data shows AI crawlers growing fast, and, at the same time, robots.txt files disallowing them more often than any other category of bot. Two facts sitting side by side: the traffic that matters most right now is also the traffic most commonly blocked, frequently by accident.

Growing fast

ChatGPT-User request volume peaked as much as 16x higher than where it started the year; PerplexityBot traffic finished the year roughly 3.5x higher; ClaudeBot crawling volume nearly doubled at its peak.

Blocked most often

Cloudflare's own review of 2025 traffic states plainly that AI crawlers were "the most frequently fully disallowed user agents found in robots.txt files" of any bot category it tracks.

Source: The 2025 Cloudflare Radar Year in Review, published December 2025.

CrawlerBelongs toRecommendation
GPTBot / OAI-SearchBotOpenAIAllow, don't disallow
ClaudeBotAnthropicAllow, don't disallow
PerplexityBotPerplexityAllow, don't disallow
Google-ExtendedGoogleAllow, don't disallow

This page covers the selection mechanism specifically. For the wider evidence base behind the Recognition Triangle, including who gets cited and how concentrated that citation share actually is, see the main research hub.

Back to the research hub
Common Questions

Questions people ask about this page.

Is there really no separate way to optimize for AI answers?
Per Google's own documentation, no. AI features on Google Search draw on the same core ranking and quality systems as regular search. That does not mean nothing changes; it means the fundamentals, authority, clarity, trust, freshness, stay the fundamentals, rather than being replaced by a new hackable layer.
Does this mean schema markup and llms.txt don't matter?
Schema and llms.txt help machines disambiguate who you are once they have already decided to look at you. They are infrastructure, not the reason you get selected in the first place. The signals shown on this page, authority, freshness, depth, are what appear to influence selection itself.
How would I know if my own site is blocking AI crawlers?
Check your robots.txt file for Disallow rules naming GPTBot, ClaudeBot, PerplexityBot, Google-Extended, or OAI-SearchBot. Many site builders and security plugins add these rules by default, which is exactly why Cloudflare found AI crawlers blocked more often than any other bot category, frequently without the site owner choosing it deliberately.
Is this page part of the same research as the main hub?
Yes, this is one of several framework-specific evidence pages under /research/, each going deeper on one condition of the Recognition Triangle. This one covers the Machine Gatekeeper specifically, the mechanism of selection. The main hub covers the wider evidence base across Clarity, Credibility, and Earned Authority.

A note on verification

All figures on this page are checked against the primary source: SE Ranking's own published analysis, Google's own developer documentation, and Cloudflare's own official blog. Where only a secondary write-up could be reached for supporting detail, that is noted in the text. If a figure here has been superseded or misquoted, tell us at info@aingworth.com and we will correct it.

See what the gatekeeper says about you today.

Run the free diagnostic to find which condition of the Triangle is leaking, or run the Recognition Check to see how often AI actually names you.