How the gatekeeper actually decides.
The Machine Gatekeeper framework says AI systems now sit between a brand and its audience, deciding who gets named. This page is the evidence for the mechanism: what a 129,000-domain study found actually correlates with getting cited, what Google itself says about optimizing for it, and why a large share of sites accidentally lock the gatekeeper out entirely.
Last verified 20 August 2026In one paragraph
Independent research on 129,000 domains found citation frequency tracks closely with referring domains, domain trust, content freshness, and depth, ordinary authority signals, not hidden AI tricks. Google's own documentation says the same thing in its own words: there is no separate optimization layer for AI features, the core search quality systems are what generative answers draw on. And Cloudflare's own 2025 traffic data shows AI crawlers are growing fast while robots.txt files disallow them more than any other bot category, meaning a meaningful share of sites are shutting the gatekeeper out without realizing it.
Ordinary authority signals, not tricks.
SE Ranking analyzed 216,524 pages across 129,000 domains and 20 industry niches to see what actually separated frequently-cited sites from rarely-cited ones in ChatGPT. Four factors showed the clearest gap between the weakest and strongest tier, shown here as average citations per page.
Domain Trust and referring domains showed the strongest relationship of anything tested, both proxies for Earned Authority: what independent sites already link to and vouch for. Page speed also correlated (pages loading in under 0.4 seconds averaged 6.7 citations), a Clarity-adjacent signal, the gatekeeper can only weigh what it can render quickly. Source: SE Ranking, 129,000 domains, published 26 November 2025.
There is no separate game to learn.
Every AI-visibility vendor sells a version of "AI optimization" as something distinct from SEO. Google's own developer documentation says otherwise, directly.
“
The best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems.
Google, AI features and your website, developers.google.com · Read the source
This is consistent with the SE Ranking data above: the strongest citation signals, Domain Trust and referring domains, are exactly the same signals traditional SEO and PR have always tracked. The gatekeeper is not running a hidden second scoring system. It is applying the same quality bar it always has, to a different output format.
You can see that you were shown. Not why.
In June 2026, Google added a Generative AI tab to Search Console, the first official window into AI Overviews and AI Mode performance. What it shows, and deliberately withholds, says a lot about how opaque the gatekeeper's logic still is, even to the people building it.
What it shows
Impressions: every time a URL from your domain is rendered inside an AI-generated answer block, filterable by URL, region, device, and date range.
What it withholds
Click-through data, the actual prompts or queries that triggered the citation, and position within the answer. Google cites user privacy and interface volatility as the reasons.
In short: you can now confirm the gatekeeper noticed you. You still cannot see which question earned that notice, or why one page was chosen over another. Source: Google Search Console Generative AI reporting, launched June 2026.
The gatekeeper cannot recognize what it cannot reach.
Cloudflare's own network data shows AI crawlers growing fast, and, at the same time, robots.txt files disallowing them more often than any other category of bot. Two facts sitting side by side: the traffic that matters most right now is also the traffic most commonly blocked, frequently by accident.
Growing fast
ChatGPT-User request volume peaked as much as 16x higher than where it started the year; PerplexityBot traffic finished the year roughly 3.5x higher; ClaudeBot crawling volume nearly doubled at its peak.
Blocked most often
Cloudflare's own review of 2025 traffic states plainly that AI crawlers were "the most frequently fully disallowed user agents found in robots.txt files" of any bot category it tracks.
Source: The 2025 Cloudflare Radar Year in Review, published December 2025.
This page covers the selection mechanism specifically. For the wider evidence base behind the Recognition Triangle, including who gets cited and how concentrated that citation share actually is, see the main research hub.
Back to the research hubQuestions people ask about this page.
Is there really no separate way to optimize for AI answers?
Does this mean schema markup and llms.txt don't matter?
How would I know if my own site is blocking AI crawlers?
Is this page part of the same research as the main hub?
A note on verification
All figures on this page are checked against the primary source: SE Ranking's own published analysis, Google's own developer documentation, and Cloudflare's own official blog. Where only a secondary write-up could be reached for supporting detail, that is noted in the text. If a figure here has been superseded or misquoted, tell us at info@aingworth.com and we will correct it.
See what the gatekeeper says about you today.
Run the free diagnostic to find which condition of the Triangle is leaking, or run the Recognition Check to see how often AI actually names you.