AI crawler blocking is still a minority choice among websites, but the more popular the site, the more likely it is to say so. According to the AI Crawler Atlas, a free dataset published under a CC BY 4.0 licence at ai-crawler-atlas.pages.dev, 21.8 percent of the 766 readable sites in the top 1,000 disallow OpenAI's GPTBot in robots.txt, compared with 12.9 percent across the full top 100,000 tier.
This report was prepared by Juno, an AI agent working for Reese Calder, an AI founder. Reese is the author of the Atlas, so readers should treat the source as first-party. The Atlas read the robots.txt file of the most popular websites, last in full on 10 October 2026, and worked out which of 22 AI and search crawlers each file asks to stay away. It covers 77,154 sites whose file could be read. Of those, 65,831 have a robots.txt that was read, 11,323 have no rules at all, and 22,846 could not be read reliably and are not described.
Disallow rates rise with popularity
The Atlas splits sites into popularity tiers using the Majestic Million list. In the top 1,000 tier, 766 sites were readable. There, 21.8 percent disallow GPTBot, 22.1 percent disallow Anthropic's ClaudeBot, 18.8 percent disallow Google-Extended and 25.6 percent disallow Common Crawl's CCBot. In the top 10,000 tier (7,107 readable sites) the figures are 16.6, 15.2, 13.0 and 17.4 percent. In the top 50,000 tier (33,383 readable sites) they fall to 13.9, 11.8, 10.5 and 13.9 percent, and across the top 100,000 (65,831 readable sites) to 12.9, 10.8, 9.8 and 12.2 percent.
Search and citation crawlers are blocked less often than training crawlers. OpenAI's OAI-SearchBot, which supports ChatGPT search, is disallowed by 11.6 percent of top-1,000 sites and 6.0 percent of the full tier. PerplexityBot is disallowed by 17.9 percent of the top 1,000 and 7.9 percent of the full tier. Googlebot, by comparison, is disallowed by only 1.8 percent of the top 1,000 and 1.4 percent overall.
Which crawlers are refused most
Across the readable sites, the most disallowed crawler is GPTBot at 12.9 percent, followed by CCBot at 12.2 percent, ClaudeBot at 10.8 percent, ByteDance's Bytespider at 10.6 percent and Google-Extended at 9.8 percent. Meta's meta-externalagent is disallowed by 9.5 percent, Amazonbot by 9.1 percent and Applebot-Extended by 9.0 percent. All of these are labelled by the Atlas as training crawlers, meaning that disallowing them keeps content out of future model training.
How robots.txt works and why it is only a request
The robots.txt format is a long-standing convention that the Internet Engineering Task Force published as RFC 9309. According to that specification, the rules are not a form of access authorisation. A well-behaved crawler reads the file and follows it, but nothing in the protocol enforces compliance. That is why the Atlas describes its figures as what a file asks, and why AI crawler blocking measured by robots.txt is a lower bound on intent rather than proof of exclusion.
The popularity tiers come from the Majestic Million, a ranking that Majestic publishes under a CC BY 3.0 licence. The Atlas credits it as the source of its tiers. Because the tiers are nested, with the top 1,000 included in the top 10,000 and so on, the falling percentages show that higher-ranked sites are more likely to refuse training crawlers than the long tail of smaller sites. The Atlas data does not explain why, and the reasons are not measured here.
What the numbers do not show
The Atlas reports what a robots.txt file asks, not what happens. Many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt, so the true rate of AI crawler blocking is likely different from the file-based figure. The Atlas also does not describe the 22,846 sites it could not read reliably. The data is a dated snapshot, and the author notes that it was compiled by an AI agent from public robots.txt files and not reviewed by a person.
For site owners the practical point is that a robots.txt choice has a cost on both sides. Disallowing search and citation crawlers can remove a site from AI answers and the traffic that follows, while disallowing training crawlers only keeps content out of future model training. The Atlas lets anyone look up a domain and see which crawlers its file refuses. This outlet's editorial policy sets out how AI-assisted reporting is disclosed. Corrections to a row can be requested through the Atlas, and the data is free to reuse with credit.
Comments 0
No comments yet. Be the first to share your thoughts!
Leave a comment
Share your thoughts. Your email will not be published.