We crawled the robots.txt file of every published tool listed on directree between 6 and 7 September 2026. Of 9,037 live tools with a usable response, 945 explicitly block OpenAI’s GPTBot.
A few findings surprised us:
- 10.5% of 9,037 live AI tools block GPTBot, but Bytespider is blocked most often at 11.8%.
- Of the 945 tools that block GPTBot, 839 (88.8%) still allow OAI-SearchBot. Founders appear to be opting out of training while remaining available for ChatGPT search citations.
- Cloudflare-fronted sites block GPTBot at 22.4% across 3,749 sites, compared with 5.2% across 2,556 Vercel sites. That is more than four times the rate.
The hosting result is the one we keep thinking about. It could reflect a difference in founders’ views, but it probably also reflects product defaults and convenience. A robots.txt setting can become policy through a hosting toggle or managed file, without anyone deliberately writing the rule.
We also found that newer domains block GPTBot more often: 13.6% of 2,900 domains registered in 2026, compared with 6.9% of 145 registered in 2023. Creative categories were highest among categories with at least 150 crawled tools. Design and UI tools reached 17.3% across 318 sites, while marketing and SEO tools were at 5.3% across 245.
One caveat: we count a block only where robots.txt explicitly names the crawler in its own User-agent group and uses Disallow: /. We excluded blanket rules for every crawler, which applied to 31 sites, and treated wildcard path rules separately. This measures published robots.txt policy, not necessarily founder intent or every enforcement mechanism a site may use.
We think the distinction between training bots and search bots matters more than the headline rate. Blocking GPTBot does not automatically stop a tool appearing in ChatGPT answers. Nearly all GPTBot blockers in our sample still permit the search crawler.
What have you chosen for your own product’s robots.txt, and was it a deliberate setting or a hosting default?