
directree
The honest software directory.
We crawled the robots.txt file of every published tool listed on directree between 6 and 7 September 2026. Of 9,037 live tools with a usable response, 945 explicitly block OpenAI’s GPTBot.
A few findings surprised us:
- 10.5% of 9,037 live AI tools block GPTBot, but Bytespider is blocked most often at 11.8%.
- Of the 945 tools that block GPTBot, 839 (88.8%) still allow OAI-SearchBot. Founders appear to be opting out of training while remaining available for ChatGPT search citations.
- Cloudflare-fronted sites block GPTBot at 22.4% across 3,749 sites, compared with 5.2% across 2,556 Vercel sites. That is more than four times the rate.
The hosting result is the one we keep thinking about. It could reflect a difference in founders’ views, but it probably also reflects product defaults and convenience. A robots.txt setting can become policy through a hosting toggle or managed file, without anyone deliberately writing the rule.
We also found that newer domains block GPTBot more often: 13.6% of 2,900 domains registered in 2026, compared with 6.9% of 145 registered in 2023. Creative categories were highest among categories with at least 150 crawled tools. Design and UI tools reached 17.3% across 318 sites, while marketing and SEO tools were at 5.3% across 245.
One caveat: we count a block only where robots.txt explicitly names the crawler in its own User-agent group and uses Disallow: /. We excluded blanket rules for every crawler, which applied to 31 sites, and treated wildcard path rules separately. This measures published robots.txt policy, not necessarily founder intent or every enforcement mechanism a site may use.
We think the distinction between training bots and search bots matters more than the headline rate. Blocking GPTBot does not automatically stop a tool appearing in ChatGPT answers. Nearly all GPTBot blockers in our sample still permit the search crawler.
What have you chosen for your own product’s robots.txt, and was it a deliberate setting or a hosting default?
Software buying is quietly moving off the search results page.
Instead of typing a query and scanning ten links, more and more people just ask ChatGPT or Perplexity "what's the best tool for X" and act on the three names it gives back. There's no page 2. If your product isn't one of those names, you don't exist for that buyer, and you'll never see it in your analytics. No impression, no click, no trace.
I kept running into founders (including me) who had no idea whether this was happening to them. So this week I shipped a small free thing on directree to answer exactly that question: is ChatGPT recommending your SaaS?
You paste your domain. It sends three real buyer questions to a live model with web search on, reads the raw answers, and shows you whether your brand is actually named, in how many of the three, and which competitors got recommended instead. No account, no charge.
The honest part matters here, and it's where I nearly shipped it wrong.
The first version was basically a lead-gen trick. It leaned on questions that already mentioned the brand, so almost everyone scored high. Great for conversions, useless as a signal, and completely against the whole reason directree exists (we're the "no fake precision scores" directory). If a tool tells you you're crushing AI search when you're not, it's lying to you.
So I tore that out. Now the questions are brand-free category questions, and detection is binary: your name is in the raw AI answer or it isn't. I tested it against a made-up brand to be sure it would honestly return "not visible" instead of inventing a flattering number. It did. That felt like the right bar.
The one that stung: I asked it "best note-taking tools" and it named Notion, Evernote, OneNote and Google Keep on its own. Notion showed up. A brand nobody's heard of did not. That's the real game now, and most founders are invisible in it without knowing.
I also wired the tool into the actual fix. When you see you're missing, there's a clear next step: a free, honest, structured listing gives answer engines a machine-readable source they can cite, plus a do-follow backlink. It's not magic and I don't pretend it is. But "consistent facts across third-party sources" is genuinely what these engines pull from, and a directory listing is one concrete way to add that.
Small tool, but it's the kind of thing I like building: it teaches you something true in a minute, even if you never list with us.
1 Like
Comment
Everyone says "don't build another directory." Fair. There are hundreds, and most blur together.
But I kept running into the same small frustration: when I'm evaluating a tool on one of these sites, I can't always tell what's actually verified versus what an AI just generated to fill the page. Integrations, pricing, "best for" claims, it all reads with the same confidence, whether it's a checked fact or a guess.
So I'm building directree around one rule: be honest about where every fact comes from.
Every field on a listing is labelled. "Observed" means we crawled and verified it, and it's shown as a fact. "AI-inferred" means a model generated it (the summary, the strengths, the likely competitors), and it's clearly marked as a model's take, never disguised as certainty. Founders can claim their listing, correct anything that's wrong, and keep a real do-follow backlink.
That's the product. Here's the part I actually lie awake thinking about.
The directory itself is the easy 20%. Anyone can build a listing page. The reason most directory projects quietly die is distribution: they can't get the pages ranked, so nobody ever sees them. That's the 80%.
And it's the one part I have an unfair advantage in. I already run an SEO/GEO pipeline across my other products (indexing automation, keyword research, the whole content engine). directree gets to plug straight into it. So instead of "build it and pray," I'm treating ranking and getting cited by AI answers as the actual core problem, and the honest listings as the thing worth ranking.
There's a newer wrinkle too. People increasingly ask ChatGPT "what should I use for X?" instead of Googling it. Which means the sources those models trust matter more every month. A directory that's structured, transparent, and clear about what's verified is exactly the kind of source I'd want an AI to pull from. That's the bet I'm making with the "GEO" side of it.
Where I am now: the directory is live, listings are getting enriched, and I'm heads-down on the boring, compounding distribution work rather than chasing a launch-day spike.
Open question for anyone who's done this: for a directory, what actually moved the needle on getting pages indexed and ranked, beyond just publishing volume? Backlinks, freshness, structured data, something else entirely? I'd genuinely like to hear what worked (or didn't) for you.
3 Likes
1 Comment
1 Comment
-
1
I think the more interesting distinction isn't that you're building another directory—it's that you're making uncertainty visible instead of hiding it.
As AI-generated information becomes more common, clearly separating observed facts from inferred opinions could become a trust signal in its own right. That's a much harder advantage to replicate than simply having more listings.
About
I wanted a tool myself to compare and read about softwares. It grew and turned out being directree.


Comment