I'm building HumanizeMyAI, and before writing any marketing I wanted the uncomfortable answer: do these tools actually work? So I ran the same passages through 9 humanizers and scored every output on 6 public detectors (GPTZero, Originality, Copyleaks, ZeroGPT, Turnitin, QuillBot). 162 measurements.
It was worse than I expected:
Most "humanizers" are just synonym-swappers. The harder they reword, the more they trip the detector's paraphrase-pattern signal.
QuillBot's humanizer scores ~95% AI on QuillBot's own detector. The vendor catches its own tool.
The paraphraser-class tools landed between ~22% and ~33% mean AI — nowhere near "undetectable."
But the deeper problem isn't the tools, it's the detectors. A Stanford study found they falsely flag non-native English writers 61% of the time vs 5% for native speakers. A lot of flagged people didn't cheat — they just don't write like the training data.
That's the actual reason I'm building this: not "beat the detector," but help the students who get wrongly flagged — using a corpus of 2,590 real student essays (58% of them ESL) instead of a black-box paraphraser. Free, 4 runs a day.
Still very early ($0 MRR, building in public). Two things I'd love takes on:
Is publishing my own benchmark that names competitors honestly smart positioning, or a fight a small player can't win?
For a free-first tool in a gray-area niche, what converts to paid without feeling scummy?
Happy to share the full detector-by-detector numbers if useful.
The thing I'd be careful with is that both questions may actually be downstream of the same decision.
Right now there are at least two very different stories a buyer could tell themselves about HumanizeMyAI, and the choice between them changes what "good positioning" looks like, what feels trustworthy, and what someone would ever pay for.
I would not make that call casually in a thread because it ends up shaping the benchmark, the offer, and the conversion path all at once.
Worth tightening before optimizing anything else.