
AI Undetectable
AI humanizer with text and image detection
I run AI Undetectable, so this is a product-development update, not an independent review.
This week we tested five complete writing samples across Basic and Stealth: general writing, an essay, an article, marketing copy, and a story. Each source was 332-442 words. We ran two repeats per combination, then repeated the 20-output set after a prompt revision. A separate 645-word essay provided two additional holdout outputs.
The useful failures were not spelling mistakes. Some otherwise fluent outputs dropped a fictional label, moved a timing detail, or shortened the source more than intended. Those problems are easy to miss when a demo contains only one sentence.
Our review checklist now separates four questions: did the response finish, did it preserve the facts, did it preserve the evidence status, and did it retain enough of the original argument? A proposed study must stay proposed. A booking deadline must not become the session's start time.
We also added a backend check that rejects a model response when its finish reason says it hit the output limit. A cut-off response should not count as a successful paid rewrite.
Two actual, unedited essay outputs and their full source are published here, with model identifiers and downloadable data:
https://aiundetectable.com/ai-humanizer-benchmark#full-length-examples
These are selected examples, not a pass-rate claim. This editorial test did not use detector scores or customer writing, and the prompt still has limitations.
For other builders of writing tools: how do you keep your factual-preservation fixtures useful as prompts change? Separate held-out topics, repeated runs, or another approach?
Disclosure: AI assistance was used to prepare this post and the evaluation work. The linked outputs are actual recorded examples, not independently validated results.

Comment