AI Undetectable

AI humanizer with text and image detection

Visit Website
September 18, 2026 What full-length rewrite tests caught that short demos missed

I run AI Undetectable, so this is a product-development update, not an independent review.

This week we tested five complete writing samples across Basic and Stealth: general writing, an essay, an article, marketing copy, and a story. Each source was 332-442 words. We ran two repeats per combination, then repeated the 20-output set after a prompt revision. A separate 645-word essay provided two additional holdout outputs.

The useful failures were not spelling mistakes. Some otherwise fluent outputs dropped a fictional label, moved a timing detail, or shortened the source more than intended. Those problems are easy to miss when a demo contains only one sentence.

Our review checklist now separates four questions: did the response finish, did it preserve the facts, did it preserve the evidence status, and did it retain enough of the original argument? A proposed study must stay proposed. A booking deadline must not become the session's start time.

We also added a backend check that rejects a model response when its finish reason says it hit the output limit. A cut-off response should not count as a successful paid rewrite.

Two actual, unedited essay outputs and their full source are published here, with model identifiers and downloadable data:

https://aiundetectable.com/ai-humanizer-benchmark#full-length-examples

These are selected examples, not a pass-rate claim. This editorial test did not use detector scores or customer writing, and the prompt still has limitations.

For other builders of writing tools: how do you keep your factual-preservation fixtures useful as prompts change? Separate held-out topics, repeated runs, or another approach?

Disclosure: AI assistance was used to prepare this post and the evaluation work. The linked outputs are actual recorded examples, not independently validated results.

Comment