
Our AI Content Scored 100% AI. Switching Models Di
Five articles came back from our editor with one line attach
Five articles came back from our editor with one line attached. "ZeroGPT says 100 percent AI. I'm not publishing these."
She was right to bounce them. I read them again the next morning, cold, and they were fine in the way a hotel lobby is fine. Nothing wrong with any of it. Nothing in it either. Every sentence ran about as long as the one before. Every paragraph opened by announcing what the paragraph was going to say.
We had produced them fast, and that part had worked. Fast and publishable content marketing turned out to be two different finish lines, and we had only crossed the first one
The two fixes that didn't work
First instinct was to write a better prompt. We added voice instructions, banned a list of words, told the model to vary its sentence lengths, fed it samples of writing we liked. The score moved a little. Not enough to publish. And the next article needed that same prompt rebuilt from memory by whoever happened to be at the keyboard, so the fix didn't survive into the following week.
Second instinct was to change the model. This is the one I hear most from founders, and it's the one I would have bet on before testing it. We ran the same brief through a different frontier model, then a third. The scores barely moved.
That result is worth sitting with, because it kills a comfortable theory. Detectors aren't sniffing out one vendor's fingerprint. If they were, switching would work. They're measuring something every one of these models does by default, which means no vendor is going to quietly solve it for you in the next release.
What the detectors are actually reading
Two things, mostly.
Predictability is the first. A detector goes through your text and asks, at each word, whether it would have guessed that word. When the answer keeps coming back yes, the writing reads as machine-made. Models reach for the safest available word every single time. People don't. We pick the odd verb, the too-specific noun, the phrasing that sits slightly wrong and is better for it.
Sentence variance is the second, and it's the louder signal. AI produces sentences of roughly equal length, one after another, on and on. Smooth and even. Humans are lumpy. A four-word sentence slams into a thirty-word one, then a fragment, then back up. That lumpiness is hard to fake with an instruction and it falls out naturally when a person actually edits.
Once we understood that, the failure made sense. The detectors had caught us not editing.
The fix was a checklist
What worked was boring. We wrote down what a publishable piece required, as a list a human runs after the draft already exists.
Delete every em-dash. Cut the transition crutches, the moreovers and the furthermores. Kill the first sentence of most paragraphs, because it usually just announces the paragraph. Swap one vague claim per section for a real number, a real tool name, or a real job title. Read the whole thing aloud and break up any run of sentences that clock in at the same length.
Nothing clever in there. The value was never in the individual rules. It was that the rules lived in a document instead of in somebody's head, so the second article got the same treatment as the first, and so did the eleventh.
The rewritten piece passed both detectors. So did the four after it. What changed between the batch that failed and the batch that passed was not the model and not the prompt. A defined pass existed, and someone ran it.
Why the workflow is the part worth building
One distinction did most of the work for us.
A prompt is single use. You write it, you get an output, and the value evaporates unless somebody saves it and the next person can find it. A workflow is an asset. It gets slightly better every time someone hits an edge case and adds a line to it.
That matters even more as AI-driven search changes how people discover and evaluate content. The model handles volume, but the workflow makes sure every piece still has a clear purpose, useful information, and a human owner at the end who signs off.
The model handles volume. The person handles judgment. Neither part is optional, and keeping them separate is what makes output consistent instead of dependent on who happened to be on shift that day.
For a solo founder this scales down cleanly. You don't need twenty of them. You need one, for whatever it is you produce most often.
What this doesn't mean
Two caveats, because process talk gets oversold badly.
The work moves rather than shrinks. You spend less time coaxing a draft out of a chat window and more time deciding what good looks like and then holding the line on it. Writing the workflow down the first time costs you an afternoon you would rather spend shipping. It pays back around the fourth run, not the first.
And a defined process won't make weak content good. Run a careful review pass on a piece nobody needed and you get a well-structured piece nobody needs. The judgment about what deserves writing at all stays with you, and no checklist rescues you from getting that part wrong.
Where to start this week
Pick the one thing you produce most often. A changelog post, a launch email, an onboarding sequence, whatever actually repeats for you.
Write down the steps you already take when you do it well. Not the aspirational version. What you genuinely do on a good day. Then add the review pass underneath: the specific checks a draft has to survive before it goes anywhere. Run the whole thing twice. That second run is where you find the three steps you forgot to write down, and those are usually the ones carrying the quality.
You'll know it's working when the output stops depending on how sharp you felt that morning.
That was the real lesson from the batch that got rejected. We had five articles and no agreed definition of done, and a detector was just the first thing rude enough to say so out loud.

Comment