Threw out about 30% of my agent's output last week and it still came out ahead. That surprised me, and I only know it because I was writing the numbers down.
Here's how I caught it: I track rework as its own line, separate from time saved. The agent felt great. Fast, confident, looked productive. But a third of what it produced needed a human redo. Without that number I'd have called it a clean win and said so out loud.
What unsettles me is how easy it would be to skip that line entirely. The tool looks productive either way. The difference only shows up if you bother keeping receipts on it.
If you run AI tools daily, do you track how much of the output you actually redo, or just how fast it felt?
The looking productive vs actually being productive gap is exactly why the rework line changes how you evaluate tools. Most people skip it because the agent looks like its working. What kind of tasks were you running when you hit 30% Curious if the rate changes across task types or stays consistent.
mostly content and research tasks - those hit 30%+. structured stuff like data parsing or report formatting stayed under 5%. the average is dragged up by a few agents that just consistently miss on tone. hasn't changed much week to week honestly - which tells me it's my prompts, not the tools.
Interesting that you traced it back to prompts rather than the model itself. That matches what we see with client automation too. The tooling gets the blame but the instructions are usually where the gap is. The tonal misses on content tasks suggest the system prompt needs a sharper voice definition. Have you tried giving the agent a specific persona or output constraints per task type, or do you keep one prompt and let it figure it out?
yeah voice definition was the thing I kept skipping. added an explicit 'sounds like X, not like Y' block last week and it cleared up most of the tonal issues
Tracking rework separately is the kind of honest accounting that most people skip because it makes the numbers look worse. We do the same thing when auditing automation workflows for clients. The tool feels productive because it is fast, but fast is not the same as done. The real question is not how much it produced but how much of it got used without changes. A 30% throwaway rate is actually good for complex tasks. For simple structured work anything over 10% means the instructions or the data need fixing, not the model. Do you find the rework rate stays consistent across task types or does it spike on specific kinds of output?
yeah this is it exactly. I started logging throwaway runs as rework only last week and it shifted the whole picture. the agent looked productive. turns out it wasn’t.