The painful journey
For a while I avoided saying this out loud: the agent wasn't good enough, and the whole product rested on it.
Everything else we could fix with normal work. Screens, speed, the bugs, all of it responds to effort. But the entire promise was that you tell it what happened and it handles what follows. And in that first period, it handled the sentence I had in my head when I wrote it, and it fell over on the million other ways a real person says the same thing.
It was too literal. Ask for one thing, get exactly that thing, and none of the obvious work around it. It didn't connect the dots that any decent assistant would connect. It would do what you said instead of what you meant, and then it would be pleased with itself.
I even remember the specific moment I stopped defending it - rationalizing why it doesn't understanding me correctly.
We rebuilt how it thinks, more than once. Not tweaks to the wording we gave it. Structural changes to what it's allowed to do, how it checks its own work, how it asks when it's unsure instead of guessing confidently. That work is still going. I don't think it ever really ends.
Question for anyone building with AI right now: how do you decide when it's good enough to put the agent in front of people?