Quick context for anyone new here: I'm building Alisio, an invoicing tool for freelancers who bill clients across borders. I have a PM/UX background, no engineering. An AI coding assistant writes every line of the codebase; I don't touch the code myself.
The question I get most is some version of: "how do you know if the code is actually right if you can't read it?"
Honest answer: I don't verify code, I verify behavior. Before I ask for a feature, I write out the exact real-world case it has to handle — the freelancer with three clients paying in three different currencies, the invoice sent right before midnight on the last day of the month, the amount typed with a comma instead of a period. Then I test against those cases myself, by hand, the same way I'd test any product I didn't build.
When something's off, I can rarely say "line 42 is wrong." What I can say is "this number doesn't match what I'd expect in this scenario" — and that's usually enough for the AI to find the actual bug faster than I could have pointed at it myself.
It's slower than being able to read a diff and just know. But it forces a level of product thinking I think makes the product more honest, not less: if I can't picture the exact case that breaks a feature, I probably don't understand that feature well enough yet either.
Alisio is still early — a live product, no paying customers yet, looking for the first freelancers willing to kick the tires and tell me what's broken. If you bill clients internationally: https://getalisio.com/?utm_source=indiehackers&utm_medium=community&utm_campaign=semana-2026-09-28
Happy to go deeper on the non-technical-founder-directs-AI process if people want specifics.
For an invoicing app, I'd add a two-client test: sign in as client B and try opening client A's invoice using its direct URL. Check the download too, not just whether the invoice disappears from the list.
Curious how you approach regressions in your application? For example, when you see the value is wrong you can tell the AI to fix it, but what happens if it changes something you didn't expect and so you don't know to validate?
This is a problem I have as a developer building a product where I've spent a fair number of cycles to automate that kind of testing. But I was curious if you have anything you are doing from that standpoint.
I like this approach. You don’t necessarily need to understand the code to catch product bugs. Defining real-world test cases and checking the actual behavior can reveal a lot, especially early on.
Yes, you are correct and thats the way forward
Nice post! Love how you’re focusing on real-world scenarios instead of diving into code. That’s a solid way to keep the product ship‑ready. 🚀
Verifying behavior with real cases is the right habit when you can't read the diff. For Alisio I'd add a couple of contact cases to that list: an invoice email on a domain with no MX, and a client phone typed in local format for the wrong country. If those scenarios fail in the UI the same way a comma-decimal fails, the AI has something concrete to fix.
This hits hard. I see a ton of folks treat the green check as proof when the real proof is the ugly edge case. Writing the case before the feature is such a sharp habit. One thing that helped me when directing agents is keeping a tiny regression checklist of past failures and running it after every change so old currency or date bugs do not sneak back in. Behavior first is the move.
Checking behavior instead of code is the right call, and I think it holds even for people who can read a diff.
We build AI apps at my company and ran into a version of this. When one general prompt did everything, every site came out, but they all looked like the same template with the words swapped. So we split the work. Separate agents do research, copy, design and deploy, and an orchestrator checks the result before anything goes live. Same idea as yours: judge the output against what was asked for, not how it was made.
One tip: keep all your real world cases in one list and rerun every one of them after each change, not just the case for the new feature. Old cases are the ones that break quietly.
Have your test cases surfaced failures that changed the product materially, or are you still waiting for real international freelancers to expose the harder edge cases?
Love this angle. Building Xstream4K right now so this hits close to home — what made you look into it in the first place?
What made me look into it was necessity, not a choice. Early on I approved features because they ran without erroring, then found out days later something broke on a case I hadn't tested, a specific currency combination, a date right at a month boundary. I can't read the diff, so the only thing left was writing out the exact case beforehand and checking the result by hand afterward. It ended up catching things a code review would've missed too, a diff can look clean and still be wrong on a case nobody thought to write down. Curious what it looks like on your side with Xstream4K, is verifying a stream a similar problem, checking what actually plays against what's expected?