9
20 Comments

I can't read code. Here's how I still catch bugs before they ship.

Quick context for anyone new here: I'm building Alisio, an invoicing tool for freelancers who bill clients across borders. I have a PM/UX background, no engineering. An AI coding assistant writes every line of the codebase; I don't touch the code myself.

The question I get most is some version of: "how do you know if the code is actually right if you can't read it?"

Honest answer: I don't verify code, I verify behavior. Before I ask for a feature, I write out the exact real-world case it has to handle — the freelancer with three clients paying in three different currencies, the invoice sent right before midnight on the last day of the month, the amount typed with a comma instead of a period. Then I test against those cases myself, by hand, the same way I'd test any product I didn't build.

When something's off, I can rarely say "line 42 is wrong." What I can say is "this number doesn't match what I'd expect in this scenario" — and that's usually enough for the AI to find the actual bug faster than I could have pointed at it myself.

It's slower than being able to read a diff and just know. But it forces a level of product thinking I think makes the product more honest, not less: if I can't picture the exact case that breaks a feature, I probably don't understand that feature well enough yet either.

Alisio is still early — a live product, no paying customers yet, looking for the first freelancers willing to kick the tires and tell me what's broken. If you bill clients internationally: https://getalisio.com/?utm_source=indiehackers&utm_medium=community&utm_campaign=semana-2026-09-28

Happy to go deeper on the non-technical-founder-directs-AI process if people want specifics.

on September 28, 2026
  1. 2

    brk's regression question is the real gap in this workflow. Hand-checking the new case works; the next agent session often never sees the old ones (comma vs period, midnight month boundary), so a "fix" quietly undoes them and you don't know to retest. I'd keep that case list in a file the agent is instructed to re-run after every change, not only in the chat that wrote the feature. Direct-URL access between two clients, as feech said, belongs on that same list.

  2. 1

    Your three-currency freelancer, comma-decimal, and midnight invoice cases are stronger than a green test suite written by the same AI.
    The diagnosis is production patterns loading/empty/error, Chapter 6 of Vibe Coding for Beginners: behavior needs a visible check at each state, not just a happy path.
    Start with one regression case file containing only the midnight invoice scenario.
    After the next AI change, can that exact invoice still be sent and reloaded?
    Kael Voss / DurableFoundations

  3. 1

    A green suite the agent wrote is agreement with itself, not proof the product case holds. I'd write the ugly real-world scenario before the feature (inputs, currencies, failure modes), treat "looks right in the transcript" as untrusted until that case fails closed or passes, and refuse "tests pass" as done without an expected answer that didn't come from the implementation. Soft "I can't read the diff so ship it" trains developers to overlook correlated bugs, while a case-first check prevents this. Curious which single case you'd force before any money-moving path first: multi-currency invoice, failed payment retry, or a client who never opens the PDF.

  4. 1

    Behavior testing is essential, but it can’t expose security, architecture, data-handling or operational risks. That’s the gap we built Agiloop to assess alongside real-world acceptance tests — especially when AI writes the code and the founder can’t inspect it directly.

  5. 1

    Same setup here: Claude writes nearly all the code behind UtilitySEO, and we verify behaviour, not diffs.

    One case I'd add to every list, because a regression file won't catch it: the save that looks right and never persisted. We got burned by a tool where the edit showed on screen, the confirmation appeared, and after a reload the old value was back. Now every check ends with "reload and look again".

    That leads to the second habit: treat "fixed" from the AI as a claim, not a result. It usually means it changed what it believed was the cause. The scenario only passes once you've seen the number yourself after a fresh load.

    For invoices, the one that bites is the total on screen vs the total in the PDF the client actually receives. Do you check the sent document, or only the app view?

  6. 1

    So you're debugging functional bugs, not technical bugs. Great!

  7. 1

    The detail I find most interesting is what writing the case first does to you, not the AI. When you have to spell out "invoice sent right before midnight on the last day of the month" before the feature exists, you've already turned a vague request into an acceptance test — and I'd guess fewer of your feature requests come back wrong at all, not just that bugs get caught faster. Two things I'd be curious about as Alisio grows: first, does every new feature get its cases, or do small changes quietly skip the ritual? A case list that starts rotting is worse than none, because it convinces you the check ran when only half of it did. Second, invoicing has a class of case that only shows up at calendar edges — a client whose fiscal year ends in March, a periodic invoice generated during the leap-day, timezones flipping the issue date over midnight. Those are exactly the ones neither you nor the AI thinks to write down, so it might be worth letting a script generate the boring cases (every month boundary of the next 12 months) and only hand-check what it finds.

  8. 1

    The thread has converged on the right fix, and it is worth putting a number on why the chat-based version fails, because the failure is combinatorial rather than a matter of discipline.

    Scenario testing has a cost that grows quadratically while its coverage grows linearly. Add one feature that interacts with two existing ones and you have three things to re-check, not one. Ten features with pairwise interactions is 45 checks, and at a realistic thirty seconds each that is 22 minutes of hand verification per change, which is the point where a solo founder stops running the list by hand. The list is not too long because you were lazy; it is too long because pairwise interactions outrun any linear habit.

    That gives a concrete threshold. Below roughly 10-15 active invariants, a written file plus discipline works and you should not build anything. Above it, the honest options are to test a rotating subset, which means a regression can sit undetected for several cycles, or to move the check to something that runs without you. There is no third option where the list stays complete and hand-run.

    The second number is the one that decides whether your approach holds. You verify behaviour, which is correct, but the verification only catches what the case describes. An invoice with three currencies and a comma decimal is one case; three currencies with a comma decimal and a midnight boundary is another, and the pairs multiply faster than you will write them. So the useful question is not how many cases you have, but how fast the set grows when a feature ships.

    I build Piramyd (https://piramyd.cloud), which is aimed at exactly that ratio, so treat the numbers above as an interested party's framing.

    Concretely: how many invariants do you have on the list right now, and when you add a feature, do you add cases or does the count stay the same?

  9. 1

    Do you keep a permanent invariant list outside the agent’s chat?

  10. 1

    Scenario-based tests are a strong substitute for code-level confidence. I'd keep a small regression checklist of those edge cases and rerun it after every change so an AI fix doesn't quietly break a working flow.

  11. 1

    For an invoicing app, I'd add a two-client test: sign in as client B and try opening client A's invoice using its direct URL. Check the download too, not just whether the invoice disappears from the list.

  12. 1

    Curious how you approach regressions in your application? For example, when you see the value is wrong you can tell the AI to fix it, but what happens if it changes something you didn't expect and so you don't know to validate?

    This is a problem I have as a developer building a product where I've spent a fair number of cycles to automate that kind of testing. But I was curious if you have anything you are doing from that standpoint.

  13. 1

    I like this approach. You don’t necessarily need to understand the code to catch product bugs. Defining real-world test cases and checking the actual behavior can reveal a lot, especially early on.

  14. 1

    Yes, you are correct and thats the way forward

  15. 1

    Nice post! Love how you’re focusing on real-world scenarios instead of diving into code. That’s a solid way to keep the product ship‑ready. 🚀

  16. 1

    Verifying behavior with real cases is the right habit when you can't read the diff. For Alisio I'd add a couple of contact cases to that list: an invoice email on a domain with no MX, and a client phone typed in local format for the wrong country. If those scenarios fail in the UI the same way a comma-decimal fails, the AI has something concrete to fix.

  17. 1

    This hits hard. I see a ton of folks treat the green check as proof when the real proof is the ugly edge case. Writing the case before the feature is such a sharp habit. One thing that helped me when directing agents is keeping a tiny regression checklist of past failures and running it after every change so old currency or date bugs do not sneak back in. Behavior first is the move.

  18. 1

    Checking behavior instead of code is the right call, and I think it holds even for people who can read a diff.

    We build AI apps at my company and ran into a version of this. When one general prompt did everything, every site came out, but they all looked like the same template with the words swapped. So we split the work. Separate agents do research, copy, design and deploy, and an orchestrator checks the result before anything goes live. Same idea as yours: judge the output against what was asked for, not how it was made.

    One tip: keep all your real world cases in one list and rerun every one of them after each change, not just the case for the new feature. Old cases are the ones that break quietly.

  19. 1

    Have your test cases surfaced failures that changed the product materially, or are you still waiting for real international freelancers to expose the harder edge cases?