One thing I’m becoming less interested in with AI app builders is the first screenshot.
A lot of tools can produce something that looks convincing very quickly now.
What I care about more is what happens when you actually use the thing:
Does the workflow make sense?
Does the logic hold up?
Can you change it without starting over?
What happens once the app has real rules instead of just nice screens?
We’ve been documenting some of these experiments while working on built.new:
I’m obviously involved with built.new, but I’d be interested in hearing from people using other builders too:
What’s the first thing you test after the “wow, it generated an app” moment wears off?
Appreciate the honesty here, most people only share the wins.
Nice progress. What is the next thing you are focusing on?
Good write-up. What would you do differently if you started again?
Nice progress. What is the next thing you are focusing on?
Interesting approach. What was the hardest part to get right?
Good point. Did you test that with users before committing to it?
Good point. Did you test that with users before committing to it?
What made you pick this stack over the alternatives?
Makes sense. Are you planning to charge for it, or keep it free for now?
Nice progress. What is the next thing you are focusing on?
Makes sense. Are you planning to charge for it, or keep it free for now?
Great breakdown. What feedback have you had from early users?
Good point. Did you test that with users before committing to it?
Really relatable. How much time do you put into this each week?
How did you decide this was worth building in the first place?
Nice work shipping it. What has been the biggest challenge since launch?
Nice work shipping it. What has been the biggest challenge since launch?
Nice work shipping it. What has been the biggest challenge since launch?
That distinction matches what I’ve seen: the first useful test is to take the generated happy path and then change one assumption—an input shape, validation rule, or failure case. If the builder can’t explain or preserve the change, the screenshot was mostly theatre. I’d also test whether a new user can recover after the first error without knowing the implementation. Do you have a repeatable “break it” checklist yet, or are you still collecting cases?
Really solid approach — curious how you're thinking about this, what's been the hardest part to figure out so far?