2
58 Comments

Vibecoded apps rarely fail on the clever code, they fail on the boring checks

We build AI-generated apps at Inithouse. Fourteen products, all vibecoded across Lovable, Cursor, and similar tools. After running professional audits on dozens of these projects - ours and others - we noticed something that keeps repeating.

The code that breaks production is almost never the complex logic. It's the stuff nobody thinks to check.

Where things actually break

We built Audit Vibe Coding - a professional audit for AI-generated (vibecoded) projects - it scores security, SEO, performance, accessibility and code quality and returns prioritized fixes. We run 47 checks across eight domains. Here's what the data keeps showing:

The average vibecoded project we audit scores around 31 out of 100. Production-ready starts at 80.

That 49-point gap? Almost none of it is in the application logic - the thing the builder spent hours prompting. The gap sits in security headers, meta tags, accessibility attributes, error boundaries, and performance optimization. The boring layer.

Here is where most projects lose the bulk of their points:

Security headers - missing Content-Security-Policy, no X-Frame-Options, absent Strict-Transport-Security. AI builders generate functional apps. They don't add the HTTP headers that stop clickjacking, XSS, or mixed content. We've seen projects with solid auth flows and zero security headers. The auth works. The iframe embedding your login page on a phishing domain also works.

SEO meta and indexability - no canonical tags, duplicate titles across pages, missing Open Graph, broken or absent sitemaps. One of our own products had working pages that Google couldn't find for three weeks because the sitemap returned a 404. The pages existed. The route to them didn't.

Accessibility - missing alt attributes, no ARIA labels on interactive elements, color contrast failures, keyboard navigation broken. Not a single vibecoded project we've audited shipped with a working skip-to-content link. Not one.

Performance - unoptimized images (4 MB hero PNGs are common), no lazy loading, render-blocking scripts in the head. One project loaded 11 MB on first paint. It worked fine on the builder's fiber connection. On a phone in a subway, it timed out.

Why AI builders skip this

The AI-generated code itself is usually fine. The code generation is often decent - solid component structure, reasonable state management, working database queries.

The issue is the prompt loop. You prompt for a feature, you get the feature. You prompt for a login flow, you get the login flow. Nobody prompts "now add all the HTTP security headers" or "make every interactive element keyboard-accessible." Those aren't features. They're hygiene, and the prompt-build-test cycle naturally skips them because they don't produce visible changes.

We've watched this in our own workflow. We'd build a product, test the happy path, ship it, and then find - weeks later - that Lighthouse scored it at 38 for accessibility or that a missing robots.txt blocked half the site from indexation.

The same pattern shows up in every AI builder we've tested. Lovable, Cursor, Bolt, Replit - they all generate working features and skip the production plumbing. It's not a flaw in any one tool. It's a structural blind spot in prompt-driven development: you get what you ask for, and nobody asks for Content-Security-Policy.

What we got wrong across our own portfolio

Looking back at our fourteen products with the lens of the audit checklist, here are patterns we wish we'd caught sooner:

Missing error boundaries. Several of our React apps had no error boundaries at all. A single failed API call in a nested component would white-screen the entire app. The fix is four lines of code. We added it to twelve products retroactively.

No rate limiting on public endpoints. Three of our products had open API routes with zero rate limiting. One of them got hit by a script doing 400 requests per second for six hours. Supabase edge function billing is usage-based. That was an expensive weekend.

Mixed content warnings blocking features. Two products loaded assets over HTTP on HTTPS pages. Browsers silently blocked the assets. The features "didn't work" in production but worked fine in development. It took us two days to track down a bug that was a protocol mismatch in an image URL.

Cookie consent missing in EU markets. We shipped products to Czech, Slovak, Polish, and German users without cookie consent banners. Not because we were ignoring GDPR, but because the AI builder didn't add one and we didn't think to prompt for it. It's not a feature. It's a legal requirement that looks like boilerplate.

Structured data absent everywhere. None of our products launched with JSON-LD schema markup. For products competing in search - and most of ours do - that's organic traffic left on the table. Adding it after launch means waiting for re-crawl, re-index, and re-ranking. Starting with it means search engines understand your pages from day one.

What we'd do differently

If we were starting the portfolio again today, we'd run the boring checks first, before the first deploy.

The sequence would be:

  1. Build the feature (prompt-driven, as usual).
  2. Before touching the deploy button, run the audit checklist on the build output.
  3. Fix everything scored critical or high - almost always security headers, meta tags, and accessibility basics.
  4. Deploy.
  5. Verify indexability within the first 48 hours (sitemap resolving, pages appearing in Search Console).

That's maybe two hours added to each launch. In exchange, you skip the three-week-later discovery that half the site is invisible to search engines or that the login page has no CSRF protection.

The boring layer is where production readiness lives

We keep saying this internally: the interesting code is almost never the problem. The auth flow works, the database queries run, the UI renders. What's missing is the layer between "it works on my machine" and "it can survive real users, real browsers, real search crawlers, and real attackers."

AI tools are getting better at application logic every month. The gap we keep finding in audits - and the reason we built the audit as a product - is that nobody is prompting for the production layer. Security, SEO, performance, accessibility. The boring stuff. The stuff that breaks real deployments.

If you're shipping vibecoded apps, check the boring layer before you ship. Not after. The clever code can wait. The Content-Security-Policy header can't.


Jakub, builder @ Inithouse. We build and audit AI-generated apps at auditvibecoding.com

on September 24, 2026
  1. 1

    Appreciate the honesty here, most people only share the wins.

  2. 1

    Really relatable. How much time do you put into this each week?

  3. 1

    Appreciate the honesty here, most people only share the wins.

  4. 1

    Interesting. How are you measuring whether it is working?

  5. 1

    Really relatable. How much time do you put into this each week?

  6. 1

    Same pattern on my side — I'm building a code health scorer (300-850, like a credit score), and the biggest deductions in AI-generated repos are almost never the feature logic. It's the agent solving the same problem three different ways, dead code nobody remembers generating, and those 4MB hero PNGs. The '31 out of 100' rings true. I built this; waitlist's open: https://muse.ai/s/waitlist-page-zxc6q7xbxii91n

  7. 1

    Appreciate the honesty here, most people only share the wins.

  8. 1

    Really relatable. How much time do you put into this each week?

  9. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  10. 1

    Interesting. How are you measuring whether it is working?

  11. 1

    The 4 MB hero PNG and the sitemap 404 are the details that make this real: a feature can pass the happy path while production users and crawlers hit a wall.

    I would add a red-flag pause before deploy: check one production URL on a phone, confirm the sitemap and CSP resolve, and force one API failure to see an error state. If any one fails, stop the launch and fix that slice before adding features.

    Free 10-min check: https://durablefoundations.gumroad.com/l/pyramid-reality-check

    Which check catches the most expensive miss in your next audit?

    Kael Voss / DurableFoundations

  12. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  13. 1

    Appreciate the honesty here, most people only share the wins.

  14. 1

    Really relatable. How much time do you put into this each week?

  15. 1

    Interesting. How are you measuring whether it is working?

  16. 1

    Appreciate the honesty here, most people only share the wins.

  17. 1

    Interesting. How are you measuring whether it is working?

  18. 1

    Really relatable. How much time do you put into this each week?

  19. 1

    Appreciate the honesty here, most people only share the wins.

  20. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  21. 1

    How did you decide this was worth building in the first place?

  22. 1

    Nice work shipping it. What has been the biggest challenge since launch?

  23. 1

    Really relatable. How much time do you put into this each week?

  24. 1

    This is useful. How are you finding your first users so far?

  25. 1

    Solid lesson. Which channel has worked best for you so far?

  26. 1

    Curious how long it took before you saw the first real results?

  27. 1

    Interesting approach. What was the hardest part to get right?

  28. 1

    Good write-up. What would you do differently if you started again?

  29. 1

    This is useful. How are you finding your first users so far?

  30. 1

    Appreciate the honesty here, most people only share the wins.

  31. 1

    Interesting. How are you measuring whether it is working?

  32. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  33. 1

    Really relatable. How much time do you put into this each week?

  34. 1

    Appreciate the honesty here, most people only share the wins.

  35. 1

    Really relatable. How much time do you put into this each week?

  36. 1

    Interesting. How are you measuring whether it is working?

  37. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  38. 1

    Really relatable. How much time do you put into this each week?

  39. 1

    Interesting. How are you measuring whether it is working?

  40. 1

    Appreciate the honesty here, most people only share the wins.

  41. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  42. 1

    Really relatable. How much time do you put into this each week?

  43. 1

    Good point. Did you test that with users before committing to it?

  44. 1

    Good point. Did you test that with users before committing to it?

  45. 1

    Really relatable. How much time do you put into this each week?

  46. 1

    Helpful post. How did you get your first bit of traction?

  47. 1

    Nice work shipping it. What has been the biggest challenge since launch?

  48. 1

    Helpful post. How did you get your first bit of traction?

  49. 1

    Really relatable. How much time do you put into this each week?

  50. 1

    Interesting. How are you measuring whether it is working?

  51. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  52. 1

    Appreciate the honesty here, most people only share the wins.

  53. 1

    Really relatable. How much time do you put into this each week?

  54. 1

    Appreciate the honesty here, most people only share the wins.

  55. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  56. 1

    Interesting. How are you measuring whether it is working?

  57. 1

    Really good writeup, thanks for sharing it. What's the next thing you're planning to try here?

  58. 1

    same pattern for us. what actually bit us in prod wasnt the game logic, it was a proxy header nobody configured and rate limits trusting the wrong ip. boring checks are boring right up until they arent