1
1 Comment

What 14 MVPs taught us about kill-criteria: which ones just survived the cull

We run 14 products at Inithouse. All built in weeks, all live, all collecting real usage data. The question we get most often from other builders: "How do you decide which ones to keep?"

The honest answer: we almost didn't. For the first few months, every product felt like it deserved more time. Sunk cost is real, especially when you built the thing yourself. So we forced ourselves to create a framework that removes gut feeling from the equation.

Here is the framework, which products passed, and what we learned about killing your own work.

The kill-criteria framework

We score every product on three axes every month:

| Signal | What we measure | Kill threshold |
|--------|----------------|----------------|
| Active usage | Monthly sessions with >30s engagement | Declining 3 months straight |
| Organic traction | GSC impressions trend (not absolute) | Flat or declining, no growth vector |
| Ops cost | Hours of manual work per week to keep it running | >2h/week with no automation path |

If a product hits the kill threshold on all three, it goes on the deprecation list. Two out of three means it gets flagged for a focused sprint: either fix the weakest signal or accept the trajectory.

One thing this framework explicitly does not do: evaluate products younger than 30 days. New launches get a stage gate instead. The question for a 2-week-old product is "did anyone come back?" not "are sessions growing month over month?" We learned this after nearly killing a product that needed 6 weeks to find its audience through search indexation.

Who survived and why

Živá Fotka turns a static photo into a short live video. It runs across 5 domains (CZ, SK, PL, EN, DE) and has the most consistent CTR in the portfolio. Usage is stable, organic traffic grows with each new locale, and it requires almost zero manual ops. This is the product that validates itself every month without us touching it.

Tarotas is our daily tarot app. It looked borderline for a while because sessions were flat. Then we noticed something: users who picked a language other than English during their first session had significantly higher return rates. That one insight opened a locale expansion path. We added 4 new languages, and the retention curve shifted. Tarotas survived because the data showed a growth vector we had not exploited yet.

Audit Vibe Coding runs security and quality audits on AI-generated codebases. It has the clearest unit economics of anything in the portfolio. Each audit is a defined deliverable with a fixed price. Organic traffic comes from developer search queries. B2B products with clean economics get the most patience in our framework.

Who is on the edge

Some products are flat. We are not going to pretend otherwise. When sessions have been stable but not growing for 8+ weeks and there is no clear experiment to run, the product moves to a watch list. We give it one more content push or distribution experiment. If the next 30 days do not change the trajectory, we sunset it.

We do not name these publicly because the decision is not final. But the pattern is consistent: products that rely entirely on paid traffic with no organic flywheel tend to stall. Products where the content itself generates distribution (like Watching Agents, where each public agent page is an indexed piece of content) tend to compound.

What the framework cannot do

Kill criteria work for products past the initial launch phase. For anything under 30 days old, we use stage gates instead:

  • Day 7: Did at least one person who was not us use it?
  • Day 14: Did anyone come back without being prompted?
  • Day 30: Is there a single acquisition channel that works without manual effort?

If a product clears all three, it graduates to the monthly kill-criteria cycle. If it fails stage gate 1, we either pivot the positioning or park it immediately. No point optimizing retention for a product nobody found in the first place.

Cross-portfolio patterns

Running 14 products taught us things that no single product could:

Multi-domain products outperform single-domain ones for locale-specific queries. Živá Fotka on zivafotka.cz gets Czech traffic that alivephoto.online never would. This insight changed how we launch new products for non-English markets.

B2B products with a clear deliverable (Audit Vibe Coding: you get a report) convert better than B2B products with continuous value (Voice Tables: you get an ongoing workspace). The deliverable creates urgency. The workspace creates evaluation paralysis.

Products where the AI output is shareable (a pet portrait from Pet Imagination, a song from Magical Song) have lower CAC than products where the output is private (a conflict resolution verdict, a personality profile). Shareability is free distribution.

What we would do differently

If we started over, three changes:

Set kill criteria before launch, not after. We built the framework retroactively. By the time we had it, emotional attachment to products was already baked in. Writing "this product dies if X does not happen by date Y" on day zero is uncomfortable but it makes the actual decision trivial.

Track ops cost from week one. We did not start measuring manual hours per product until month 3. By then, some products had accumulated invisible maintenance debt: manual email replies, content updates, deployment babysitting. If we had tracked it earlier, two products would have been flagged sooner.

Kill faster. Three months of flat metrics is too generous. Six weeks of stagnation with no testable hypothesis is a clearer signal. The products we eventually sunset all showed the same pattern: a brief spike at launch, a plateau, and then hope masquerading as a strategy.

The portfolio today

We keep building. The studio model means a killed product frees up attention for a new experiment. The framework is not about being ruthless. It is about being honest with the data so that sunk cost does not eat the resources that could go toward the next product with actual traction.

Full portfolio: inithouse.com

Jakub, builder @ Inithouse

on June 26, 2026
  1. 1

    This framework removes gut feeling from the kill decision, which is the hard part. Most founders delay killing products because of the emotional attachment - you built it, shipped it, got early users. This data-driven approach (especially the ops cost axis) gives you permission to move on. The multi-domain insight for Zivá Fotka is gold too - that one shift in thinking probably unlocked months of runway elsewhere. Did you find that the 30-day stage gate changed your launch approach fundamentally, or was it more of a "we wish we had done this sooner" realization?