23
132 Comments

10 testers, almost nobody came back on day 2. Fix activation first or get more testers first?

My cofounder and I are 17. We built an iOS app that turns one photo of your fridge or a grocery receipt into recipes you can cook with what you already have.

First cohort was 10 testers. Almost nobody opened it again the next day.

We think we know why. The first run asks you to pick a mode before you have seen anything work, and the recipe list still surfaces things that need ingredients you do not have. So the one thing the app exists to prove never actually gets proven in the first session.

The argument we keep having:

  1. Fix the first run, then recruit cohort 2. Slower, but the retention number would mean something.
  2. Recruit 20 to 30 people now and treat cohort 1 as a funnel measurement, because 10 people is not enough signal to redesign around.

We went with 2, and we are wiring day 1 and day 7 events before the next batch of testers goes out.

For anyone who has shipped a consumer app with a bad first cohort: did fixing activation before adding users actually work, or did you need the volume first to even see where the real drop-off was?

on September 22, 2026
  1. 1

    Good point. Did you test that with users before committing to it?

  2. 1

    The framing of the debate is what I'd push back on: at n=10 the retention number was never going to be the useful output, so "more volume to see the drop-off" mostly buys you a more precise version of a signal you already have. You've already stated the diagnosis in the post — the mode picker fires before anyone has seen the app work, and the recipe list surfaces dishes needing ingredients the user doesn't have — and neither of those conclusions gets stronger with 30 testers. The instrumentation I'd add isn't day 1 / day 7, it's a first-session funnel: photo taken → items recognized → at least one recipe shown → at least one recipe shown with zero missing ingredients → recipe opened. That last step is your actual activation event, because it's the only one where the app has kept its promise, and I'd bet cohort 1's number on it is close to zero. Concretely for cohort 2, I'd kill the mode picker entirely (pick a default, let people switch later) and hard-filter the first session to zero-missing-ingredient recipes even if the results are boring — boring-but-cookable beats aspirational-but-impossible on day 1. Two questions: how many of the 10 ever reached a cookable-tonight recipe, and have you written down what activation rate cohort 2 has to hit for you to call the fix successful? Deciding that number before you recruit is what stops 20-30 people from becoming another ambiguous data point.

  3. 1

    Good point. Did you test that with users before committing to it?

  4. 1

    Great insights here. I’d be interested to know what you would do differently if you started again.

  5. 1

    I’d fix the first run before bringing in a bigger group. You already found a pretty significant issue: people aren’t getting to experience the main reason they downloaded the app. I’d want to see what happens when those first few minutes actually deliver on the promise, then see what the next cohort does.

  6. 1

    Good point. Did you test that with users before committing to it?

  7. 1

    I like that you’re wiring day 1 and day 7 events before expanding the cohort; otherwise more users would mostly create a larger mystery. Could you define activation as “a recipe cooked” rather than just a recipe generated, so the first session tests the real value?

  8. 1

    Great to see , young founders, how are you getting this testers?

  9. 1

    Genuinely impressive you two are doing this at 17 — and instrumenting day 1/day 7 events before the next batch, rather than just guessing again, is the right instinct.

    I don't have consumer-app retention experience to speak from directly, so take this as an outside view rather than a lesson learned: my own rule of thumb in technical work is proving something small is actually correct before scaling it up, since problems get way harder to isolate once there's more noise in the system. Your plan for cohort 2 feels like a reasonable middle ground though — you're not skipping the diagnosis, just widening the sample before drawing conclusions from 10 people. Curious what the wider cohort actually shows.

  10. 1

    Good point. Did you test that with users before committing to it?

  11. 1

    With only 10 testers I'd just message the ones who didn't come back and ask what they expected to happen. Five short replies will probably tell you more than day 7 events from 30 people. And if a recipe needs stuff they don't have, that's likely the moment they closed it.

  12. 1

    Good point. Did you test that with users before committing to it?

  13. 1

    Went through the same thing with our first cohort. The thing that unblocked us wasn't more testers, it was naming the one event that means the promise actually landed and logging every step up to it. Ten people is enough to see where the first session breaks; it only stops being enough when you start counting percentages. Fix the first run and instrument it in the same pass, those aren't alternatives.

  14. 1

    Your diagnosis sounds right to me: if the first session never proves the core promise, day-2 retention mostly measures the onboarding. Could you let the very first run skip the mode choice and use a sample fridge photo, so people see one good recipe before they're asked to decide anything? Then cohort 2 tells you about the product, not the setup screen.

  15. 2

    How did you decide this was worth building in the first place?

  16. 1

    Thanks for sharing the numbers and the process behind them. It makes the journey much easier to understand. Looking forward to seeing where you take this.

  17. 1

    The day 1/day 7 event tracking seems like the real fix here regardless of which path was "right," at least cohort 2 gives you real numbers instead of two competing guesses. Curious what day 7 ends up looking like once it's live.

  18. 1

    One thing to watch in cohort 2: from your replies it sounds like you're now doing both, removing the mode picker AND recruiting 20 to 30 more testers. That's reasonable, but it means if retention improves you won't know whether it was the new first run or just different people.

    Two cheap ways around it. Either keep the old first run for a handful of testers so you have something to compare against, or write down right now what number would make you call the fix a success, before cohort 2 comes in. The second one costs nothing and stops you from reading whatever comes back as a win.

    The receipt-scan-as-re-entry idea from the thread seems like the sharpest insight here. Days between the first cookable recipe and the next scan is a much more honest number than day 2 for a dinner app.

  19. 1

    Aloha, I'm about to step into the same waters as you. I'm working towards gathering my first cohort of between 50-100 users to test retention over a 2-4 week span. I'm planning to let everyone know that the trail period is free and seeing if anyone actually converts to paid during the test. We're still pondering offering a discounted subscription to the App but we decided to only test 1 condition at a time. Have you sprung your second cohort yet and how did it go?

  20. 1

    Good point. Did you test that with users before committing to it?

  21. 1

    Since you're considering a sample fridge on screen one, I'd keep its success event separate from a recipe generated from the tester's own scan. Otherwise the sample could improve your 'first cookable' number while the camera or ingredient-confirmation step still loses people. One field on the event—sample vs own photo vs receipt—would let you see that distinction. For the small group you watch live, also note whether they reached the result unaided or after a prompt. That should make cohort 2 easier to interpret without adding another onboarding screen. I haven't tested the app; this is based on your updates here.

  22. 1

    Good point. Did you test that with users before committing to it?

  23. 1

    fix activation first, no question. 10 testers with 90% churn tells you the product isn't delivering on its promise fast enough — more testers just means more people confirming the same problem. the specific issue you described (picking a mode before seeing value) is a classic time-to-magic bottleneck. strip the first run down to one action: photo → recipe with ingredients you have. make that work in under 60 seconds. retention follows from a moment that actually lands, not from a bigger sample size.

    1. 1

      under 60 seconds from photo to a recipe made only of what you have is the target. mode picker is going, filter becoming default, and time from open to that recipe is the first number we log.

  24. 1

    One extra thing I’d separate from activation is usage cadence. Dinner planning may not naturally be a day-2 habit, so even a good first session can look like weak retention. I’d pair the ‘first cookable recipe’ event with the next natural trigger — a new grocery receipt, a refreshed fridge scan, or a saved recipe becoming relevant — and measure whether people return at that moment. If users reach a cookable result but ignore the next trigger, that’s a retention problem; if they never reach it, it’s still activation. That gives the next cohort a cleaner question than day-2 alone.

    1. 1

      the cadence point is the one we keep circling. dinner is not a daily job for most of cohort 1 and grocery runs are weekly, so the receipt scan is probably the real re-entry point rather than day 2. plan is to log first cookable, then log the next scan (fridge or receipt) as its own event and look at days between them instead of a day 2 yes or no. reach cookable and never scan again is the retention bucket, never reach it is still activation.

  25. 1

    With only 10 testers, I’d fix activation before recruiting more—but treat it as one narrow experiment, not a full onboarding rewrite. On Monday, pick the first moment that proves value (for this app: photo → one recipe someone would actually cook), remove the mode choice and ingredient cleanup from that path, then watch whether testers complete it and return the next day. Talk to the 2–3 who did not return; ask what they expected to happen after the first scan. Ship one change, talk to users, and kill it if day-2 return doesn’t move after a week.

    I keep a free Monday Launch Checklist for that weekly triage:
    https://eastwestkonnex.gumroad.com/l/monday-launch-checklist

    What’s the one activation event you’re measuring?

    1. 1

      one event: first recipe shown with zero missing ingredients, with seconds from app open attached. everything else on day 1 is secondary to that until most people hit it.

      1. 1

        Zero-missing-ingredients recipe plus time-to-first is a clean activation line. With only 10 testers I'd also tag where the miss happens — pantry incomplete, picker slow, or bounce before pantry — because the same missed event wants three different fixes. Otherwise day 2 still dies and you don't know which lever to pull.

  26. 1

    Choosing (2) was right, but there's a third option cheaper than both: re-run the same 10 testers after the fix instead of recruiting fresh ones. A paired comparison — same people, before and after — controls for everything except your change, which 20 new strangers can't do at this stage. And keep the first-run fix as small as possible: auto-pick the mode (or infer it from the photo) so nobody chooses before seeing value, and label each recipe 'you have everything' vs 'one trip away' — that turns the missing-ingredients problem from a failure state into a shopping list. Big redesigns change too many variables to learn anything from one cohort.

    1. 1

      rerunning the same 10 is happening, just slower than i want because they are friends and family and half need a nudge. the 'you have everything' vs 'one trip away' label is a good middle ground. karthik shipped the only-what-i-have filter as a toggle and it is becoming the default, so one-trip-away would live as a second tab instead of being mixed into the first list. passing that to him.

  27. 1

    Good point. Did you test that with users before committing to it?

  28. 1

    Good point. Did you test that with users before committing to it?

  29. 1

    Good point. Did you test that with users before committing to it?

  30. 1

    Good point. Did you test that with users before committing to it?

  31. 1

    Ten people with wired day-1 and day-7 events beats 30 people guessing. You're already solving the hard part - the day-1 drop is visible because you're measuring it. That clarity is worth way more than volume right now. Most teams would ship cohort 2 and still not know whether it was the mode selection or the recipe list. You'll know.

  32. 1

    Interesting. How are you measuring whether it is working?

  33. 1

    I'd fix the first run first.

    10 users is obviously a very small sample, but if almost nobody came back AND you already see a pretty clear problem in the first session, getting 30 more people through the same flow probably won't tell you much more.

    I had similar issue with one of my apps. Getting more installs felt like progress, but in reality I was just sending more people into the same weak funnel :)

    I'd make the first value moment as fast as possible, then bring the next 20-30 testers and compare.

  34. 1

    I’d fix the first-run promise before recruiting more. Give each tester a fast “aha” moment with only recipes that are actually cookable from the scan, then instrument the first session as a funnel: scan completed → recipe opened → ingredients checked → recipe started. A larger cohort is useful after that path is trustworthy; otherwise you mostly measure confusion.

  35. 1

    I’d keep the next cohort small while instrumenting the first-session funnel, rather than choosing between learning and volume. Define activation as the first “aha” (a recipe they can actually cook) and track time-to-aha plus the exact ingredient mismatch; then recruit more users once that step is reliably working. That gives you better retention signal without waiting for a huge sample.

  36. 1

    honestly with only 10 testers you don't have enough signal yet to know if it's a real activation problem or just noise. 2 people not coming back could be "the onboarding is broken" or it could just as easily be "these 2 people were never your target user in the first place."

    before touching the funnel, i'd want to know why they didn't come back, not just that they didn't. did they hit an error, get confused, get what they needed in one session and had no reason to return, or just forget? those are completely different problems with completely different fixes, and "more testers" vs "fix activation" assumes you already know which one it is.

    get on a call with the ones who didn't come back if you can, even a short one. actual words from an actual person will tell you more than staring at a 2-person dropoff number ever will.

    1. 1

      agreed, and i have been messaging each of them. what came back so far: no errors, no complaints, two described the mode picker unprompted, nobody described a recipe. so the wall is the first screen for now. the target user point is fair too, all 10 were my own network, which is exactly why cohort 2 is strangers from here and the tester subs.

  37. 1

    Fixing the first run before a bigger cohort makes sense here. The first session never showed the one thing the app is for, so more testers would mostly repeat that. Day 1 and day 7 on the next batch should show whether the first run was the drop-off.

  38. 1

    Update for everyone who weighed in: we went with fix first run, then measure a second cohort, and the recruiting post for that cohort is up here https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f . TestFlight link is in it along with the three things I want reported back (time to first cookable recipe, what the scan got wrong, would you open it tomorrow). If you commented in this thread you already know where the holes are, so you are exactly who I want breaking it. I test yours back at the same depth.

  39. 1

    honestly option 1. if the first run never shows the magic, 20 more testers will just confirm that. skip the mode picker, show one recipe from the photo first and ask questions after

    1. 1

      yes. mode picker is going, and recipes stay hidden until the scan finds at least 5 items so the first thing on screen is a short list you can actually cook, not a form.

  40. 1

    At n=10, I’d use cohort 1 to define one activation contract before adding volume: photo in → first genuinely cookable recipe in under a minute → save or start cooking. Instrument each drop-off, watch 3–5 sessions, and ask every non-returner one specific question. Ship the highest-frequency fix, then rerun a small cohort; treat day-2 retention as secondary until first value is reliable.

    Disclosure: I sell a short $17 PDF checklist myself, so I’m biased. It’s the same kind of lightweight validation/planning aid: https://dropmountainltd-ux.github.io/checklist/

  41. 1

    How did you decide this was worth building in the first place?

  42. 1

    I’ve run into the same thing with a dividend calendar / stock screener. The first session is easy to overpack, then the user never gets to the one job they came for. I’m trimming ours back to a single US-ticker scan and making the next useful action obvious. The free version is here if useful: https://tools.evolutionfreedomltd.co.uk/stocks

  43. 1

    Your diagnosis is probably right. With a cohort this small, I would watch two sessions and write down the exact seconds between install and the first usable recipe, plus where each person stalls. If nobody reaches that first meal, new testers only give you more versions of the same answer.

    1. 1

      agreed, and the seconds-to-first-usable-recipe number is the first thing we are asking cohort 2 to report, along with the exact screen they stalled on. recruiting post is up: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  44. 1

    Fix activation first. Always. More testers before that just gives you more people to watch the same problem happen.

    The question I'd ask before anything else: did anyone in the 10 reach one moment where the app delivered on its promise? Not just "opened the app" or "saw a recipe" — but the specific moment where someone saw a recipe they could actually cook with what they had. If even one or two people hit that moment and didn't come back, you have a retention problem. If nobody hit that moment, you have an activation problem. Those need completely different fixes.

    With 10 testers you can't trust percentages, but you can trust "did this happen or not." So go back through those sessions and ask: did the product demonstrate the thing it promises? If the answer is unclear, that's your diagnosis right there.

    The trap with recruiting more testers is it feels like progress. More people, more data, more signal. But at 10 you usually have enough to find the break — you just have to look harder at what actually happened in session one rather than averaging across it. What was the last thing each person did before they stopped?

    1. 1

      your split is the one we are using now: activation problem if nobody hit a cookable recipe, retention problem if they hit it and still left. cohort 1 data says almost nobody hit it, so activation first. cohort 2 will be measured on exactly that moment, post is here if you want in: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  45. 1

    Fix activation first. More testers just churns more people on the same screen. With my TestFlight betas the thing that actually moved the needle was watching one tester use it live — the stall point is never where you think it is. Pick the one action that means "this works" and cut everything before it.

    1. 1

      the stall point never being where you think it is matches cohort 1 exactly. i assumed the recipe list and it turned out to be the mode picker two screens earlier. watching a few cohort 2 sessions live is in the plan. if you have an iphone and want to be one of them: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  46. 1

    Activation, 100%. More testers on top of a broken first run just scales the drop-off. You already nailed it: nobody should pick a mode before seeing the app do its one thing, and the first recipe list should only show what they can actually cook right now. Land that first win in session one, then worry about volume.

  47. 1

    Fix activation first. More testers just multiply a leaky first session. I'd watch one clear aha moment (first wall placed / first win vs AI) before spending time on acquisition. Once that sticks, then scale testers.

  48. 1

    How did you decide this was worth building in the first place?

  49. 1

    Love this angle, honestly. What made you look into it in the first place?

  50. 1

    Really like how openly you're working through this. "First cookable" is a great way to frame the activation moment.

    One thing I haven't seen mentioned yet: right now the first real value only shows up after a camera permission prompt and a photo. Some people will drop off there just because their fridge isn't in front of them when they install. What if the first screen showed a sample fridge photo with a couple of fully cookable recipes, before asking for anything? The promise gets proven in 10 seconds, and the real scan becomes "ok now try it with yours."

    On the return side, since you already think daily use doesn't match the behavior, the receipt scan might be your natural re-entry point. People get home from the store with a receipt in hand, and that's exactly when "what can I make with this?" matters most. A prompt tied to that moment could tell you more than a generic day 2 push.

    Curious to see cohort 2's numbers, good luck to you both!

    1. 1

      this is the one idea in the thread we had not considered. a sample fridge on screen one proves the promise before we ask for the camera, and it costs nothing to build. passing it to karthik today, it fits inside the three screen first run he is rebuilding. if you want to see the current version before that lands: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  51. 1

    Measurement clarity beats cohort size. Ten people with event streams that show when the app actually proved its value (first cookable recipe visible + ready to use in < 30s) beats 100 guessing whether the "opened again" drop was UI friction or just wrong timing. Your choice to wire Day 1 / Day 7 events before the next batch is perfect - you'll see exactly which moment testers drop. That moment is your real activation problem, not the headcount.

  52. 1

    Fix activation with the 10 you already have — more testers before that just scales the confusion.

    Watch 2-3 sessions (recorded is fine). Day-2 drop is usually an onboarding cliff: the 'first win' takes too long to reach, so cut every step between signup and it.

    One underused lever: a 40-60s replay/demo embedded in the day-1 email and the empty state. On our own product the cohort that watches it returns day-2 at roughly 2x. Happy to share the script structure.

    1. 1

      yes, would love the script structure. we are recording a 60 to 90 second demo this week anyway so the timing works. and if you have an iphone the cohort 2 post is up with the link: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  53. 1

    Your reply about keeping track of the remaining ingredients raises a practical question: how will the app know what someone actually cooked or used up? A recipe opened isn't necessarily a meal made. I'd test a simple confirmation after cooking before building around that return visit. Otherwise yesterday's fridge photo could lead to another recipe with missing ingredients, even if the first run worked.

    1. 1

      honest answer is it does not know yet. a recipe opened is the only signal we have, so the inventory is only as fresh as the last scan. the cheap version of your suggestion, one tap after cooking that removes those items, is going on the list right behind the first run fix. until then the second scan matters more than the second open.

  54. 1

    Definitely fix activation first. When we went through closed testing, day 2 drop-off was our biggest signal that the initial onboarding friction was too high.

    Pouring more testers into a leaky funnel just burns through relationships without giving you useful feedback. What worked for us was:

    1. Shortening the time-to-first-value (let them see the core result in < 30 seconds without mandatory sign-ups).
    2. Setting up a direct channel (a quick 2-question DM or form) asking the dropped-off testers where they got stuck. Almost every time, it was a subtle UX confusion or missing file format we assumed was obvious.

    Once your day 2 retention stabilizes for even 4-5 testers, scaling up to 20+ testers becomes ten times smoother.

  55. 1

    It’s a real chicken-and-egg dilemma. My take is that it’s really difficult to understand what users want. The general answer would be to talk to users, as described by Y Combinator and Rob Fitzpatrick.

    The other thing is having the right data. How many people check out your app? How many download it? What do they do when they open the app? Can you reduce the friction?

    I help people make better decisions about warehouse design and improve their understanding of warehouse operations. For this, I’m working on www.rackflow.app , an online editor combined with a simulation mode that lets users see how a warehouse operates.

    I’ve conducted 16 user interviews so far, and my takeaway was:

    1. Improve the usability of the editor so that it works more like a drawing editor.
    2. Build more trust; no one wants to risk their credibility.
    3. Most random traffic is just looking around; people don’t have a real task they want to solve. But feedback is below 5 %...

    My internal reports showed me that most users registered but didn’t figure out how to change the model or even start the simulation. So I’m currently improving the information and documentation to attract people who are actively looking for a solution, add a preconfigured demo warehouse, and I’m also reworking parts of the user interface to improve the overall interaction and experience.

    I wish you good luck! ;-)

  56. 1

    I’d fix the first-run promise before adding a big second cohort. Ten users is too small for a clean retention curve, but it is enough to see that asking for a choice before showing the core win is risky. I’d make the first session prove “here’s a recipe you can cook now”, instrument that step, then bring in 10–20 more testers. That gives you a better funnel without waiting forever.

  57. 1

    We run a small consumer web app, and the thing that finally made day 2 make sense for us was separating two questions: "did they get the aha in session one?" and "do they have any reason to open it tomorrow?" Those fail for different reasons, and a single day-2 number blends them together.

    For the second one, it helped to design a tomorrow-shaped hook on purpose: something that is literally different each day (for you maybe "here's what you can still cook with what's left"). If the app shows the same thing tomorrow as today, even a great first session won't bring people back.

    I'd fix the first run before cohort 2, but keep it tiny. 5 people you can watch live will tell you more than 30 you can only read in a dashboard.

    1. 1

      the tomorrow-shaped hook is the part i had not designed on purpose. what is different tomorrow is what is left and what is oldest, so 'what you can still cook with what is left' fits without inventing anything. and yes on keeping the watched group tiny, a handful of live sessions on top of the wider cohort. if you have an iphone and want to be one of the watched ones: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  58. 1

    Late to the thread but this is the exact trap I almost fell into. One data point that helped me separate the two questions: benchmark consumer D1 retention is roughly 20-25%, so 10 testers giving you ~0% on day 2 is not a sample-size problem, it's a signal. I'd fix first because volume only multiplies whatever your baseline is. The framework I'd borrow is the Sean Ellis test: of those 10, how many would be disappointed if the app vanished? If the answer is zero, 30 more testers just confirm it louder. Your "first cookable" event is the right north star — I'd add one guardrail metric though: time from open to that moment. If that shrinks and day-2 doesn't move, the problem is relevance, not activation.

    1. 1

      the sean ellis question is one i have not asked the 10 yet, and i think the honest answer is zero disappointed, which says the same thing you did. time from open to first cookable is going in as the guardrail. if that number shrinks and day 2 stays flat, the problem is relevance and no amount of onboarding fixes it.

  59. 1

    I would fix the known first-session problem before bringing in a much larger group. Ten people may not provide a reliable retention percentage, but they are enough to reveal that the product’s main promise was never demonstrated.

    As co-founder of SynDiary, I’m facing a related challenge. The product becomes more valuable as personal data grows, but a new user still needs one worthwhile result at the beginning. For us, information they already have such as their calendar may be the easiest starting point.

    I would fix the first wall, test again with a small group of strangers, and speak personally with everyone who does not return. Silence rarely explains whether someone encountered friction or simply saw no reason to come back.

  60. 1

    You already diagnosed the real problem in your own post the app's core promise never gets proven in session one, so no amount of volume fixes that. Option 2 without also fixing the obvious first-run issue risks just measuring the same failure at a bigger sample size.

    Worth doing both, cheaply: ship the smallest possible fix to session one (even hardcoding a guaranteed-good first recipe using only common pantry items) before recruiting cohort 2, so the funnel data you're about to wire up isn't measuring a problem you already know exists.

    1. 1

      hardcoding a guaranteed-good first recipe from common pantry items lands close to what shovra suggested above, a sample fridge on screen one. both prove the promise before the camera prompt. the filter fix ships first, the sample screen is going to karthik with the three screen first run.

  61. 1

    I’ve seen a similar distinction with business software: getting someone into the product is much easier than getting them through a complete real workflow.

    For example, a user can sign up, look around the dashboard and even try a feature, but that doesn’t necessarily mean they experienced the actual value. The stronger signal is whether they complete the job they came there to do and would use the workflow again.

    So I’d separate “activation” into two steps: reaching the first useful result, and successfully completing the real task. Then retention tells you whether that result was valuable enough to repeat.

    With only 10 testers, I’d probably use the cohort mainly to discover friction and interview the people who didn’t return, rather than treating the retention percentage as a reliable measurement.

  62. 1

    You went with 2, and I think that was right — but I'd check the premise before you redesign the first run, because I measured the same situation this week and the data said something different from what I assumed.

    We had 4 people sign up. 3 of the 4 completed the entire first run — added a site, installed the snippet, ran their first check. So activation was not the problem at all. None of them came back. What actually happened was that after the first run there was nothing scheduled for them: the free tier had no recurring work, so the product had no new data to show, and their first contact from us was a weekly digest six days later showing exactly the numbers they'd already seen. Day 2 had nothing in it.

    So the thing I'd instrument before touching the UI is two separate events: "completed first run" and "returned". If your 10 testers are failing the first one, fix activation. If they're passing it and still not coming back, activation is fine and the gap is that day 2 is empty — which is a very different fix (something happens while they're gone, and a reason to look) and much cheaper than a redesign.

    With 10 testers you can just ask the 7 who didn't return. n=10 isn't going to give you a statistically clean answer either way, and one sentence from a person who left beats a funnel chart at that size.

    (I build an AI-search visibility tool, so the numbers above are our own — small sample, take them as one data point.)

    1. 1

      fair point and worth checking. our case looks different because most of the 10 never reached a cookable recipe at all, so we cannot yet tell whether a return hook is missing. once first run is fixed and day 1 and day 7 events are logging, your scenario is the next thing we test for. cohort 2 post if you want to be one of the data points: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  63. 1

    Your framing suggests a simple gate for cohort 2: don’t add volume until most users can reach the “cookable recipe” moment in session one. I’d also separate “didn’t return” into no need vs. failure by asking one forced-choice question at exit plus an open-text follow-up. That makes the small sample useful without pretending it’s statistically precise. Are day-1 events already capturing the missing-ingredient path?

    1. 1

      not yet. cohort 1 had zero events, everything is reconstructed from conversations. day 1 events go in this week and missing-ingredient count on each recipe shown is one of them, since that is the exact thing being fixed. the exit question is a good add, one forced choice (no need, did not work, forgot) plus a text box.

  64. 1

    Fix activation first. If almost nobody is coming back on day 2, adding more testers will mostly give you more of the same signal.

    First figure out where users are dropping off, make the core value obvious as quickly as possible, and get a few testers to reach that “aha” moment consistently. Once people are sticking around, then bring in more testers to validate whether the improvement holds at a larger scale.

  65. 1

    I’d treat this as two experiments rather than choosing one forever: first narrow the first-run path so the promised outcome is testable (photo to recipe with no missing ingredients), then send a small cohort and interview every non-returner. Day-1/day-7 events will tell you where people stop, but the conversations will tell you whether the outcome was valuable enough to repeat. If the cookable-recipe moment rarely happens, more users mostly add noise.

  66. 1

    I'd fix the first-run promise before scaling the cohort. For cohort 2, log one event when a user gets a recipe with zero missing ingredients, then ask a few non-returners what blocked them. That should help separate activation from low repeat need before you add more volume.

  67. 1

    Worth flagging before cohort 2 goes out: you're now testing three different hypotheses at once (activation, frequency mismatch, wrong metric) with 20-30 people. That's a small sample to cleanly separate three explanations — if the split between "reached cookable" and "didn't" comes out roughly even, you might not have enough people in each bucket to tell signal from noise. Might be worth deciding in advance how big a gap between groups you'd actually trust as real, so you're not stuck debating whether 6-out-of-14 vs 2-out-of-11 means anything.

    1. 1

      fair, at 20 to 30 people the split will not be clean. the pre-commit i am comfortable with: if fewer than half reach first cookable in session one, activation is still broken and nothing else gets analyzed. if more than half reach it and day 7 is still near zero, that becomes the relevance question. anything in between and we watch more sessions instead of arguing about the numbers.

      1. 1

        That's a solid gate for the two clear ends. The one edge worth naming in advance: "watch more sessions" for the middle band needs its own stopping rule, or it can quietly become the permanent answer — more sessions arrive, the split stays messy, and you keep watching without ever re-hitting the >50%/day-7 test. Might be worth deciding now how many more sessions "watch more" buys you before you force a call either way, even if that call is just "still ambiguous, ship the activation fix anyway since it's needed regardless."

  68. 1

    This is useful. How are you finding your first users so far?

    1. 1

      cohort 1 was my own network, 10 friends and family, which is part of why the signal was weak. cohort 2 is coming from this thread, comments in the reddit tester subs, a few discord beta servers and two beta listing sites. this thread has been the best of those by a wide margin. post with the link is here if you want in: https://www.indiehackers.com/post/cohort-2-i-need-20-of-you-to-install-peeka-this-week-and-try-to-break-the-first-run-316476523f

  69. 1

    At 10 testers, the useful question may not be “fix first or recruit more?” It may be: did anyone reach the one moment that proves the app works?

    Your first cohort is small, but it already exposed a concrete break. The first run asks people to choose a mode before the product has shown anything working. Then the recipe list can include dishes that need ingredients they do not have. So the promise — cook with what you already have — may never be demonstrated in session one.

    I would not treat the 10-person cohort as a retention rate. Ten is enough to find friction, not enough to redesign around a percentage. But it is also too early to recruit 20-30 people just to collect a cleaner record of the same break.

    The smallest decisive step is to define one activation event: “photo in → at least one recipe visible that only uses ingredients the app detected → recipe opened.” Then run a small fixed flow with 5-10 new testers: skip mode selection until after that moment, and hide recipes that require missing ingredients.

    Track only four things per person:

    1. Did they complete the photo step?
    2. Did they reach the first cookable recipe, and how long did it take?
    3. Did they open or save it?
    4. Did they return on day 2 without a personal reminder?

    Then use these decision lines:

    • If at least 70% reach and open a cookable recipe, but fewer than 10% return unprompted on day 2, activation is probably no longer the main bottleneck. Test whether people simply lack a next-day reason to cook.
    • If fewer than 50% reach the cookable recipe, do not scale the cohort yet. Find the exact exit point: mode choice, photo permission, wait time, ingredient confirmation, or recipe quality.
    • If people reach the recipe but do not open it, the issue may be trust or recipe relevance, not activation.

    Your plan to wire day 1 and day 7 events is right. I would add the cookable-recipe moment before recruiting the next batch, because otherwise day 2 can still collapse several different failures into one number.

    A few facts would sharpen the diagnosis: of the original 10, how many saw at least one recipe they could actually cook? Did any return later without being prompted? What did they use before the app when they had leftover ingredients? And what would make tomorrow’s open obviously useful — remaining ingredients, a meal plan, or a saved recipe?

    These observations come only from what you publicly wrote. This is a scored diagnosis under uncertainty, not a verdict; validation changes confidence, not certainty, and the decision remains yours.

    1. 1

      Straight answers to your four, with the caveat that anything from cohort 1 is reconstructed from conversations and screenshots, not logged.

      How many saw a recipe they could actually cook: I cannot prove it for a single person. When I went back and asked, people described the mode picker and the photo step. Nobody described a recipe. That is the tell.

      Did anyone return unprompted: no. Every second open I know about happened after I texted someone.

      What they used before: nothing. They open the fridge, look at it, decide or order out. I am not replacing a system, I am replacing a glance, which is a harder thing to beat.

      What would make tomorrow's open obviously useful: my guess is remaining ingredients, because after you cook one thing the fridge changed and the app is the only thing that knows how. Meal plan feels like a different product wearing our clothes.

      Your decision lines are going in as written, 70 percent reach with under 10 percent unprompted return means stop blaming onboarding. That is the part I did not have.

      If you want to see the first run instead of my description of it: https://testflight.apple.com/join/AZm9hsB8

  70. 1

    The decision to fix activation first is right, but the measurement system you're building into cohort 2 is doing more work than the cohort size. You already know 10 people is too small to separate activation from fit. But "first cookable recipe" as your measurement zero is the real insight here.

    The trap is treating day 2 return as the truth and then trying to interpret why it happened. If you wire events first, you'll see whether the people who actually got a working recipe came back, separate from the people who bounced at mode select. That's the measurement that lets you iterate.

    One that stings in your reply: you ran cohort 1 blind. Events logged before testers arrive, not after. That's the leverage point. Small cohorts can't show you statistically significant funnels, but they can show you whether your instrumentation actually captures the moment the app works. Ten people with clean event streams beat 100 with guessing.

    1. 1

      Agreed, and the part you named is the one that actually hurt.

      Cohort 1 went out blind. I can fix every other mistake in that list. That data is gone permanently, so all I have is people telling me weeks later what they think they remember doing.

      Cohort 2 does not get a link until first cookable, starts cooking and save or complete are all firing, plus time from app open on the first one. If the people who hit first cookable come back and the people who bounced at mode select do not, that is an activation problem and I keep working on the product. If both groups look the same, the honest read is that nobody needs this often enough and I have a much bigger question to answer.

      Link if you want to watch it break in person: https://testflight.apple.com/join/AZm9hsB8

  71. 1

    Thanks for sharing the numbers, that makes it much easier to follow.

  72. 1

    Nice work shipping it. What has been the biggest challenge since launch?

  73. 1

    Appreciate the honesty here, most people only share the wins.

  74. 1

    Nice progress. What is the next thing you are focusing on?

  75. 1

    Fixing activation first was the right call, but "fix the first run" can be too broad to act on. When I shipped my own iOS tool, the trap was treating onboarding as one screen instead of one moment — the single action that proves the thing works. For you that's clearly photo in, a cookable recipe out, zero missing ingredients. I'd cut mode selection entirely for cohort 1 and hard-filter to fully-cookable results, even if that means three recipes instead of thirty. Ten testers can't show you where a funnel leaks, but they can tell you whether that one magic moment lands. Did any tester actually reach a cook-it moment on day one?

    1. 1

      Not one I can prove, and that is the whole answer.

      I went back and asked. People described the mode picker and taking the photo. Nobody described a recipe, let alone cooking one. With no events wired I cannot say whether they hit a cookable result and ignored it or never got there, but the way they talk about it points at never got there.

      Hard filter to fully cookable is going in exactly as you describe, three recipes instead of thirty. Mode selection is getting killed, not deferred. The bet is that a short list of things you can genuinely make beats a long list where you have to go shopping first.

      https://testflight.apple.com/join/AZm9hsB8 if you want to see whether the fix actually lands.

      1. 1

        Thanks for checking with them, and for the invite. One edge case to watch with the hard filter: zero matches. Make that explicit and offer a way to correct the detected ingredients, so it doesn't become a new dead end. Seeing a few testers reach a recipe without coaching would be a useful next check.

  76. 1

    Curious how long it took before you saw the first real results?

    1. 1

      Two weeks, and the first real result was a negative one.

      Cohort 1 went out Sep 7. Day 2 return was visible as close to zero almost immediately, but it took until about a week ago to work out why, because I had no events running and had to reconstruct it from conversations.

      Posting the bad number publicly got me more usable input in a few hours than two weeks of staring at it alone did.

  77. 1

    It was a test. That is what testers are for. If you learned something that needs to be fixed, then fix it before you test again. Otherwise, you are just going to verify flaws that you already know about.

    1. 1

      Fair, and this is the strongest version of the argument against what I chose.

      The reason I went the other way: 10 people all from my own network is not a test, it is 10 friends being polite. I know one flaw for certain, the mode picker, and I am fixing that regardless. What I do not know is whether fixing it moves anything, and with a sample that small and that biased I cannot tell a broken first run from nobody wanting this.

      So both at once. Product fixes and instrumentation land first, then 20 to 30 strangers, not friends. If day 2 is still flat after that I stop blaming onboarding and start questioning the need.

  78. 1

    Clear and practical, thanks. Did anything surprise you along the way?

  79. 1

    Nice work shipping it. What has been the biggest challenge since launch?

  80. 1

    Appreciate the honesty here, most people only share the wins.

  81. 1

    Ten people is enough to ask, not enough to measure. Message every one of them and ask what happened the second day. You'll get more out of ten honest answers than from wiring up events for cohort two. The thing that moved day two for me wasn't a fix in the product, it was a person saying hello.

    1. 1

      I have started doing exactly that and you are right that it beats the dashboard I do not have yet.

      What came back so far: nobody thought leaving was worth reporting. No complaints, no bug reports, just silence, and every one of them was friendly about it when asked. Two described the mode picker unprompted, which is how I know that screen is the wall.

      The part of your comment I keep thinking about is the person saying hello. Cohort 1 got a link and nothing else. Cohort 2 gets me in the thread with them.

      I am still wiring the events, because I want both. Ten honest answers tell me why, events tell me how many.

      1. 1

        We're in the same spot with DOER right now. Enough curiosity for people to come and see what it is, sign up, then not come back. So today I wrote personally to about 40 of our early members who drifted and asked them straight: if DOER wasn't what you expected, what did you expect? And I asked them to tell me on the platform, not by email, because the point is to get them back in the door, not just get an answer.

        Silence is the normal signal. Nobody writes to say a thing was fine, they just don't come back. That's why asking works, you got the mode picker named twice without prompting, which no dashboard would have handed you. And the fact they were all friendly when asked is worth noticing. They don't hate it, they just had no reason to open it again. Being in the thread with cohort two fixes that better than any event you wire. Good luck with them.

        1. 1

          And if you want to compare notes as we both work through it, I'm around. Would be good to hear how cohort two goes.

  82. 1

    I read your post about your first 10 testers, and I’m curious about one thing.

    You mentioned that the app was surfacing recipes that required ingredients users didn’t actually have. Did your testers specifically tell you that this was why they didn’t come back, or is that something you and your cofounder identified yourselves?
    I’m asking because if that’s already a known issue in the core experience, I wonder how much adding another 20–30 users can tell you before fixing it. If I opened an app that promised recipes based on what I already have, but it suggested ingredients I don’t have, I probably wouldn’t come back either :)
    Or are you using the larger cohort specifically to find out whether that issue is actually what’s causing the drop-off?

    It also made me wonder whether you have a very simple way for users to report a specific bad result directly from the recipe itself — something like “I don’t have this ingredient.”

    I had to think about a similar feedback problem with REZYCO from a different angle. My compatibility model can only work with the answers people give it, and of course a person can always be dishonest. So I gave users a way to report someone if they discover that their questionnaire doesn’t reflect reality.
    In your case, that kind of feedback might help distinguish “the user just didn’t come back” from “the app gave them a result they knew was wrong.” Do you have anything like that built in?

    1. 1

      We identified it. No tester said it out loud, which is the uncomfortable part of your question.

      What they actually described when I went back and asked was the mode picker and the photo step. Nobody got far enough to complain about recipe quality. So the missing ingredient problem is real and I can see it in the build, but I have no evidence it is what caused the drop. It might be the second wall behind the first one.

      That is why cohort 2 is still going out, but not blind. If people clear the fixed first run and still leave, the recipe list is the problem and I will have the events to show it.

      On feedback: no, nothing like that exists in the build, and after reading your REZYCO answer I think that is a real gap. A one tap "I do not have this" on the ingredient line does two jobs at once, it corrects the inventory the detection got wrong and it logs a bad result with the reason attached. Right now a wrong scan and a bored user look identical in my data.

      Stealing it. Thank you for the actual answer instead of a take.

      https://testflight.apple.com/join/AZm9hsB8 if you want to poke at it, iPhone only.

      1. 1

        That actually makes a lot of sense. If nobody got far enough to really experience the recipe results, then I agree — you don’t know yet whether that’s the reason they left or, as you put it, the second wall behind the first one.
        And I really like your point that a wrong scan and a bored user currently look identical in the data. That’s exactly why I added reporting to REZYCO — sometimes knowing that someone stopped engaging tells you almost nothing about why.
        Please steal the idea 😄 I’ll be curious to see what you learn from it.
        And good luck with the app! As a woman, I can definitely relate to the two eternal questions: “What should I cook?” and “What can I cook with what I actually have?” 😄
        So I genuinely hope you make this work.

  83. 1

    Clear and practical, thanks. Did anything surprise you along the way?

    1. 1

      Two things.

      Nobody complained. I expected bug reports and got silence, which is worse, because a complaint tells you where to look and a silent exit tells you nothing. When I went back and asked, people were friendly about it. None of them had thought of leaving as something worth reporting.

      And the thing I was scared of was not the thing. I assumed detection accuracy would sink us, spinach coming back as kale, that kind of error. Not one person who left mentioned accuracy. They left before they got far enough for accuracy to matter, at a mode picker screen that exists for internal reasons and does nothing for them.

      The one that stings: we ran the whole cohort without day 1 and day 7 events wired. Everything else here I can go back and fix. That data I cannot get back.

  84. 1

    I’d fix the first-run promise before recruiting a larger cohort. With the current flow, more users mostly gives you a cleaner measurement of a broken experience; log the moment a user gets a recipe they can actually make, starts cooking, and returns to save or complete one, then compare those steps in cohort 2. Keep recruiting small while you iterate, and expand once that core action is happening consistently.

    1. 1

      That is the exact event we are wiring. We call it first cookable, the moment the app shows a recipe where every ingredient came from the photo, with time from app open attached. Nobody in cohort 1 has a timestamp for it, which tells you how buried it currently is.

      Starts cooking and returns to save or complete are going in as steps 2 and 3, so we get a 3 step funnel instead of one retention number that just says people left.

      Keeping cohort 2 at 20 to 30 for the reason you gave. Wide enough to see the shape of the drop, small enough that I can still talk to every person who quits.

      1. 1

        “First cookable” is a great name for that moment. Wiring the 3-step funnel—open → cookable → starts cooking → save/complete—is exactly the right instrumentation, especially when cohort 1 has zero timestamps.

  85. 1

    With only 10 testers I'd fix activation first — the sample's too small to tell if it's the product or just the wrong users. What does the day-1 experience actually look like right now?

    1. 1

      Day 1 as it stands: you open it, the first screen asks you to pick a mode before you have seen the app do anything, then you take a fridge photo, wait a few seconds for detection, confirm or correct the items it found, and land on a recipe list that still includes recipes needing two or three things you do not have.

      So the one promise, cook with what is already in there, is technically delivered on screen four and unproven on screens one through three.

      Three changes going in: kill the mode question and default to fridge only, gate the list so a recipe does not appear unless you have every ingredient, and cut first run to three screens. If day 2 is still flat after that, then it is the users or the need, and I will stop blaming the onboarding.

  86. 1

    Makes sense. Are you planning to charge for it, or keep it free for now?

    1. 1

      Free through the beta, no paywall in the build at all.

      I am not putting one in until people come back on their own, because a paywall on top of a retention problem just hides the retention problem behind a conversion number.

      Where it probably lands: free for a limited number of scans, paid for unlimited, and a separate track where food programs, clinics and campus pantries pay for seats for the people they serve. That second one is the real business, but it only works if the consumer app holds people first.

  87. 1

    Appreciate the honesty here, most people only share the wins.

    1. 1

      Easy to be honest this early, there is not much to protect yet.

      The unflattering version in full: cohort 1 was 10 testers starting Sep 7, day 2 return close to zero, and we ran the whole thing without day 1 or day 7 events wired, so I am reconstructing what happened from screenshots and conversations instead of data.

      Posting the bad number here got me more useful input in three hours than two weeks of staring at it did.

  88. 1

    With the first cohort showing weak day-2 retention, what behavior in cohort 2 would distinguish an activation problem from users simply not needing the app frequently?

    1. 1

      Best question in the thread.

      The split we are using: cut day 7 return by whether the person ever hit a cookable recipe in session 1, not by cohort. If it is activation, the people who reached it come back at a much higher rate and the drop sits before that event. If it is frequency, both groups look the same and people just show up when the fridge runs low, which for most households is every four or five days, not daily.

      Second tell is what a returning user does. Open, look, leave without a scan means the app is not worth reopening. Open, scan, then bounce at the recipe list means the recipes are the problem, not the hook.

      Third one I care about: day 2 return is arguably the wrong metric for us. Nobody needs to solve dinner twice in 24 hours. If week 1 return with two or more scans looks healthy while day 2 stays flat, the honest read is that I picked a metric that does not match the behavior.

      1. 1

        That metric distinction is worth digging into. Could be worth continuing this by email — what’s easiest on your side?

        1. 1

          Email works, nitish@nourishly.app.

          Happy to send you what the funnel actually looks like once cohort 2 is instrumented, including the parts that do not flatter us. If you have run this split before on your own product I would rather hear that first, since right now I am designing the measurement off reasoning rather than experience.

          1. 1

            Thanks! I’ve just sent it over.

            Looking forward to hearing your thoughts whenever you have a chance.

  89. 1

    Interesting take. Would you still recommend this approach to someone starting today?

    1. 1

      Only in our exact situation, and I would not generalize it.

      We went for volume because 10 people is not a sample. At that size you cannot separate a broken first run from bad fit, so redesigning around it is just guessing with extra steps. If we already had 200 users and day 2 was flat, I would fix the product and add nobody.

      The part I would tell anyone starting today is the thing we got wrong: have your events logged before the testers arrive. We ran an entire cohort blind and now I am reverse engineering what happened from screenshots and conversations. Cohort 2 does not go out until day 1 and day 7 are firing.