My cofounder and I are 17. We built an iOS app that turns one photo of your fridge or a grocery receipt into recipes you can cook with what you already have.
First cohort was 10 testers. Almost nobody opened it again the next day.
We think we know why. The first run asks you to pick a mode before you have seen anything work, and the recipe list still surfaces things that need ingredients you do not have. So the one thing the app exists to prove never actually gets proven in the first session.
The argument we keep having:
We went with 2, and we are wiring day 1 and day 7 events before the next batch of testers goes out.
For anyone who has shipped a consumer app with a bad first cohort: did fixing activation before adding users actually work, or did you need the volume first to even see where the real drop-off was?
Your reply about keeping track of the remaining ingredients raises a practical question: how will the app know what someone actually cooked or used up? A recipe opened isn't necessarily a meal made. I'd test a simple confirmation after cooking before building around that return visit. Otherwise yesterday's fridge photo could lead to another recipe with missing ingredients, even if the first run worked.
Definitely fix activation first. When we went through closed testing, day 2 drop-off was our biggest signal that the initial onboarding friction was too high.
Pouring more testers into a leaky funnel just burns through relationships without giving you useful feedback. What worked for us was:
Once your day 2 retention stabilizes for even 4-5 testers, scaling up to 20+ testers becomes ten times smoother.
It’s a real chicken-and-egg dilemma. My take is that it’s really difficult to understand what users want. The general answer would be to talk to users, as described by Y Combinator and Rob Fitzpatrick.
The other thing is having the right data. How many people check out your app? How many download it? What do they do when they open the app? Can you reduce the friction?
I help people make better decisions about warehouse design and improve their understanding of warehouse operations. For this, I’m working on www.rackflow.app , an online editor combined with a simulation mode that lets users see how a warehouse operates.
I’ve conducted 16 user interviews so far, and my takeaway was:
My internal reports showed me that most users registered but didn’t figure out how to change the model or even start the simulation. So I’m currently improving the information and documentation to attract people who are actively looking for a solution, add a preconfigured demo warehouse, and I’m also reworking parts of the user interface to improve the overall interaction and experience.
I wish you good luck! ;-)
I’d fix the first-run promise before adding a big second cohort. Ten users is too small for a clean retention curve, but it is enough to see that asking for a choice before showing the core win is risky. I’d make the first session prove “here’s a recipe you can cook now”, instrument that step, then bring in 10–20 more testers. That gives you a better funnel without waiting forever.
We run a small consumer web app, and the thing that finally made day 2 make sense for us was separating two questions: "did they get the aha in session one?" and "do they have any reason to open it tomorrow?" Those fail for different reasons, and a single day-2 number blends them together.
For the second one, it helped to design a tomorrow-shaped hook on purpose: something that is literally different each day (for you maybe "here's what you can still cook with what's left"). If the app shows the same thing tomorrow as today, even a great first session won't bring people back.
I'd fix the first run before cohort 2, but keep it tiny. 5 people you can watch live will tell you more than 30 you can only read in a dashboard.
Late to the thread but this is the exact trap I almost fell into. One data point that helped me separate the two questions: benchmark consumer D1 retention is roughly 20-25%, so 10 testers giving you ~0% on day 2 is not a sample-size problem, it's a signal. I'd fix first because volume only multiplies whatever your baseline is. The framework I'd borrow is the Sean Ellis test: of those 10, how many would be disappointed if the app vanished? If the answer is zero, 30 more testers just confirm it louder. Your "first cookable" event is the right north star — I'd add one guardrail metric though: time from open to that moment. If that shrinks and day-2 doesn't move, the problem is relevance, not activation.
I would fix the known first-session problem before bringing in a much larger group. Ten people may not provide a reliable retention percentage, but they are enough to reveal that the product’s main promise was never demonstrated.
As co-founder of SynDiary, I’m facing a related challenge. The product becomes more valuable as personal data grows, but a new user still needs one worthwhile result at the beginning. For us, information they already have such as their calendar may be the easiest starting point.
I would fix the first wall, test again with a small group of strangers, and speak personally with everyone who does not return. Silence rarely explains whether someone encountered friction or simply saw no reason to come back.
You already diagnosed the real problem in your own post the app's core promise never gets proven in session one, so no amount of volume fixes that. Option 2 without also fixing the obvious first-run issue risks just measuring the same failure at a bigger sample size.
Worth doing both, cheaply: ship the smallest possible fix to session one (even hardcoding a guaranteed-good first recipe using only common pantry items) before recruiting cohort 2, so the funnel data you're about to wire up isn't measuring a problem you already know exists.
I’ve seen a similar distinction with business software: getting someone into the product is much easier than getting them through a complete real workflow.
For example, a user can sign up, look around the dashboard and even try a feature, but that doesn’t necessarily mean they experienced the actual value. The stronger signal is whether they complete the job they came there to do and would use the workflow again.
So I’d separate “activation” into two steps: reaching the first useful result, and successfully completing the real task. Then retention tells you whether that result was valuable enough to repeat.
With only 10 testers, I’d probably use the cohort mainly to discover friction and interview the people who didn’t return, rather than treating the retention percentage as a reliable measurement.
You went with 2, and I think that was right — but I'd check the premise before you redesign the first run, because I measured the same situation this week and the data said something different from what I assumed.
We had 4 people sign up. 3 of the 4 completed the entire first run — added a site, installed the snippet, ran their first check. So activation was not the problem at all. None of them came back. What actually happened was that after the first run there was nothing scheduled for them: the free tier had no recurring work, so the product had no new data to show, and their first contact from us was a weekly digest six days later showing exactly the numbers they'd already seen. Day 2 had nothing in it.
So the thing I'd instrument before touching the UI is two separate events: "completed first run" and "returned". If your 10 testers are failing the first one, fix activation. If they're passing it and still not coming back, activation is fine and the gap is that day 2 is empty — which is a very different fix (something happens while they're gone, and a reason to look) and much cheaper than a redesign.
With 10 testers you can just ask the 7 who didn't return. n=10 isn't going to give you a statistically clean answer either way, and one sentence from a person who left beats a funnel chart at that size.
(I build an AI-search visibility tool, so the numbers above are our own — small sample, take them as one data point.)
Your framing suggests a simple gate for cohort 2: don’t add volume until most users can reach the “cookable recipe” moment in session one. I’d also separate “didn’t return” into no need vs. failure by asking one forced-choice question at exit plus an open-text follow-up. That makes the small sample useful without pretending it’s statistically precise. Are day-1 events already capturing the missing-ingredient path?
Fix activation first. If almost nobody is coming back on day 2, adding more testers will mostly give you more of the same signal.
First figure out where users are dropping off, make the core value obvious as quickly as possible, and get a few testers to reach that “aha” moment consistently. Once people are sticking around, then bring in more testers to validate whether the improvement holds at a larger scale.
I’d treat this as two experiments rather than choosing one forever: first narrow the first-run path so the promised outcome is testable (photo to recipe with no missing ingredients), then send a small cohort and interview every non-returner. Day-1/day-7 events will tell you where people stop, but the conversations will tell you whether the outcome was valuable enough to repeat. If the cookable-recipe moment rarely happens, more users mostly add noise.
I'd fix the first-run promise before scaling the cohort. For cohort 2, log one event when a user gets a recipe with zero missing ingredients, then ask a few non-returners what blocked them. That should help separate activation from low repeat need before you add more volume.
Worth flagging before cohort 2 goes out: you're now testing three different hypotheses at once (activation, frequency mismatch, wrong metric) with 20-30 people. That's a small sample to cleanly separate three explanations — if the split between "reached cookable" and "didn't" comes out roughly even, you might not have enough people in each bucket to tell signal from noise. Might be worth deciding in advance how big a gap between groups you'd actually trust as real, so you're not stuck debating whether 6-out-of-14 vs 2-out-of-11 means anything.
This is useful. How are you finding your first users so far?
At 10 testers, the useful question may not be “fix first or recruit more?” It may be: did anyone reach the one moment that proves the app works?
Your first cohort is small, but it already exposed a concrete break. The first run asks people to choose a mode before the product has shown anything working. Then the recipe list can include dishes that need ingredients they do not have. So the promise — cook with what you already have — may never be demonstrated in session one.
I would not treat the 10-person cohort as a retention rate. Ten is enough to find friction, not enough to redesign around a percentage. But it is also too early to recruit 20-30 people just to collect a cleaner record of the same break.
The smallest decisive step is to define one activation event: “photo in → at least one recipe visible that only uses ingredients the app detected → recipe opened.” Then run a small fixed flow with 5-10 new testers: skip mode selection until after that moment, and hide recipes that require missing ingredients.
Track only four things per person:
Then use these decision lines:
Your plan to wire day 1 and day 7 events is right. I would add the cookable-recipe moment before recruiting the next batch, because otherwise day 2 can still collapse several different failures into one number.
A few facts would sharpen the diagnosis: of the original 10, how many saw at least one recipe they could actually cook? Did any return later without being prompted? What did they use before the app when they had leftover ingredients? And what would make tomorrow’s open obviously useful — remaining ingredients, a meal plan, or a saved recipe?
These observations come only from what you publicly wrote. This is a scored diagnosis under uncertainty, not a verdict; validation changes confidence, not certainty, and the decision remains yours.
Straight answers to your four, with the caveat that anything from cohort 1 is reconstructed from conversations and screenshots, not logged.
How many saw a recipe they could actually cook: I cannot prove it for a single person. When I went back and asked, people described the mode picker and the photo step. Nobody described a recipe. That is the tell.
Did anyone return unprompted: no. Every second open I know about happened after I texted someone.
What they used before: nothing. They open the fridge, look at it, decide or order out. I am not replacing a system, I am replacing a glance, which is a harder thing to beat.
What would make tomorrow's open obviously useful: my guess is remaining ingredients, because after you cook one thing the fridge changed and the app is the only thing that knows how. Meal plan feels like a different product wearing our clothes.
Your decision lines are going in as written, 70 percent reach with under 10 percent unprompted return means stop blaming onboarding. That is the part I did not have.
If you want to see the first run instead of my description of it: https://testflight.apple.com/join/AZm9hsB8
The decision to fix activation first is right, but the measurement system you're building into cohort 2 is doing more work than the cohort size. You already know 10 people is too small to separate activation from fit. But "first cookable recipe" as your measurement zero is the real insight here.
The trap is treating day 2 return as the truth and then trying to interpret why it happened. If you wire events first, you'll see whether the people who actually got a working recipe came back, separate from the people who bounced at mode select. That's the measurement that lets you iterate.
One that stings in your reply: you ran cohort 1 blind. Events logged before testers arrive, not after. That's the leverage point. Small cohorts can't show you statistically significant funnels, but they can show you whether your instrumentation actually captures the moment the app works. Ten people with clean event streams beat 100 with guessing.
Agreed, and the part you named is the one that actually hurt.
Cohort 1 went out blind. I can fix every other mistake in that list. That data is gone permanently, so all I have is people telling me weeks later what they think they remember doing.
Cohort 2 does not get a link until first cookable, starts cooking and save or complete are all firing, plus time from app open on the first one. If the people who hit first cookable come back and the people who bounced at mode select do not, that is an activation problem and I keep working on the product. If both groups look the same, the honest read is that nobody needs this often enough and I have a much bigger question to answer.
Link if you want to watch it break in person: https://testflight.apple.com/join/AZm9hsB8
Thanks for sharing the numbers, that makes it much easier to follow.
Nice work shipping it. What has been the biggest challenge since launch?
Appreciate the honesty here, most people only share the wins.
Nice progress. What is the next thing you are focusing on?
Fixing activation first was the right call, but "fix the first run" can be too broad to act on. When I shipped my own iOS tool, the trap was treating onboarding as one screen instead of one moment — the single action that proves the thing works. For you that's clearly photo in, a cookable recipe out, zero missing ingredients. I'd cut mode selection entirely for cohort 1 and hard-filter to fully-cookable results, even if that means three recipes instead of thirty. Ten testers can't show you where a funnel leaks, but they can tell you whether that one magic moment lands. Did any tester actually reach a cook-it moment on day one?
Not one I can prove, and that is the whole answer.
I went back and asked. People described the mode picker and taking the photo. Nobody described a recipe, let alone cooking one. With no events wired I cannot say whether they hit a cookable result and ignored it or never got there, but the way they talk about it points at never got there.
Hard filter to fully cookable is going in exactly as you describe, three recipes instead of thirty. Mode selection is getting killed, not deferred. The bet is that a short list of things you can genuinely make beats a long list where you have to go shopping first.
https://testflight.apple.com/join/AZm9hsB8 if you want to see whether the fix actually lands.
Thanks for checking with them, and for the invite. One edge case to watch with the hard filter: zero matches. Make that explicit and offer a way to correct the detected ingredients, so it doesn't become a new dead end. Seeing a few testers reach a recipe without coaching would be a useful next check.
Curious how long it took before you saw the first real results?
Two weeks, and the first real result was a negative one.
Cohort 1 went out Sep 7. Day 2 return was visible as close to zero almost immediately, but it took until about a week ago to work out why, because I had no events running and had to reconstruct it from conversations.
Posting the bad number publicly got me more usable input in a few hours than two weeks of staring at it alone did.
It was a test. That is what testers are for. If you learned something that needs to be fixed, then fix it before you test again. Otherwise, you are just going to verify flaws that you already know about.
Fair, and this is the strongest version of the argument against what I chose.
The reason I went the other way: 10 people all from my own network is not a test, it is 10 friends being polite. I know one flaw for certain, the mode picker, and I am fixing that regardless. What I do not know is whether fixing it moves anything, and with a sample that small and that biased I cannot tell a broken first run from nobody wanting this.
So both at once. Product fixes and instrumentation land first, then 20 to 30 strangers, not friends. If day 2 is still flat after that I stop blaming onboarding and start questioning the need.
Clear and practical, thanks. Did anything surprise you along the way?
Nice work shipping it. What has been the biggest challenge since launch?
Appreciate the honesty here, most people only share the wins.
Ten people is enough to ask, not enough to measure. Message every one of them and ask what happened the second day. You'll get more out of ten honest answers than from wiring up events for cohort two. The thing that moved day two for me wasn't a fix in the product, it was a person saying hello.
I have started doing exactly that and you are right that it beats the dashboard I do not have yet.
What came back so far: nobody thought leaving was worth reporting. No complaints, no bug reports, just silence, and every one of them was friendly about it when asked. Two described the mode picker unprompted, which is how I know that screen is the wall.
The part of your comment I keep thinking about is the person saying hello. Cohort 1 got a link and nothing else. Cohort 2 gets me in the thread with them.
I am still wiring the events, because I want both. Ten honest answers tell me why, events tell me how many.
We're in the same spot with DOER right now. Enough curiosity for people to come and see what it is, sign up, then not come back. So today I wrote personally to about 40 of our early members who drifted and asked them straight: if DOER wasn't what you expected, what did you expect? And I asked them to tell me on the platform, not by email, because the point is to get them back in the door, not just get an answer.
Silence is the normal signal. Nobody writes to say a thing was fine, they just don't come back. That's why asking works, you got the mode picker named twice without prompting, which no dashboard would have handed you. And the fact they were all friendly when asked is worth noticing. They don't hate it, they just had no reason to open it again. Being in the thread with cohort two fixes that better than any event you wire. Good luck with them.
And if you want to compare notes as we both work through it, I'm around. Would be good to hear how cohort two goes.
I read your post about your first 10 testers, and I’m curious about one thing.
You mentioned that the app was surfacing recipes that required ingredients users didn’t actually have. Did your testers specifically tell you that this was why they didn’t come back, or is that something you and your cofounder identified yourselves?
I’m asking because if that’s already a known issue in the core experience, I wonder how much adding another 20–30 users can tell you before fixing it. If I opened an app that promised recipes based on what I already have, but it suggested ingredients I don’t have, I probably wouldn’t come back either :)
Or are you using the larger cohort specifically to find out whether that issue is actually what’s causing the drop-off?
It also made me wonder whether you have a very simple way for users to report a specific bad result directly from the recipe itself — something like “I don’t have this ingredient.”
I had to think about a similar feedback problem with REZYCO from a different angle. My compatibility model can only work with the answers people give it, and of course a person can always be dishonest. So I gave users a way to report someone if they discover that their questionnaire doesn’t reflect reality.
In your case, that kind of feedback might help distinguish “the user just didn’t come back” from “the app gave them a result they knew was wrong.” Do you have anything like that built in?
We identified it. No tester said it out loud, which is the uncomfortable part of your question.
What they actually described when I went back and asked was the mode picker and the photo step. Nobody got far enough to complain about recipe quality. So the missing ingredient problem is real and I can see it in the build, but I have no evidence it is what caused the drop. It might be the second wall behind the first one.
That is why cohort 2 is still going out, but not blind. If people clear the fixed first run and still leave, the recipe list is the problem and I will have the events to show it.
On feedback: no, nothing like that exists in the build, and after reading your REZYCO answer I think that is a real gap. A one tap "I do not have this" on the ingredient line does two jobs at once, it corrects the inventory the detection got wrong and it logs a bad result with the reason attached. Right now a wrong scan and a bored user look identical in my data.
Stealing it. Thank you for the actual answer instead of a take.
https://testflight.apple.com/join/AZm9hsB8 if you want to poke at it, iPhone only.
That actually makes a lot of sense. If nobody got far enough to really experience the recipe results, then I agree — you don’t know yet whether that’s the reason they left or, as you put it, the second wall behind the first one.
And I really like your point that a wrong scan and a bored user currently look identical in the data. That’s exactly why I added reporting to REZYCO — sometimes knowing that someone stopped engaging tells you almost nothing about why.
Please steal the idea 😄 I’ll be curious to see what you learn from it.
And good luck with the app! As a woman, I can definitely relate to the two eternal questions: “What should I cook?” and “What can I cook with what I actually have?” 😄
So I genuinely hope you make this work.
Clear and practical, thanks. Did anything surprise you along the way?
Two things.
Nobody complained. I expected bug reports and got silence, which is worse, because a complaint tells you where to look and a silent exit tells you nothing. When I went back and asked, people were friendly about it. None of them had thought of leaving as something worth reporting.
And the thing I was scared of was not the thing. I assumed detection accuracy would sink us, spinach coming back as kale, that kind of error. Not one person who left mentioned accuracy. They left before they got far enough for accuracy to matter, at a mode picker screen that exists for internal reasons and does nothing for them.
The one that stings: we ran the whole cohort without day 1 and day 7 events wired. Everything else here I can go back and fix. That data I cannot get back.
I’d fix the first-run promise before recruiting a larger cohort. With the current flow, more users mostly gives you a cleaner measurement of a broken experience; log the moment a user gets a recipe they can actually make, starts cooking, and returns to save or complete one, then compare those steps in cohort 2. Keep recruiting small while you iterate, and expand once that core action is happening consistently.
That is the exact event we are wiring. We call it first cookable, the moment the app shows a recipe where every ingredient came from the photo, with time from app open attached. Nobody in cohort 1 has a timestamp for it, which tells you how buried it currently is.
Starts cooking and returns to save or complete are going in as steps 2 and 3, so we get a 3 step funnel instead of one retention number that just says people left.
Keeping cohort 2 at 20 to 30 for the reason you gave. Wide enough to see the shape of the drop, small enough that I can still talk to every person who quits.
“First cookable” is a great name for that moment. Wiring the 3-step funnel—open → cookable → starts cooking → save/complete—is exactly the right instrumentation, especially when cohort 1 has zero timestamps.
With only 10 testers I'd fix activation first — the sample's too small to tell if it's the product or just the wrong users. What does the day-1 experience actually look like right now?
Day 1 as it stands: you open it, the first screen asks you to pick a mode before you have seen the app do anything, then you take a fridge photo, wait a few seconds for detection, confirm or correct the items it found, and land on a recipe list that still includes recipes needing two or three things you do not have.
So the one promise, cook with what is already in there, is technically delivered on screen four and unproven on screens one through three.
Three changes going in: kill the mode question and default to fridge only, gate the list so a recipe does not appear unless you have every ingredient, and cut first run to three screens. If day 2 is still flat after that, then it is the users or the need, and I will stop blaming the onboarding.
Makes sense. Are you planning to charge for it, or keep it free for now?
Free through the beta, no paywall in the build at all.
I am not putting one in until people come back on their own, because a paywall on top of a retention problem just hides the retention problem behind a conversion number.
Where it probably lands: free for a limited number of scans, paid for unlimited, and a separate track where food programs, clinics and campus pantries pay for seats for the people they serve. That second one is the real business, but it only works if the consumer app holds people first.
Appreciate the honesty here, most people only share the wins.
Easy to be honest this early, there is not much to protect yet.
The unflattering version in full: cohort 1 was 10 testers starting Sep 7, day 2 return close to zero, and we ran the whole thing without day 1 or day 7 events wired, so I am reconstructing what happened from screenshots and conversations instead of data.
Posting the bad number here got me more useful input in three hours than two weeks of staring at it did.
With the first cohort showing weak day-2 retention, what behavior in cohort 2 would distinguish an activation problem from users simply not needing the app frequently?
Best question in the thread.
The split we are using: cut day 7 return by whether the person ever hit a cookable recipe in session 1, not by cohort. If it is activation, the people who reached it come back at a much higher rate and the drop sits before that event. If it is frequency, both groups look the same and people just show up when the fridge runs low, which for most households is every four or five days, not daily.
Second tell is what a returning user does. Open, look, leave without a scan means the app is not worth reopening. Open, scan, then bounce at the recipe list means the recipes are the problem, not the hook.
Third one I care about: day 2 return is arguably the wrong metric for us. Nobody needs to solve dinner twice in 24 hours. If week 1 return with two or more scans looks healthy while day 2 stays flat, the honest read is that I picked a metric that does not match the behavior.
That metric distinction is worth digging into. Could be worth continuing this by email — what’s easiest on your side?
Email works, nitish@nourishly.app.
Happy to send you what the funnel actually looks like once cohort 2 is instrumented, including the parts that do not flatter us. If you have run this split before on your own product I would rather hear that first, since right now I am designing the measurement off reasoning rather than experience.
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.
Interesting take. Would you still recommend this approach to someone starting today?
Only in our exact situation, and I would not generalize it.
We went for volume because 10 people is not a sample. At that size you cannot separate a broken first run from bad fit, so redesigning around it is just guessing with extra steps. If we already had 200 users and day 2 was flat, I would fix the product and add nobody.
The part I would tell anyone starting today is the thing we got wrong: have your events logged before the testers arrive. We ran an entire cohort blind and now I am reverse engineering what happened from screenshots and conversations. Cohort 2 does not go out until day 1 and day 7 are firing.