
Nobody tells you, before you ship your first Android app, that finishing the app is the easy part of getting it live.
In November 2023, Google quietly changed the rules for new personal developer accounts: before you can request production access for any app, you have to run a closed test with at least 12 testers, all opted in and active, for 14 consecutive days in a row. Miss a day with too few active testers and the clock resets. This isn't a checkbox — it's enforced automatically, and there's no way around it for a new account.
I found this out the way most solo developers do: after finishing an app, at the exact moment I wanted to ship it. I had no audience, no mailing list, no Discord full of people waiting to help. Just an app, and a Google Play Console screen telling me I needed 12 real humans to open it and keep it installed for two weeks straight.
The obvious options were all bad in the same way. Asking friends and family works until you run out of friends and family, and most people ghost after day 3 anyway because there's nothing in it for them beyond doing you a favor. Paying for testers is against Google's own policy and gets flagged as exactly the kind of fake engagement the requirement was designed to filter out in the first place. Posting "please test my app" cold in forums mostly gets ignored, because everyone reading it is a developer with the identical problem, not a potential tester with free time.
That last part is what actually gave me the idea. Every single person who might see a "need testers" post is, by definition, another developer who also needs 12 testers. That's not a room full of people who owe you nothing — it's a room full of people who need exactly what you need, at exactly the same time. The obvious move was reciprocity: I test your app for real, you test mine back, and neither of us has to pay or wait on a stranger's goodwill.
So I built PeerPlay to make that swap automatic instead of something you have to manually track in a spreadsheet with however many other developers you're juggling at once. A handshake system that registers you as a tester on someone else's app the moment they agree to test yours. A verification layer that opens the target app for real and enforces a genuine minimum session, because Google's fraud detection is very good at catching apps that get opened and immediately closed, and a swap only works if both sides are actually testing, not just installing.
It's a small, unglamorous problem — most people outside of Android development have never heard of the 12-tester requirement at all. But it's a real wall that stops real apps from shipping, and I hit it myself before I built anything to fix it, which is the only reason I trust that the mechanic actually addresses the problem instead of just sounding like it does on a landing page.
If you've hit this wall on Play Store, or an equivalent gate on another platform, I'm curious what you actually did about it — swap, pay, beg, or something else entirely.
es una locura luego la gente en redes sociales te hace parecer que es facil
Totalmente. En redes solo se ve el lanzamiento, no las dos semanas buscando a doce personas que abran la app. Por eso quise escribir la parte aburrida.
The real risk is that your whole market exists because of one line in Google's policy, and the same team that wrote it can delete your business on a Tuesday. The second one is churn by design: every developer who clears 14 days has no reason to open PeerPlay again. I would work out now what the product is worth to them on day 15, because that is where the actual business lives.
Both fair, and I won't argue the first one. Google already changed the number once — it was 20 testers before it dropped to 12. The same pen can take it to zero. I don't have a hedge for that beyond keeping costs low enough that it wouldn't sink me.
On churn, I'd push back a little. The wall isn't once per developer, it's once per app, and a fair share of people come back for their next release. So it's episodic rather than dead after day 15. But you're right that episodic isn't a business on its own, and "come back in six months when you ship again" is a thin reason to exist in between.
I don't have a good day-15 answer yet. That's the honest version.
The biggest takeaway for me is that building the app can actually be easier than getting the right people to test it consistently. The reciprocity approach makes sense because both developers have the same problem and can create value for each other. I also like the distinction between solving the cold-start problem and solving tester follow-through—getting 12 people to opt in is one challenge, keeping them engaged for the full testing period is another. Turning the testing process into a structured feedback loop could make those 14 days much more valuable for both the developer and the testers.
Thanks — and agreed on the feedback loop. The hard part is keeping it light enough that it doesn't become one more reason to drift off before day 14.
I’m feeling this while recruiting a few 20-min usability sessions for an ADHD/freelancer admin tool. The ask is tiny: bring one overdue invoice, try enter → next action → chase message while I quietly observe. Free, no sales call. If you’re that kind of user and want to help, reply here.
Good luck with it. One small thing that might help: have a sample overdue invoice ready as a fallback. A lot of people will hesitate to pull up a real one with a client's name and amount on it in front of a stranger, even for 20 minutes. Offering a stand-in removes the one part of the ask that isn't actually tiny.
Good catch — thanks. Updated the ask: I’ll give a sample overdue (fake client/amount/dates) so nobody has to surface real invoices. The session stays enter → next action → chase message.
Nice, that should lower the barrier a lot. Good luck with the sessions.
What worked for me on the 12-tester gate: recruit people who already need the app, not swap partners. Reciprocal installs often never open. Ask for one screenshot of your core flow in 48 hours so you know they actually ran it.
If you can find people who already need the app, that's strictly better — no argument. They open it, they have opinions, and some of them stick around after day 14.
The catch is that the people hitting this wall are mostly solo devs with no audience yet. "Recruit people who need it" assumes you can reach them, and for a lot of first-time developers that's the exact thing they don't have. Swapping is what you do when that option isn't on the table.
The screenshot idea is good though. Cheap proof that someone got past the install screen, and it doubles as the first real look at how a stranger moves through your core flow.
Get an LLC and a DUNS number.
At least if you live in the US, that's probably easier than jumping through Google hoops.
Genuinely the cleanest exit, and I'd rather say that than pretend otherwise — organization accounts skip the 12/14 requirement entirely.
It's just not free. Formation fees, a registered agent, annual state fees, and D-U-N-S plus Google's org verification can take a few weeks on its own. Outside the US it gets messier depending on the country. For someone shipping one side project, two weeks of swapping is often the cheaper path. For someone planning several apps, the LLC pays for itself quickly.
Worth noting it only skips the first gate though. You still land in production with zero ratings and nobody finding you.
The reset rule is the part that deserves a warning label. Twelve is the floor, not the target, so if you run exactly twelve you are one uninstall away from restarting a fourteen day clock you have already paid for. I would aim for sixteen to eighteen opted in and treat the extra as insurance against normal churn. Does PeerPlay show the host how far above the floor they are on a given day, or only whether they cleared it? Knowing you are sitting at thirteen on day nine is the moment you can still do something about it.
Yes — the host sees the live count each day, not just pass/fail. Agreed that it's the most useful number in the whole run: thirteen on day nine is fixable, eleven on day thirteen isn't.
And agreed on the buffer. Twelve is the floor, and treating it as the target is how people end up restarting. Sixteen to eighteen is roughly where I'd point people too.
The distinction between getting 12 opted-in testers and getting useful feedback is important, especially with the example of only 4 of 19 staying active. I would make the exchange show a daily active streak and a short feedback prompt after each session, so the test period produces evidence rather than just an install count. What did your 117 cycles reveal about when people usually disappear: after their own app clears the gate, or earlier?
Two clusters, roughly.
The first is the first few days — people who signed up, opted in once, and never really started. That's mostly a commitment problem, and the visible quit counter seems to filter some of it at the front door.
The second, and the bigger one, is right after their own app clears. Their incentive ends on their day 14, their obligation to you runs to yours, and those almost never line up. Once they've got what they came for, the swap stops feeling like a swap.
The middle of the 14 days is actually the most stable part, which surprised me.
On the feedback prompt — I like the goal, but I'd be careful with "after each session." A daily prompt for 14 days is its own reason to drift off. One short prompt around the midpoint and one at the end probably gets you most of the evidence without adding friction to the thing you're trying to protect.
The 14-day gate isn't the only wall, though. After it, the stuff that's held up my Play submissions is the boring paperwork: a privacy policy URL, an account deletion page or flow (required if your app has accounts), and duplicate version codes when the build pipeline doesn't bump them. The target API level requirement also keeps moving up. Since the closed test is a clock you're waiting on anyway, I'd knock all of that out while it ticks so it's the only thing left.
Good list, and the parallel-work point is the right instinct — the clock is the one thing you can't compress, so everything else should be done before it stops.
Two that caught me on top of yours:
Account deletion needs a web URL, not just an in-app flow. People build the in-app button, tick the box, and get bounced because there's no publicly reachable page someone can use without installing the app.
And the Data safety form has to match what the app actually does. If you declare no data collection but you're pulling in an analytics SDK that quietly grabs a device identifier, that's a mismatch, and it's one of the easier ways to get rejected on something you didn't even know you were doing.
Target API level moving every year is the one I've made peace with. It's annoying but at least it's scheduled.
I didn't know this rule existed until reading your post, and I'm planning to ship an Android app in the next few months — so this is genuinely useful to know ahead of time instead of hitting it cold like you did.
The reciprocity insight makes a lot of sense. It reminds me of early-stage SEO / backlink building for web products, where the same principle applies: the people most likely to help you are the ones who need the identical thing from someone else at the same time, not random strangers with nothing to gain.
Question: does the swap only work if both apps are in a similar category, or does PeerPlay match anyone regardless of what the app actually does? I'd imagine testing something completely unrelated to your interests for 14 days straight is where most swaps quietly die.
No category matching — anyone can test anyone. Partly by design, partly because the pool isn't big enough to segment without making everyone wait.
Your intuition is reasonable but I don't think it's where swaps die. Google's requirement is that 12 testers stay opted in for 14 consecutive days; it isn't checking that anyone used the app daily. So the actual burden is lower than it sounds, and boredom isn't really the failure mode. What kills swaps is the other person's clock finishing before yours — they got what they came for, so they stop.
Category matching would help with something else though: whether you get useful feedback rather than just a warm body. Someone who builds in your space notices things a random developer won't. That's a real argument for it, just not the retention one.
One practical thing since you're shipping soon: the rule applies to personal developer accounts created after November 2023, and the 14 days runs as a wall clock, not as work. Start the closed test early and let it tick while you finish everything else — otherwise you're sitting there watching a calendar with a finished app.
Finding those first testers is the part nobody warns you about. I'm
solo with no audience and getting even a handful of people to actually
try something takes more effort than building the feature did. How did
you approach the strangers — cold outreach, or somewhere they already
gather?
Where they already gather, and badly at first.
I started in the Reddit threads where developers post their opt-in links asking for 12 testers. Every one of those is someone with the same problem as you, which means they're the most motivated audience you'll ever find — but they're all asking, nobody's offering. Cold outreach into that is a waste of time.
What changed things was leading with the offer instead of the ask. "I'll test yours first, here's mine if you want it." Same threads, same people, much better response. Going first costs you 14 days of tapping an app once a day, which is nothing compared to two weeks of being ignored.
Then it compounds slightly — the people you tested for are the easiest ones to ask next time.
The honest caveat: this only works because the 12/14 rule creates a crowd of people with an identical, urgent problem. If you're solo with no audience and no forcing function like that, I don't think I'd have managed it.
the accidental upside of that rule is that google is forcing you to watch the number most saas people ignore until it hurts: how many people are still there on day 14. i work on retention tooling for b2b saas and "has this account reached the thing it signed up for by day 14" turns out to be the earliest reliable churn signal there is, earlier than usage dips, earlier than support tickets. most teams only look at it months later, on the renewal.
so the swap idea is smart, but i'd keep the 14-day dashboard after you ship too. if the swapped testers drop off on day 3 you've learned something about onboarding, not just about google's policy.
Agreed on keeping the dashboard, but I'd push back slightly on what the drop-off is telling you.
In a swap, the tester isn't there because your onboarding hooked them. They're there because they need their own 14 days. So when they vanish on day 3, the most likely explanation is that their incentive moved, not that your first-run experience failed. Attributing that to onboarding would have me optimising the wrong thing.
Where it does become signal is comparative. Same pool, same motivation, same two weeks — if one app holds its testers and another bleeds them, the difference is the app. That's a cleaner read than absolute retention, and it's basically a free control group that most solo devs never get.
Which is a slightly better version of your point, I think. The 14-day number matters, but only once you can compare it against someone else's.
Honest answer to your question: I don't have a clever way through that gate. What I can add is what's waiting on the other side of it, because I measured it this week and it surprised me more than the 14-day clock did.
We have three apps live on Play. I ran 32 search queries where the words I typed literally appear in our own listing title. The app came back in 1 of them.
Searching the brand name returns it instantly. So it is indexed — Play knows the app exists and can match the text. It just won't surface it in broad search. The Console hints at why: average rating over the last 28 days is a dash. Zero ratings, across all three apps.
Device acquisitions for the main app over roughly the last month: 7 total. Four from Play "explore", two from direct links, one unattributed.
So the shape of it is: the 12-tester gate lets you ship, and then there's a second unlabelled gate that decides whether anyone can find you — and that one appears to run on ratings and engagement rather than text relevance. Nobody warns you about that one either, and in one way it's worse than the first: no checklist, no clock, no failure message. It just quietly returns nothing.
The practical consequence, which might actually be useful to you given what PeerPlay does: those 12 testers are worth more to you as 12 ratings than as 12 activity days. Same people, same two weeks, but only one of those outcomes moves the thing that gates discovery afterward. If the swap can nudge the reciprocal side toward leaving an honest rating as well as opening the app, that's a bigger favour than the gate itself requires — and it's the part I'd have wanted waiting for me on day 15.
(Where the numbers come from: I build SIGNUM HQ, a free US-market data app — options flow, dark pool share, GEX, max pain. iOS and Android: https://www.signumhq.com/app?from=indiehackers )
The second gate is real and I hadn't seen anyone lay it out with numbers before. 32 queries, 1 hit, zero ratings across three apps — that's a clearer picture of it than anything in the docs.
On the practical suggestion though, I don't think it can work, for two reasons.
Mechanically: testers in a closed test can't leave public ratings at all. Test versions don't accept public reviews, and whatever feedback they do leave stays private — it never touches your listing. So those 12 people can't become 12 ratings during the 14 days. They'd have to come back after you're in production and rate voluntarily, weeks later, with nothing tying them to you.
And that's where the second problem starts. Coordinating ratings through a swap is exactly what Play policy calls manipulating ratings and rankings. It doesn't matter that the ratings would be honest — the coordination is the thing being prohibited. I'd be running a service that quietly gets its own users delisted.
So the gap you found is real and I've got nothing for it. The only legitimate version is asking your own users after launch, uncoordinated, which is the advice everyone gives and nobody finds useful.
What I can do is make the private feedback during the test worth more, since that part is allowed and most developers ignore it — Google asks about tester engagement in the production application, and "12 people opened it" reads differently than actual written feedback.
The "4 out of 19, passed anyway" line is the whole post for me. Google's check counts installs, not attention — so a dev can clear the bar and still not know if the app works, which is a weirdly common shape: the metric that's easiest to satisfy is rarely the one that means anything.
Following the thread with aryan_sinh — did the quit counter change behavior on its own, or mostly just make ghosting visible after the fact? Visible-but-unpunished and actually-preventing-it feel like they'd need different fixes.
Mostly visible after the fact, and you're right that those are different fixes.
The counter works forward, not backward. It doesn't stop the person ghosting you today — it just helps the next person decide whether to swap with them. Which is useful, but it's reputation, not prevention.
Prevention would mean the incentive is still live on day 12, and right now it isn't: they've already got what they came for. Shame doesn't close that gap. Something structural would — holding part of their completion until yours finishes, or staggering the swap so the clocks overlap. Haven't built either.
And yes on the metric point. Google checks a streak of installs because that's what's cheap to verify. Everyone optimizes to what's measured, including me — I built the thing that makes the streak easier to get.
Rough shape: your campaign doesn't close the moment you hit day 14 — part of it stays open until the people who tested for you have finished their own runs. You get credit immediately for the days you did, but the final piece unlocks on their timeline, not yours.
That keeps a reason to stay past day 9, because leaving early costs you something that hasn't landed yet.
The obvious problem is that it punishes people whose partner ghosts them, which is the opposite of what I want. So it probably needs a cap — hold for a fixed window, then release regardless. At which point it's a nudge rather than an escrow, and I'm not sure a nudge is strong enough to matter.
That's roughly why it isn't built yet.
The cap doesn't have to be uniform time, though — it could be conditional on the partner, not the clock. If their app shows any activity (they're still testing, just behind schedule), hold longer, because the eventual payout is still likely. If they've gone fully dark, release early, because holding punishes someone for a partner who was never coming back anyway. That splits "slow" from "gone," which is closer to what you actually want to penalize. Still needs some signal of partner activity to key off, but that seems more buildable than a pure time-based nudge.
Slow vs gone is the right split — that's the distinction I was fumbling toward and didn't land.
The catch is the signal. If "activity" means the developer clicking "I tested today," it's self-reported, and a click is the one thing a ghoster will still do right up to the end. It measures intent to look active, not testing.
And I can't see much more than that. Whether someone actually opened another developer's app lives in that developer's Play Console, not in mine. So the honest version is keying off things I can observe inside PeerPlay — feedback left, messages answered, their own campaign still being tended — which is weaker than real app activity, but at least can't be faked with one tap.
Still more buildable than a flat time cap, agreed.
Good instinct moving off self-report, but the three you listed aren't equally hard to fake either — a low-effort message ("looks good!") costs almost nothing to send and would still count as "messages answered." Feedback that references something specific about the actual app is a lot more expensive to fake than feedback that exists. Might be worth weighting by effort rather than just presence — not just did they leave feedback, but does the feedback contain anything that could only come from someone who actually opened the app.
Agreed — "looks good!" is barely better than a click.
The question is who judges whether feedback is specific. I can't reliably tell from the outside whether a comment could only come from someone who opened the app, but the developer receiving it can, instantly. So the simplest version is letting the host mark feedback as useful or not, and weighting from that.
Someone earlier in the thread suggested asking for a screenshot of the core flow within 48 hours. That's the other half: cheap proof of opening, then the host's judgment for whether they actually engaged.
4 out of 19 actually being there, and he still passed, is the bit that stuck with me. Google is checking a streak of installs, you are chasing real sessions, and those two jobs fight. Friends ghosting by day 3 is the same thing at a smaller scale. The quit counter being on the record instead of a punishment sounds better than the usual guilt email. Has anyone actually stuck past day 9 because of that counter, or do they still vanish once the swap looks fair?
Honest answer: I can't prove attribution. No clean before-and-after, so anything I claim about the counter working is guesswork.
But "once the swap looks fair" is the whole thing. People don't vanish at random — they vanish when their own campaign completes. Their incentive ends on their day 14, yours runs to your day 14, and those clocks never line up.
My guess is the counter filters who joins rather than who stays. Weaker claim, but the one I can defend.
I'm doing a waitlist to see the interest in the idea of what I'm building. Should I just open it for beta and have people test it?
What is idea?
Veris is my idea. A social media platform with the goal of eliminating bots, sharing genuine content, and using AI to check shared information. Check out the link on my profile
No future...
The gap between "I finished building" and "I actually have users" is where most solo founders quit, and you leaned into it hard. Hunting down 12 strangers manually to get real feedback instead of launching cold is exactly the kind of thing that separates builders who make it from those who don't. Respect the grind.
Thanks. Though I'll be honest — the 12 strangers weren't a choice. Google won't grant production access without them, so it was hunt them down or don't ship at all.
The part I did choose was building a place for it instead of posting in Reddit threads for two weeks. That bit was probably procrastination dressed up as strategy.
While I have not had this exact experience on the Android store for a mobile app, I have had somewhat of a similar experience. Once you finish building your app or website, finding the first group of users can often be the most difficult. The distribution for these new web apps or mobile apps can be fragmented (there is no one consolidated place to reach the most beta testers/users). Something I am doing about this is offering onboarding or interviews to validate the product I have built. I would also be open to any recommendations about where to best find places to find other apps to test (in IndieHackers platform or elsewhere).
The fragmentation is exactly why PeerPlay exists — there was no consolidated place, so everyone was cobbling together Reddit threads and guilting friends into installing things.
So to answer directly: that's what the free tier is. You browse apps currently in closed testing, test the ones you pick, and yours goes into the same pool. No payment, no account-selling. https://peerplay.vmcreate.rs
Your onboarding-and-interviews instinct is the right one though. Google's rule checks a headcount, not whether anyone used the app — one developer told me he had 4 consistent testers out of 19. He passed, but he didn't get what he actually needed.
The reciprocity mechanic is the real product thesis. Have you seen swaps where both sides consistently complete the required testing period, or is that still the weak point?
It's still the weak point, and I don't think it stops being one.
What holds up is the first move. People who test back immediately get testers immediately — one developer put it plainly in his completion note: got testers fast because he tested their apps first. The exchange works when it's genuinely mutual.
Where it breaks is the middle of the 14 days. Signing up is cheap, day 9 isn't. One developer told me he had 4 consistent testers out of 19 — he hit Google's headcount but not much else. That gap between "enrolled" and "actually there on day 12" is the whole problem.
So I've spent more engineering time on anti-ghosting than on anything else in the app. There's a visible quit counter now, deliberately not punitive, just on the record. It helps. It doesn't solve it.
My read after 117 cycles: reciprocity reliably fixes the cold-start problem, but it doesn't fix follow-through. Those are two different problems and I conflated them when I started.
The 117-cycle distinction between solving cold-start and solving follow-through is the important signal. If you’re open to it, what’s the best email to reach you on?