8
24 Comments

Google made me find 12 strangers willing to test my app for 14 days. Nobody warned me how hard that actually is.

Nobody tells you, before you ship your first Android app, that finishing the app is the easy part of getting it live.

In November 2023, Google quietly changed the rules for new personal developer accounts: before you can request production access for any app, you have to run a closed test with at least 12 testers, all opted in and active, for 14 consecutive days in a row. Miss a day with too few active testers and the clock resets. This isn't a checkbox — it's enforced automatically, and there's no way around it for a new account.

I found this out the way most solo developers do: after finishing an app, at the exact moment I wanted to ship it. I had no audience, no mailing list, no Discord full of people waiting to help. Just an app, and a Google Play Console screen telling me I needed 12 real humans to open it and keep it installed for two weeks straight.

The obvious options were all bad in the same way. Asking friends and family works until you run out of friends and family, and most people ghost after day 3 anyway because there's nothing in it for them beyond doing you a favor. Paying for testers is against Google's own policy and gets flagged as exactly the kind of fake engagement the requirement was designed to filter out in the first place. Posting "please test my app" cold in forums mostly gets ignored, because everyone reading it is a developer with the identical problem, not a potential tester with free time.

That last part is what actually gave me the idea. Every single person who might see a "need testers" post is, by definition, another developer who also needs 12 testers. That's not a room full of people who owe you nothing — it's a room full of people who need exactly what you need, at exactly the same time. The obvious move was reciprocity: I test your app for real, you test mine back, and neither of us has to pay or wait on a stranger's goodwill.

So I built PeerPlay to make that swap automatic instead of something you have to manually track in a spreadsheet with however many other developers you're juggling at once. A handshake system that registers you as a tester on someone else's app the moment they agree to test yours. A verification layer that opens the target app for real and enforces a genuine minimum session, because Google's fraud detection is very good at catching apps that get opened and immediately closed, and a swap only works if both sides are actually testing, not just installing.

It's a small, unglamorous problem — most people outside of Android development have never heard of the 12-tester requirement at all. But it's a real wall that stops real apps from shipping, and I hit it myself before I built anything to fix it, which is the only reason I trust that the mechanic actually addresses the problem instead of just sounding like it does on a landing page.

If you've hit this wall on Play Store, or an equivalent gate on another platform, I'm curious what you actually did about it — swap, pay, beg, or something else entirely.

on September 19, 2026
  1. 1

    I didn't know this rule existed until reading your post, and I'm planning to ship an Android app in the next few months — so this is genuinely useful to know ahead of time instead of hitting it cold like you did.

    The reciprocity insight makes a lot of sense. It reminds me of early-stage SEO / backlink building for web products, where the same principle applies: the people most likely to help you are the ones who need the identical thing from someone else at the same time, not random strangers with nothing to gain.

    Question: does the swap only work if both apps are in a similar category, or does PeerPlay match anyone regardless of what the app actually does? I'd imagine testing something completely unrelated to your interests for 14 days straight is where most swaps quietly die.

    1. 1

      No category matching — anyone can test anyone. Partly by design, partly because the pool isn't big enough to segment without making everyone wait.

      Your intuition is reasonable but I don't think it's where swaps die. Google's requirement is that 12 testers stay opted in for 14 consecutive days; it isn't checking that anyone used the app daily. So the actual burden is lower than it sounds, and boredom isn't really the failure mode. What kills swaps is the other person's clock finishing before yours — they got what they came for, so they stop.

      Category matching would help with something else though: whether you get useful feedback rather than just a warm body. Someone who builds in your space notices things a random developer won't. That's a real argument for it, just not the retention one.

      One practical thing since you're shipping soon: the rule applies to personal developer accounts created after November 2023, and the 14 days runs as a wall clock, not as work. Start the closed test early and let it tick while you finish everything else — otherwise you're sitting there watching a calendar with a finished app.

  2. 1

    Finding those first testers is the part nobody warns you about. I'm
    solo with no audience and getting even a handful of people to actually
    try something takes more effort than building the feature did. How did
    you approach the strangers — cold outreach, or somewhere they already
    gather?

    1. 1

      Where they already gather, and badly at first.

      I started in the Reddit threads where developers post their opt-in links asking for 12 testers. Every one of those is someone with the same problem as you, which means they're the most motivated audience you'll ever find — but they're all asking, nobody's offering. Cold outreach into that is a waste of time.

      What changed things was leading with the offer instead of the ask. "I'll test yours first, here's mine if you want it." Same threads, same people, much better response. Going first costs you 14 days of tapping an app once a day, which is nothing compared to two weeks of being ignored.

      Then it compounds slightly — the people you tested for are the easiest ones to ask next time.

      The honest caveat: this only works because the 12/14 rule creates a crowd of people with an identical, urgent problem. If you're solo with no audience and no forcing function like that, I don't think I'd have managed it.

  3. 1

    the accidental upside of that rule is that google is forcing you to watch the number most saas people ignore until it hurts: how many people are still there on day 14. i work on retention tooling for b2b saas and "has this account reached the thing it signed up for by day 14" turns out to be the earliest reliable churn signal there is, earlier than usage dips, earlier than support tickets. most teams only look at it months later, on the renewal.

    so the swap idea is smart, but i'd keep the 14-day dashboard after you ship too. if the swapped testers drop off on day 3 you've learned something about onboarding, not just about google's policy.

    1. 1

      Agreed on keeping the dashboard, but I'd push back slightly on what the drop-off is telling you.

      In a swap, the tester isn't there because your onboarding hooked them. They're there because they need their own 14 days. So when they vanish on day 3, the most likely explanation is that their incentive moved, not that your first-run experience failed. Attributing that to onboarding would have me optimising the wrong thing.

      Where it does become signal is comparative. Same pool, same motivation, same two weeks — if one app holds its testers and another bleeds them, the difference is the app. That's a cleaner read than absolute retention, and it's basically a free control group that most solo devs never get.

      Which is a slightly better version of your point, I think. The 14-day number matters, but only once you can compare it against someone else's.

  4. 1

    Honest answer to your question: I don't have a clever way through that gate. What I can add is what's waiting on the other side of it, because I measured it this week and it surprised me more than the 14-day clock did.

    We have three apps live on Play. I ran 32 search queries where the words I typed literally appear in our own listing title. The app came back in 1 of them.

    Searching the brand name returns it instantly. So it is indexed — Play knows the app exists and can match the text. It just won't surface it in broad search. The Console hints at why: average rating over the last 28 days is a dash. Zero ratings, across all three apps.

    Device acquisitions for the main app over roughly the last month: 7 total. Four from Play "explore", two from direct links, one unattributed.

    So the shape of it is: the 12-tester gate lets you ship, and then there's a second unlabelled gate that decides whether anyone can find you — and that one appears to run on ratings and engagement rather than text relevance. Nobody warns you about that one either, and in one way it's worse than the first: no checklist, no clock, no failure message. It just quietly returns nothing.

    The practical consequence, which might actually be useful to you given what PeerPlay does: those 12 testers are worth more to you as 12 ratings than as 12 activity days. Same people, same two weeks, but only one of those outcomes moves the thing that gates discovery afterward. If the swap can nudge the reciprocal side toward leaving an honest rating as well as opening the app, that's a bigger favour than the gate itself requires — and it's the part I'd have wanted waiting for me on day 15.

    (Where the numbers come from: I build SIGNUM HQ, a free US-market data app — options flow, dark pool share, GEX, max pain. iOS and Android: https://www.signumhq.com/app?from=indiehackers )

    1. 1

      The second gate is real and I hadn't seen anyone lay it out with numbers before. 32 queries, 1 hit, zero ratings across three apps — that's a clearer picture of it than anything in the docs.

      On the practical suggestion though, I don't think it can work, for two reasons.

      Mechanically: testers in a closed test can't leave public ratings at all. Test versions don't accept public reviews, and whatever feedback they do leave stays private — it never touches your listing. So those 12 people can't become 12 ratings during the 14 days. They'd have to come back after you're in production and rate voluntarily, weeks later, with nothing tying them to you.

      And that's where the second problem starts. Coordinating ratings through a swap is exactly what Play policy calls manipulating ratings and rankings. It doesn't matter that the ratings would be honest — the coordination is the thing being prohibited. I'd be running a service that quietly gets its own users delisted.

      So the gap you found is real and I've got nothing for it. The only legitimate version is asking your own users after launch, uncoordinated, which is the advice everyone gives and nobody finds useful.

      What I can do is make the private feedback during the test worth more, since that part is allowed and most developers ignore it — Google asks about tester engagement in the production application, and "12 people opened it" reads differently than actual written feedback.

  5. 1

    The "4 out of 19, passed anyway" line is the whole post for me. Google's check counts installs, not attention — so a dev can clear the bar and still not know if the app works, which is a weirdly common shape: the metric that's easiest to satisfy is rarely the one that means anything.
    Following the thread with aryan_sinh — did the quit counter change behavior on its own, or mostly just make ghosting visible after the fact? Visible-but-unpunished and actually-preventing-it feel like they'd need different fixes.

    1. 1

      Mostly visible after the fact, and you're right that those are different fixes.

      The counter works forward, not backward. It doesn't stop the person ghosting you today — it just helps the next person decide whether to swap with them. Which is useful, but it's reputation, not prevention.

      Prevention would mean the incentive is still live on day 12, and right now it isn't: they've already got what they came for. Shame doesn't close that gap. Something structural would — holding part of their completion until yours finishes, or staggering the swap so the clocks overlap. Haven't built either.

      And yes on the metric point. Google checks a streak of installs because that's what's cheap to verify. Everyone optimizes to what's measured, including me — I built the thing that makes the streak easier to get.

      1. 1

        Rough shape: your campaign doesn't close the moment you hit day 14 — part of it stays open until the people who tested for you have finished their own runs. You get credit immediately for the days you did, but the final piece unlocks on their timeline, not yours.

        That keeps a reason to stay past day 9, because leaving early costs you something that hasn't landed yet.

        The obvious problem is that it punishes people whose partner ghosts them, which is the opposite of what I want. So it probably needs a cap — hold for a fixed window, then release regardless. At which point it's a nudge rather than an escrow, and I'm not sure a nudge is strong enough to matter.

        That's roughly why it isn't built yet.

  6. 1

    4 out of 19 actually being there, and he still passed, is the bit that stuck with me. Google is checking a streak of installs, you are chasing real sessions, and those two jobs fight. Friends ghosting by day 3 is the same thing at a smaller scale. The quit counter being on the record instead of a punishment sounds better than the usual guilt email. Has anyone actually stuck past day 9 because of that counter, or do they still vanish once the swap looks fair?

    1. 1

      Honest answer: I can't prove attribution. No clean before-and-after, so anything I claim about the counter working is guesswork.

      But "once the swap looks fair" is the whole thing. People don't vanish at random — they vanish when their own campaign completes. Their incentive ends on their day 14, yours runs to your day 14, and those clocks never line up.

      My guess is the counter filters who joins rather than who stays. Weaker claim, but the one I can defend.

  7. 1

    I'm doing a waitlist to see the interest in the idea of what I'm building. Should I just open it for beta and have people test it?

      1. 1

        Veris is my idea. A social media platform with the goal of eliminating bots, sharing genuine content, and using AI to check shared information. Check out the link on my profile

  8. 1

    The gap between "I finished building" and "I actually have users" is where most solo founders quit, and you leaned into it hard. Hunting down 12 strangers manually to get real feedback instead of launching cold is exactly the kind of thing that separates builders who make it from those who don't. Respect the grind.

    1. 1

      Thanks. Though I'll be honest — the 12 strangers weren't a choice. Google won't grant production access without them, so it was hunt them down or don't ship at all.

      The part I did choose was building a place for it instead of posting in Reddit threads for two weeks. That bit was probably procrastination dressed up as strategy.

  9. 1

    While I have not had this exact experience on the Android store for a mobile app, I have had somewhat of a similar experience. Once you finish building your app or website, finding the first group of users can often be the most difficult. The distribution for these new web apps or mobile apps can be fragmented (there is no one consolidated place to reach the most beta testers/users). Something I am doing about this is offering onboarding or interviews to validate the product I have built. I would also be open to any recommendations about where to best find places to find other apps to test (in IndieHackers platform or elsewhere).

    1. 2

      The fragmentation is exactly why PeerPlay exists — there was no consolidated place, so everyone was cobbling together Reddit threads and guilting friends into installing things.

      So to answer directly: that's what the free tier is. You browse apps currently in closed testing, test the ones you pick, and yours goes into the same pool. No payment, no account-selling. https://peerplay.vmcreate.rs

      Your onboarding-and-interviews instinct is the right one though. Google's rule checks a headcount, not whether anyone used the app — one developer told me he had 4 consistent testers out of 19. He passed, but he didn't get what he actually needed.

  10. 1

    The reciprocity mechanic is the real product thesis. Have you seen swaps where both sides consistently complete the required testing period, or is that still the weak point?

    1. 1

      It's still the weak point, and I don't think it stops being one.

      What holds up is the first move. People who test back immediately get testers immediately — one developer put it plainly in his completion note: got testers fast because he tested their apps first. The exchange works when it's genuinely mutual.

      Where it breaks is the middle of the 14 days. Signing up is cheap, day 9 isn't. One developer told me he had 4 consistent testers out of 19 — he hit Google's headcount but not much else. That gap between "enrolled" and "actually there on day 12" is the whole problem.

      So I've spent more engineering time on anti-ghosting than on anything else in the app. There's a visible quit counter now, deliberately not punitive, just on the record. It helps. It doesn't solve it.

      My read after 117 cycles: reciprocity reliably fixes the cold-start problem, but it doesn't fix follow-through. Those are two different problems and I conflated them when I started.

      1. 1

        The 117-cycle distinction between solving cold-start and solving follow-through is the important signal. If you’re open to it, what’s the best email to reach you on?