2
2 Comments

The headline question for any pilot isn't "did it produce signups."

When I started designing the first concierge pilot for what I'm building, I had to figure out what success would actually look like before I ran one. The thing I kept landing on:

The headline question isn't whether the pilot produces trial-starts. Trial-starts are easy to game. You can buy attention, gate it behind a reward, run a thread, and signups will spike. That number tells you almost nothing about whether you've built something real.
The actual question: do the signups from the pilot behave like real users once the incentive is gone?
A pilot cohort that activates, uses the product for a week, and converts at rates close to the founder's organic baseline — that's signal. A cohort that spikes on signup, drops to zero on day 2, and disappears — that's noise in better clothes. Same number of "trial-starts," opposite meaning.

The framework I'm using:
Tag the pilot cohort with a UTM or unique signup link. Compare their behavior to the founder's organic baseline on the same activation events — first key action, day-3 return, day-7 retention. If the curves match, the gating mechanism filtered for real intent. If they diverge sharply once the perk is spent, you built an offerwall and dressed it up.
This is the single number I'd want to know before claiming a pilot worked. It's also the one most founders fudge — because measuring against your own baseline is uncomfortable when the baseline is honest.
Curious how others handle this in early pilots: what's the equivalent metric you track to tell whether the signups you got are real?

on June 3, 2026
  1. 1

    Comparing the pilot cohort to your own baseline is the right instinct. Most people skip that part because the baseline can be uncomfortable.

    The signal I’d look for is the second voluntary action with no nudge.

    The first action can be curiosity, politeness, or the incentive. But when someone comes back and does something useful again without a reminder or reward, the behavior starts to belong to them.

    So I’d care less about “did they sign up?” and more about “did anyone come back and do the real thing twice on their own?”

    Small number, hard to fake.

  2. 1

    This is the right way to think about pilots.

    The metric I’d care about is not just whether the pilot cohort matches organic users on day 3 or day 7. It’s whether they reach the same first meaningful action without needing extra hand-holding.

    Because if the pilot produces signups that only activate when someone is manually pushing them, the channel may be creating attention, not intent.

    I’d probably look at three things together:

    first key action completion
    day-3 return
    whether they repeat the core action without another prompt

    That tells you whether the pilot brought in users who understood the product, not just users who accepted the offer.

    For founder pilots, I’d also make the post-pilot report very simple: cohort source, activation rate, retention curve, and where the pilot users behaved differently from baseline.

    Happy to put a tighter version in writing if useful. The main thing I’d map is the pilot scorecard, activation events, and what founders should compare before calling a pilot successful.