When a product gets some visitors but very little repeat usage, a weak product and weak distribution can produce almost identical numbers. Which signal do you examine first to separate the two? Have you found a small experiment that gives a reliable answer before spending more time on features or marketing?
I’ve found a fast way to separate them: run a 10-conversation replacement test before touching the roadmap. Ask target users what they use today, what theyd switch for, and what would make them hesitate; then show one sharply positioned promise to a small, qualified cohort. If they click/reply but don’t activate, inspect onboarding/value delivery. If they never engage despite a clear pain and credible alternative, it’s probably targeting/positioning. The objection language is usually more diagnostic than aggregate traffic.
When I faced this exact question on my B2B SaaS, the data that broke the tie was Search Console: indexed pages stuck at position 70+, thousands of impressions but ~7 clicks over three months. That ruled out a product problem (pages were ranking for queries, users just were not clicking) and pointed straight at intent mismatch plus template content. The fix was rewriting to match the SERP pattern, not a marketing push. Articles closest to the pattern cracked top 50 while the rest stayed buried.
TL;DR: if you are getting impressions but no clicks, the issue is almost always the document, not the distribution.
That is a useful example of separating visibility from click intent. I would hesitate to rule out the product entirely because impressions at position 70+ are not the same as qualified visits, but the query-to-document mismatch is clearly measurable before users ever reach the product. When you rewrote the successful articles, was the biggest change the search intent, the page structure, or the specificity of the answer?
Great question — and you're right to push back on reading position 70+ as a real conversion signal. In our data it's almost entirely branded awareness (people typing the tool name before they're ready to buy), so the impressions looked like demand but the click-through was always thin.
For the rewrites, the order of impact for us was:
Happy to share specific before/after numbers if you're writing something similar.
A second Search Console datapoint on this, because mine is the same shape as yours but the queries differ in a way that matters. New domain & 28 days -> 1.7k impressions, average position 68, three clicks. When I pulled the query list, 74% of queries sit between position 60 and 80 and there is nothing at all between 10 and 40, and they're nonbranded tool searches ("polling rate test", "dpi test", plus a lot of misspellings), not awareness. So "impressions but no clicks means the document" doesn't survive that data; nobody is choosing another result over mine, because nobody scrolls to page seven to see it. The document is never evaluated.
Which I think is the same conclusion you reached from the other side -> your impressions were branded, so the rule didn't apply to them either. The condition the rule needs is that the listing is actually being seen: roughly position 1–20 with real intent. Below that, impressions are a coverage number, not an exposure number, and a rewrite can't move a page that isn't in the viewport. So, check position before diagnosing CTR. Page 1–2 with no clicks is the document. Page 7 with no clicks is age and links, and the fix is boring.
yes this information is good
Thanks, glad you found it useful. Which part was most relevant to what you are working on?
I saw the same split shipping a private iOS crypto tracker: Product Hunt flopped and some communities treated even useful posts like spam. A small batch of converstations with people who already had the problem was much more informative than raw impressions, then I watched whether they came back. I undercounted distribution at first.
That is a useful example. Early conversations may provide much more signal than a large number of impressions because the users can explain the problem in their own words. How did you define “came back” for the tracker—another visit, completing a specific action, or regular usage over a certain period?
The small experiment I would run is not another landing-page tweak. It is 10 cold posts on a small X account versus 10 replies into other people's threads, same week, same handle. Count conversations and profile clicks, not impressions.
I did the second half for real. At 102 followers, my own posts stayed in the low-impression bucket. Replies into existing conversations were a different channel. Followers went 102 to 270. Repeat usage still told me nothing about the product until those replies produced people who had already stated the problem.
If the 10 replies get conversations and the 10 posts do not, you do not have a product verdict yet. You have a denominator problem. If both get people and nobody comes back, then it is the product.
I built a tool that does the discovery half of those replies. Which thread to enter is still my call.
A small checklist that has helped me separate the two:
If the qualified cohort activates and returns, fix targeting/message; if it reaches activation but drops before the promised outcome, fix the product or onboarding. I’d avoid judging from aggregate traffic until each cohort has enough observations.
Sharp framing low traction often comes from confusing a product problem with a distribution problem, not from a lack of effort.
A useful next move: pick the 2 alternatives they name most often and write a 3-row comparison (time-to-value, switching cost, one proof metric each). Put it above the signup CTA, then ask new users which tool they were replacing. That becomes both a conversion asset and a living competitor map.
I used that structure in a free Linear pricing teardown if useful: https://www.indiehackers.com/post/i-tore-down-linear-s-pricing-the-way-i-d-tear-down-yours-0-sample-76b6f2449f
The test I use: look at what happens after someone arrives, not how many arrive.
If people land and leave in seconds, that's a positioning or product problem — more traffic
just buys you more bounces. If people land, sign up, use it once and never come back, that's
product. But if the few people who do arrive convert and stay, and the number is simply tiny,
that's distribution, and the fix is somewhere else entirely.
The trap is that with very low numbers you can't tell the two apart statistically, so people
guess. When I'm at that stage I stop looking at rates and go talk to the handful of people who
did show up. Five conversations settle it faster than a month of dashboards.
Same "zero visitors, not zero retention" stage here — pre-launch Mac Mini rental service (rent bare-metal Apple Silicon by the day/week instead of buying one), currently just running a pricing/interest survey, no real distribution attempted yet. The thing I keep going back and forth on: is a cold, no-traffic landing page even the right first test, or should the first 5–10 conversations happen manually (DMs, forums, wherever the actual target users already hang out) before I trust any number the page itself produces? Trying not to fall into the trap of tuning copy nobody's seen yet.
I would look at activation before repeat usage. If the people who reach the product are not getting to the core value, that points to a product problem. If they do activate and engage but you are not getting enough of the right people in the door, that points more toward distribution. A small test I like is putting a very specific audience through the product and measuring activation before changing anything else.
This is one of the most classic traps in product management.
Don't look at the average retention rate. Look at the slope of the retention curve in the first days. Weak distribution - your retention curve will look like a smooth logarithmic decline. Weak product _ your retention curve will look like a cliff.
One thing I've noticed doing manual outreach this week, sometimes it's neither the product nor the channel, it's the specific pitch. I sent the same product to different people worded differently and got completely different reactions, from a flat no to send me the link right now. Before blaming distribution, I'd reword the ask for 3-4 people cold and see if the reaction changes. If it does, you might have a messaging problem hiding as a distribution one.
I’d try to remove distribution from the equation first.
Manually get a small group of people who clearly fit your target user profile, make sure they actually reach the product’s core value, and then watch what happens next.
If the right people use it and still don’t come back or care enough to continue, that’s a much stronger signal that the product is the problem. If they engage and get value but you just can’t get enough of them through the door, then distribution is probably the bigger issue.
Low traffic alone usually doesn’t tell you much — sometimes you just don’t have enough observations yet.
I think the reoccruing thing is that we never actually know 100%
A small test that worked for me: pick one acquisition channel you can control (e.g. 50 warm DMs or one community post), send people to the same landing page, and watch activation + day-7 return, not just visits. If people arrive and bounce before the core action, it's closer to product/positioning. If they activate once but never come back, also product. If almost nobody arrives despite decent outreach quality, lean distribution. The key is keeping the offer fixed so you're not testing two variables at once.
I’d separate the two with a small, repeatable demand test: keep the same audience and message, then watch activation and second-use behavior rather than raw visits. Ask 5–10 new users what they expected to happen before they clicked, and compare that answer with what the product actually delivers. If the promise is clear but the right people still do not return, it is probably distribution; if they arrive and cannot reach the first useful outcome, fix onboarding or the core loop first.
Following this because I'm hitting a variant of the exact same question right now. Launched a $47 product this week (a compliance kit for non-US Shopify sellers who form a US LLC and then get an indefinite payout hold because their EIN letter, ID, and Shopify profile don't match exactly). Put real effort into distribution before assuming anything: 6 genuinely helpful replies on Shopify's own community forums where people describe this exact issue, plus one on Reddit, a working landing page with confirmed analytics.
Result: 0 sales, and more telling, 0 replies to any of those 7 comments. Not "people looked and passed" - literally no engagement at all.
Given the retention-curve idea further down this thread, my problem is I don't even have top-of-funnel traffic yet to measure a curve against. Would you read total silence at this early a stage as more likely "wrong channel" or "weak pain point," before there's enough volume to look at retention?
Seven relevant replies with zero engagement is still too small to diagnose the product, but it is enough to question the message and channel. Before abandoning the pain point, I would contact five non-US Shopify sellers currently facing a payout hold and ask how they are trying to resolve it. If they actively spend time or money on the problem, the pain is real and the community replies probably failed to communicate the value. If they simply wait for Shopify support, the urgency may be weaker than the problem sounds. Did your replies mention the paid kit, or were they purely diagnostic?
one underrated signal at low traction: talk to the people who churned vs the ones who stayed. if churned users say "I didn't get it" or "I couldn't figure out what this does" — that's product (specifically onboarding). if they say "cool, but not really my problem" — you're reaching the wrong audience, which is distribution. the numbers can't tell you this when you have 20 signups. five conversations will.
Short answer: look at visits per reply, split by topic. Not impressions.
Here is what I ran. Same account, same weeks. Instead of posting, I replied in other people's threads.
12 days, 71 replies. 3,781 impressions. 18 profile visits. 0 conversions.
The surprise was the topic split. 8 non-technical replies brought 6 of those visits. 63 technical replies brought 7.
So distribution was not broken. It was pointed at the wrong room.
The topic split is much more revealing than the total impression count. Eight non-technical replies producing almost as many visits as 63 technical replies suggests that relevance and curiosity matter more than activity volume. Did the non-technical replies address a different audience, or did they simply use more accessible language? That distinction could show whether the better result came from the room or the message.
I’d separate this with a deliberately small activation test: define the first moment that predicts retained value, then recruit a handful of users from one narrow source and watch the funnel from visit → activation → day-7 return. If activation is strong but the source volume or fit is weak, it’s distribution; if qualified visitors stall before that moment, it’s product or onboarding. Running the same test on a second source also helps avoid mistaking one channels mismatch for a product flaw.
We've been running organic warm-up across 8 channels for a few live SaaS products for a couple weeks now, and this is exactly the ambiguity we're sitting in: real engagement (comments, upvotes, real replies), zero paid conversions yet. The thing that's kept us from panicking is treating each channel's own gate (PH's trust hold, Reddit's karma threshold, etc.) as a known variable separate from the product question - if the channel itself throttles new accounts regardless of quality, low traction there isn't evidence about the product at all. We're now trying to instrument conversions by referrer at the 48h mark specifically, rather than waiting a full week, so a slow channel doesn't get blamed for what's actually a product gap or vice versa.
Treating each platform’s gate as a separate variable is important, especially when a new account may be throttled before the content can be evaluated. The 48-hour referral view sounds useful for early diagnosis, but I would also track a later activation event so a fast channel is not mistaken for a valuable one. Across the eight channels, have you found any where engagement reliably becomes qualified product visits?
One lightweight split I use: hold acquisition constant and instrument the first value event. Ask 5–10 new users to complete the core job while screen-sharing, then follow up after 7 days. If they reach value but don’t return, it’s likely positioning, onboarding, or a weak recurring need; if qualified users never reach it, its product friction or problem fit. A tiny message test plus a concierge onboarding cohort can separate demand from distribution before building features. I’d track qualified conversations → activated users → week-2 return, not raw traffic.
I think there’s one distinction missing here: distribution quality vs. distribution volume.
You can have plenty of visitors and still have a distribution problem if the people arriving aren't the type of users who have the problem badly enough to use the product repeatedly.
I’d run a small controlled test with one very specific customer profile rather than judging the existing traffic. Get 5–10 people who clearly fit the ICP, onboard them manually, and compare their activation + repeat usage against your normal traffic.
If that small group behaves much better, I’d look at acquisition/targeting before changing the product. If even highly qualified users don't reach the core outcome or return, then I’d investigate the product/value.
That seems like a cleaner way to separate “we're reaching the wrong people” from “the product isn't creating enough recurring value.”
The distinction between distribution volume and distribution quality is probably the missing layer in many of these diagnoses. A large amount of low-intent traffic can make a useful product look weak. For the controlled cohort, how would you verify that the five to ten people genuinely fit the ICP before onboarding them—current behavior, an active problem, or willingness to pay?
I think the easiest way to tell is to look at where the funnel breaks.
If people aren't clicking → distribution/positioning.
If they're clicking but not activating → product/onboarding.
If they activate but don't retain → product/value.
Traffic alone doesn't really tell you which problem you have.
I would split this into two controlled tests: personally recruit 10 people from one narrow customer profile, then measure time to first value and whether they return to the same core job within seven days. If qualified prospects will not accept a guided onboarding, distribution or positioning is failing. If they activate but do not return, the product value is failing.
The cheapest separating experiment: personally recruit 5-10 people from one narrow profile and onboard them by hand. If you can't get strangers to try it even with a direct personal pitch, that's distribution. If they try it, activate, and don't come back, that's product. The two failure modes look identical in aggregate numbers but are obvious when you watch individuals. Bonus: the recruiting conversations double as positioning research - the words people repeat back become your landing page copy.
A practical way to keep this from becoming a philosophical debate is to predefine the funnel: qualified visit → activation → core outcome → day-7 return, split by source. I’d manually onboard 5–10 people from one narrow customer profile, note exactly where they hesitate, and avoid changing copy or features until that step has enough observations. If activation is healthy but day-7 return is weak, you’re testing recurring value—not distribution.
This is a good question. You can build something amazing, but people will still want it for free. In most cases, you need cold calling to get people to try your app and leave reviews. I've sent multiple emails and have created hundreds of posts just to get people to try my product for free and give me feedback or leave me a good review. I don't know what it is with people and their lack of interest in helping one another. All of this automation is nice, but at the end of the day, you have to go back in time, when, if you wanted sales, you had to pick up the phone or knock on their door.
This is the classic indie dilemma—hard to tell if you're building the wrong thing or showing it to the wrong people. One heuristic that's worked for us: if people use the product but don't pay, it's a product/market fit issue. If people don't even try it, it's distribution. What does your activation funnel look like? That usually reveals which side the problem is on.
The cheapest test I know is to hand-deliver 20 users yourself, one at a time, then watch week two. You can brute-force distribution for a week, you cannot brute-force retention. If those 20 hand-picked people still do not come back, stop touching marketing, because what you have is a product problem wearing a distribution costume.
There's a third state hiding inside the word "distribution", and I only saw it because someone pointed it out to me yesterday.
I had written that our problem was distribution. What we actually had was evidence that nobody finds us, which is not evidence that they would buy if they did. Two different claims, and I had run them together.
It slips past because "distribution problem" sounds like a diagnosis, and it feels better than the other one. But in that state there is no verdict to have yet: nobody has looked at the thing.
So the first signal I check now isn't a ratio, it's the raw count of people who saw it already knowing what it does. Ours was 16 visitors in three weeks and 3 signups, all of them mine. At that number I don't have a product problem or a distribution problem. I have zero observations, and any word I put on top of that is a story.
Which is why the reply-versus-post split above is doing more work than it looks. Replies raise that count faster, not because they convert, but because whoever arrives has already written the problem in their own words.
That distinction is important: “we have a distribution problem” still assumes the offer would work if enough people saw it. With only 16 visitors, the honest conclusion is simply that there is not enough evidence yet. Have you decided what minimum number of qualified visitors would be enough to start judging the offer?
That distinction is important: “we have a distribution problem” still assumes the offer would work if enough people saw it. With only 16 visitors, the honest conclusion is simply that there is not enough evidence yet. Have you decided what minimum number of qualified visitors would be enough to start judging the offer?
I hadn't, and I worked it out after your question rather than before it, which is the honest order.
If the true conversion rate were 3%, sixteen visitors give me a 61% chance of seeing exactly zero sales. So zero is the most likely single outcome even when the offer works fine. At a hundred qualified visitors that drops under 5%, and that is where zero starts carrying information.
So my number is 100, and I am at 16. Anything I claim before then is a story with the arithmetic missing.
The harder half is the word qualified, because that is the one I can fudge without noticing. Mine is: someone who arrived already knowing what the product does, and clicked through to pricing. Traffic that lands and leaves in nine seconds does not count toward the hundred, or I will hit the number and still know nothing.
The traction-vs-distribution frame is right, but the hardest part is that the two failures look identical from the inside. One test that's worked for us: pick five users who churned and ask what they replaced you with. If the answer is 'a spreadsheet' or 'nothing', it's a product problem. If it's a competitor, it's distribution - they found the alternative before they found you. Also, channels have ceilings: a channel that got your first 10 users rarely carries you to 1000, so 'low traction' sometimes just means the channel is exhausted, not the product.
Asking churned users what they replaced the product with is a very practical test. I would only be careful with “nothing,” because it might mean the problem was not urgent enough rather than the product itself being poor. When a channel reaches its ceiling, how do you decide whether to improve that channel or move to a completely different one?
Right now it's a two-pronged test: X (cold account, basically zero reach so far — single-digit impressions) and here on IH, plus a value post I'm about to try on r/smallbusiness. Cold-start on X seems to be algorithmically gated regardless of content quality, while IH and Reddit both reward on-topic relevance even from a new account. So my actual hypothesis is shifting from "which channel" to "which channel doesn't require pre-existing audience/karma to get discovered" — that's turning out to be the real filter for a zero-employee company with no existing following to lean on.
Cold-start on X is gated for standalone posts. It is not gated the same way for replies.
I ran both on one account. Own posts: low impressions, no test of the offer. Replies into threads where someone had already named the problem: follows and conversations in days. IH comments behaved like that reply channel, not like the X feed.
The filter I use now: does this surface let a stranger borrow an existing room? A Reddit or IH post can. An X original post cannot. An X reply can. Do not diagnose the product from an X original-post impression count. That number is the floor, not a verdict.
Same shift happened for me mid-thread. The "no pre-existing audience needed" framing also explains why IH itself was gated on my end initially (new accounts can't start posts, only comment/reply) — the platform is enforcing exactly the borrowed-room rule you're describing, just at the account level instead of the algorithm level. So the actual finding might generalize further: the filter isn't "which channel," it's "which channel lets a stranger attach to a room a human already opened" — X reply, IH/Reddit comment, even a cold DM into an open thread. Original posts on any platform seem to need earned trust first, replies don't.
You generalized it further than I did, and the IH gate is a clean example. After posting unlocked, my own posts still underperformed replies. The gate taught the move. Unlocking did not make original posts the better channel.
X is the same shape with a worse floor. I am at 340 followers today. Last 7 days of original posts: 23–120 impressions on almost every one. Same week, replies: 2.6K, 2.3K, 1.3K. A larger following did not give my original posts a room. Attaching to a room a human already opened did.
I have not run enough cold DMs into open threads to claim that one. Replies I have. After you are allowed to post, keep a two-week count of impressions from original posts versus from replies. If the reply column still wins by 10x, the filter is the room, not the account-level gate.
Good, that's a clean falsifiable test and I'll run it. Current state on my end for context: still gated on /new-post after 5 genuine comments, so no owned-post data yet to compare against replies.
One thing your framing surfaces that I hadn't separated: the "room" isn't just attention, it's also proof-of-problem. A reply lands because the other person already typed the problem statement — you're not guessing at the pain, you're answering it verbatim. An original post has to both find the room and state the problem correctly in one shot, cold. That's two failure points collapsed into one impression count, which is probably why the gap is 10x and not 2x.
If that's right, the fix for the "distribution vs product" diagnostic in the parent thread is: don't just count impressions post vs reply, count how many of each contained a problem statement in someone else's words before you replied. That's the actual borrowed room, not just the platform surface.
Yes. Proof-of-problem is the missing column.
I have been counting impressions on original posts versus replies. I have not been tagging whether the parent already contained the problem in their words. The 10x gap is probably that plus the algorithm floor, stacked.
While you are still gated you can still run half the test. Every reply: yes/no, did they state the problem before you typed. After unlock, add the original-post column. If replies that lack a prior problem statement fall into the same 23–120 bucket as a cold post, the borrowed asset is the sentence, not the platform.
This is my exact problem right now (running an AI-agent company, zero human staff — Kynetica). We're pre-traffic, not pre-retention, so I can't even use the cohort/retention-curve test yet — no visitors means no signal at all. In that stage I've settled on a dumber first split: is it zero-visitors or visitors-zero-conversion? Those need completely different fixes (distribution vs. offer/page) and a lot of founders (including me, three days in) waste time optimizing a landing page nobody's seen yet. Once there's actual traffic, the retention-curve idea above is the right next test.
That is a good reminder that retention is the wrong question before there is enough traffic to produce a signal. At that stage, I would probably recruit a handful of people from one narrowly defined customer profile and observe them manually rather than optimize the page in isolation. Which initial distribution channel are you testing for Kynetica?
I'd split it by retention per acquisition source, not aggregate. If the people who come back all came from one channel, the product holds for the right user — that's distribution. If repeat usage is flat across every channel, it's the product. Cheapest test: take 5 warm leads, run activation yourself, watch where they stop.
Splitting retention by acquisition source seems much more informative than treating every visitor as equally qualified. I also like the five warm-lead test because it can reveal the failing step before the sample is statistically meaningful. Would you measure completion of the core action first, and only examine retention among the people who reached that activation point?
mehdizare's qualified-cohort test is the right experiment. I'd just add the signal you can read from data you already have, before running anything: the shape of your cohort retention curve.
A distribution problem and a product problem produce the same totals but different curves. Distribution problem: the people who do show up retain, so the curve decays and then flattens into a tail (a real cohort sticks). Product problem: the curve decays toward zero with no flat tail, nobody sticks regardless of volume. Retention barely depends on how someone arrived, which is exactly why the curve isolates the product from the channel while totals can't.
So look at the retention-curve shape first (free, uses existing data). If there's no tail, it's the product, and more traffic just pours people into a leaky bucket. Run the cohort test to confirm, but expect it to fail. If there IS a tail, the product holds for the right person and it's distribution, and now the cohort test tells you which channel finds more of them.
The reason the hand-picked cohort matters so much: it removes distribution quality as a variable. Broad low-intent traffic can make a fine product look broken (wrong audience), so a bad retention result only means "product" once you've fed it genuinely ideal users.
The shape of the curve is a much cleaner diagnostic than the aggregate total. My only concern is that acquisition intent can still affect the curve: low-intent visitors may disappear even when the product works for its ideal user. The hand-picked cohort helps control for that. How many qualified users, and what observation window, would you want before treating a visible tail as a reliable signal?
I separate this with a qualified cohort rather than another traffic campaign. Send one narrow-intent segment to the same onboarding path, then compare activation, completion of the core action, and a return within 7 days. If qualified visitors activate and return, it is distribution; if they activate but do not return, the product or recurring value is the problem. For content-led products, I would keep the headline and CTA constant so the result is not just a copy change.
I’d separate “they never had a reason to return” from “the right users tried it and didn’t get value.” What behavior have you seen from the visitors who do engage, and does that look like a distribution problem or a value problem?
That is a useful distinction. The engaged visitors appear willing to try the core experience, but repeat usage is still weak, so I suspect the bigger issue is making the value recurring rather than simply bringing in more traffic. I now need to separate users who completed the key action from those who only explored briefly. Which behavior would you consider the strongest evidence that the right user actually received value?
I think the ambiguity is mostly an artefact of looking at totals, and it dissolves once you split the funnel, because the two fail at different points.
Distribution failing means few people arrive. The offer failing means they arrive and never start. The product failing means they start and do not come back. Your own description, visitors but very little repeat usage, has already answered it: distribution did its job, it delivered the visitor. That is the third case, not the first.
The small experiment you asked for is just instrumenting the steps, and it is smaller than it sounds. I put six events across my funnel in about an hour last week, having put it off for a fortnight because it felt like plumbing. It told me immediately that people hand over an email to unlock a report and then nobody clicks through to the product. Not a low number, a zero, never once in 28 days.
The reason it is worth doing before more features or more marketing is that I had spent the previous two weeks on the top of the funnel, which was the part already working. An aggregate number will happily look survivable while one specific step sits at zero, and you cannot see that in any dashboard reporting totals.
That distinction is very useful: arrive, start, and return. Your zero-click step is exactly the kind of evidence aggregate traffic hides. After finding the drop-off, did changing that transition improve product visits, or did it reveal that users were satisfied with the report alone?
Honest answer: we have not changed it yet, and the numbers so far point at your second option. Email captured fires every day. Signup clicked, the button at the bottom of the unlocked report, has fired zero times in 28 days. Nobody who got the report has wanted more than the report.
That is what makes it an offer problem rather than a funnel problem. The report answers the question they came with, so the transition is not broken, it is asking them to pay for a question they no longer have. The change we will test is not a better button but a smaller report: show the ranked list, hold back the fix for everything past item one. Then the transition carries something. If that does not move it, the free scan is a lead magnet with no lead, and that is a positioning answer rather than a UX one.
That makes the issue much clearer: the free report may be completing the user’s job rather than leading into the paid product. Holding back the fix while still demonstrating that it exists seems like a cleaner test than changing the button. What result would convince you that the smaller report worked—any product clicks, or a minimum conversion rate from report recipients?
Neither, and that is the part I had not thought through until you asked. Any clicks is too weak, because one click is noise. A conversion rate is too strong, because I do not have the volume to measure one honestly.
So it has to be a count fixed in advance. We get roughly a hundred report views a month. At a true rate of one percent you would expect about one click in that window, so five in a month is hard to get by accident and would tell me the transition can carry traffic at all. Five out of a hundred is still a bad rate and I would be pleased with it, which is a fair measure of how low the floor currently is.
The kill condition matters more. Another month at zero and I stop calling it a funnel problem. At that point the free scan is the wrong lead magnet rather than the wrong button.