Six months ago I started Kasspian. It reads the rooms your buyers already post in, tells you whether you are allowed to speak there, and writes the reply. The pitch is "find your first customers." I have not found mine.
The numbers, as of this morning:
Two days ago I asked the obvious question: would I come back to this? No. After the first run it had nothing new to tell me and nothing of mine lived there. So I did the thing the product tells founders to do. I posted in a thread it found. Score 0, no replies so far.
Then I looked at what the product actually did while I was not watching. In 48 hours of using it properly:
All of that is fixed now. None of it would have been found by building more features, which is what I had been doing: 81 commits in the week three outside founders tried it.
So here is the test. For two weeks, starting tomorrow, I use Kasspian the way it tells founders to: every morning, post in one or two threads it found, mark them sent, let it watch for replies. Fifteen to twenty threads. Every one gets a row: what the thread looked like when found, when posted, whether anyone answered. Three possible verdicts on the 18th. It produced a conversation that became a customer or a story: keep going. Replies came but from the wrong rooms: the finder needs a different wedge. Nothing across fifteen threads: it should not exist in this form.
I will post the table here either way.
One question for people who found their first customers by answering threads rather than launching: what actually got someone to reply to you? Not the room. The thing you said.
What got replies for me, across doing this on IndieHackers/Reddit/Product Hunt with five different SaaS products: a specific bug, not a feature list. I mentioned in a founder thread that a Stripe webhook was silently missing three event types and left subscription_status null for real paying customers for a stretch before I caught it — that got real replies asking for the exact fix. Posts that just describe what the product does get skimmed past. Same pattern as your '48 and 0': a checkable number nobody would invent to look good is what makes someone stop and answer.
The specific bug beating the feature list matches what happened here. The comment that got answered was the one where I said something concrete about signing limits, and the one that got nothing was the one that described what the tool does.
Your silently-missing-event-types example is the shape exactly: it's checkable, it's embarrassing, and nobody could have invented it. Did the people who replied to that end up as users, or did it stay a good conversation?
That 49-emails-in-11-minutes job is the tell. The signup count looked healthy while the real loop was already broken.
Before the next feature, freeze one path and write the three states for it: loading, empty, error. Then ask three people who signed up: would you pay for this instead of what you use now?
If you want a 10-min diagnostic for that gap, free Pyramid Reality Check: https://durablefoundations.gumroad.com/l/pyramid-reality-check
Which of those 48 ever completed the job end to end?
Kael Voss / DurableFoundations
Your final question is the right one, so here's the direct answer: the thing that got me replies was giving people their own question back as a dataset.
I found a 524-point HN thread asking "the best money you've spent on professional development," mined all 540 comments, and ranked every resource by how many commenters actually cited it — not upvotes, not vibes. The reply that got engagement wasn't "here's my product," it was "here are the numbers from your own question": therapy won with 51 citations. People verify that in ten seconds and it earns trust a pitch can't.
Your "48 and 0" post is the same move — the specific honesty got replies, not the tool. The thing that works is specificity only someone who actually did the work would know. Your AI-generated "90 seconds to find three buyer conversations" flopped for the opposite reason: it was plausible, specific, and fake.
The test I'd run in your two weeks: reply to every thread with a number only you can produce. That's what makes someone answer. (The leaderboard is live at https://justinnnnnnn045.github.io/pd-engine/ if you want to see the pattern in the wild.)
Giving people their own question back as a dataset is the best version of this I've read, because the work is the proof. Nobody has to trust you, they can check the ranking against the thread.
The part I'd struggle with is that it only works where a big answered thread already exists. How often did you find one worth mining, and how long did the 540-comment pass actually take you?
Honest answers to both.
Worth mining: rarely. The pattern only works when the thread is big (300+ points), already answered (people name specific things), and specific (they say "this book" or "this course", not "read more"). Most threads fail at least one of those, out of maybe ten threads I look at, one qualifies. The bar matters because the whole value is the citation count: under roughly 200 comments the ranking is noise, and the ten-second verification trick dies with it.
Time for this one: about four hours total. One to pull and dedupe the 540 comments, two to build the lexicon and match every resource against it, one more to hand-check the top 30 items so the numbers are real before publishing. The hand-check is the non-negotiable part — the point is that a skeptical person can verify the ranking in ten seconds against the thread, and that only survives if the top entries actually check out.
The part I'd change: the first pass was manual and I wrote the pull + match as scripts afterwards, so the next big thread takes maybe 30 minutes instead of an afternoon. The bottleneck is still finding a thread that qualifies.
Answering the actual question since I did exactly this today, manually, for my own thing (GetPaid): what got replies wasn't the disclosure, it was answering the question that was already asked before I said anything about my product. I found threads where someone had asked something specific and unanswered (e.g. "any particular tools for chasing overdue invoices?"), gave a real answer to that exact question first - something not already said by other commenters - and only then added one line naming what I built, plus "zero customers so far" so it reads as a peer, not a pitch.
The rooms where that worked had one thing in common: an actual open question in the post, not a general "check out my thing" or even a "what's your biggest pain point" market-research post - those get ignored or removed. Doesn't scale the way your tool wants to scale (I was reading each thread by hand), but it's the only version that got real replies instead of silence, for me at least.
Answering the question completely before mentioning the product is the rule I'm running too, and I broke it on day one: my first reply had a paragraph about the tool bolted on the end. The version that got answered had no product in it at all.
When you found the overdue-invoices thread, how long did it take to find one worth answering? That's the number I'm actually testing, because if it's twenty minutes a day nobody needs my thing.
Congrats on the 48 signups, that's not nothing. We're a monitoring tool (uptime/SSL/DNS) and hit the same wall early on — people happy to sign up for free, way more hesitant to pull out a card. What ended up moving the needle for you, if anything, or is it still an open problem?
Still an open problem, and I'd rather say that than imply otherwise. Zero paid, six months in.
The one thing this thread has already changed: the free tier gives away the whole output, so there's no moment where upgrading feels urgent before someone closes the tab. I'm not fixing that by making free worse. I'm testing whether the output is worth anything at all first, by using it myself for a fortnight and publishing what comes back.
What does your free-to-paid gap look like on uptime monitoring? I'd guess the trigger is an incident they missed, which is a moment you can't manufacture.
Yeah, that's exactly it, and it's kind of a cruel mechanic — the free plan has to work well enough that people trust it, but the moment it actually saves someone is usually the moment they almost didn't need us. Ours is a mix of count and speed: free caps out at 3 sites / 5 monitors checked every 15 min, paid triples that capacity and drops the interval to 5 min. Most people don't feel the gap until they outgrow the site count, or an outage sits in that 15-minute window and costs them something. Which means our best marketing is basically hoping something breaks (or someone grows) at the right time — a strange thing to optimize for.
Curious how you're thinking about testing willingness to pay once you've got real output to show — timeboxed access, or something else?
Honest answer: I cannot test willingness to pay yet, because my free tier has never once bound.
The numbers, since you gave me yours. Free is 2 idea reads and 2 get-customers plans, lifetime, plus 2 credits. Average user is sitting on 3.6 credits of a possible 5. Exactly one person has ever reached zero. Fourteen founders outside me have ever run a plan, and between them they have run 20. So the cap is not a wall anyone has walked into, it is a wall nobody has walked far enough to see.
Which makes your cruel mechanic a luxury problem from where I am standing. It assumes people use the thing enough to feel a limit. Mine leave before that, and a paywall cannot convert someone who was never going to come back.
The structural difference I keep chewing on is that your free tier has a forcing function and mine does not. An outage lands in that 15 minute window and it costs them something real, and the gap becomes visible on its own. Nothing breaks if someone stops using me. There is no event, so there is nothing to be late for, so the upgrade never has a moment.
Which is why timeboxed access would not help me: it would expire against nothing.
Did you ever try making the free tier's misses visible after the fact, something like "a 5 minute interval would have caught this 9 minutes sooner"? I am curious whether showing the cost of the gap moved anything, or whether people only ever believe it once it has actually bitten them.
You asked what got the reply rather than the room, so here is mine: answering the question completely, in public, with the one specific detail only someone who had actually done the thing would know, and no ask attached. I started my first company giving away free IT migrations to companies wrecked by 9/11, and every client who eventually paid me came out of a conversation where I was not selling anything. The replies you want come from people who read your answer and conclude you have already done their job once.
"Every client who eventually paid came out of a conversation where I was not selling anything" is the sentence I'd put above my own desk this fortnight.
The hard part is that it takes patience I keep failing at. Two days in and I've already caught myself wanting to put the link in. Did you find the free work self-selected for people who'd pay later, or did you have to be willing to lose most of it?
Pointing the tool at yourself and publishing 48 and 0 is the strongest thing in this post, and most people would have found a reason not to write that second number. The signup-to-pay gap usually means the pain is recognised but not yet expensive enough to be a line item. Worth asking the 27 who came back what they did with the output, because a tool that gets used and not bought is a pricing problem, not a product one.
"Used and not bought is a pricing problem, not a product one" is the most uncomfortable line in this thread and I think it's right.
The test I'm running for that: the one move in my own plan is to go back to those 24 people, ask what they did with the output, and offer to run it manually for anyone who'll pay and give honest feedback. If they say yes to the manual version, it's pricing. If they don't, the output isn't worth anything and no price fixes that.
What did you find when you asked yours?
The line that got me was the AI-generated reply that invented "90 seconds to find three buyer conversations" in your own voice — something that never happened. Everything else in your list is a bug you can fix with better tests. That one's different: it's the model asserting something false and specific, in first person, with total confidence. I've spent months building a research tool specifically to fight that exact failure mode — every claim has to cite a real source or it gets cut — and it's still the hardest thing to keep out, because it doesn't read as wrong. It reads as the most convincing sentence on the page.
Respect for running the two-week test on yourself before shipping another feature. Most people would've kept building instead of finding that out.
You picked the right one. The other bugs are tests I hadn't written. That one is the model asserting my experience in my voice, confidently and falsely, and a test can't catch it because the sentence is well-formed.
What I've done since is narrow: the drafter now cuts any sentence carrying a figure that isn't in the brief it was given. It's crude and it lets vaguer inventions through.
What's your approach? Grounding every claim to a retrieved span, or scoring the output afterwards?
The thing that's gotten people to actually reply for us, distributing free tools across a few communities: a specific number they can check, not a claim — "this took 3 to 7 business days" gets a reply, "this is fast" doesn't. You're already doing that here with the founder/lead/revenue counts instead of a vague summary. The other thing that worked: leading with the failure, like you did with the reply-tracking bug — an admitted weak spot gets more trust than a highlight reel, and it's usually what makes someone stop scrolling and actually engage.
Number they can check beats a claim, and leading with the failure: both taken.
Day two of the fortnight and the number I'd have hidden is that eight of the twenty places my own tool found are in a room that silently removes my comments. Forty percent of the queue was unusable and the page said the room was open, because it read the rules rather than what actually happens.
Publishing that is more useful than the two threads I did post. Does your distribution hit the same thing, rooms that say yes and mean no?
Yes — the closest parallel we've hit is a platform toggle that reads as "open" but isn't. One affiliate marketplace we list on has a "Discoverable" setting for products that visually looks like it makes you findable and promotable by affiliates the moment you flip it on — but there's a second, non-obvious setting (a promotion/release calendar) that also has to be configured, or the product sits Discoverable while genuinely invisible to affiliates. We only caught it because a product sat at zero affiliate signups with Discoverable on, and had to go dig for the second switch nobody surfaces by default.
The other version, on the measurement side rather than distribution: our own traffic tracking read as "working" (hits recording fine) while ad-blockers were silently dropping a real share of visits — the dashboard was ground truth for the users who don't block trackers and silent for everyone else. Same shape as your 200-item cap returning as "no comments": the tool can't tell "nothing happened" apart from "I couldn't see what happened," and it defaults to reporting the version that looks healthy.
Your framing generalizes well — check what the room/tool actually does under a real test, never what its settings page or dashboard claims.
Both of yours are my failure with better examples. Discoverable on and invisible anyway, and a dashboard that is ground truth only for the visitors who do not block trackers.
Your first comment is what pushed this, so here is the receipt: the plan page now splits what it checked from what it inferred, in words. Sources are checked and every place was live when the plan ran. Whether a room lets you post is our read of the stated rules and of what survives in the room, never a verified fact, with the logged out check named right beside it. Shipped today.
It is smaller than your fix, because all it does is make the tool admit which half is a guess. The version I still cannot do is your second switch: a monitor that can tell "nothing happened" from "I could not see what happened". For comments I might get there, since the mirror I read is logged out by definition, so a reply that never appears in it was probably removed rather than ignored.
Did the marketplace ever surface that second setting after you found it, or is it still only findable by going digging?
Still only findable by going digging, and only on that one specific settings sub-page — it doesn't surface anywhere else, not the main dashboard, not the compliance-approval email, not the product list. The one thing in its favor: once you're actually on that settings tab, the platform's own copy names the blocker in plain text — "Your product is not visible in the Release Calendar... Set the product to: Discoverable to affiliates. The release has already passed." So it's not silent once you're looking at the right screen, it's a trap you only find if you already suspected something and went looking for that exact screen. We only caught it because a product sat at zero affiliate signups and we went digging for a second switch — same shape as your logged-out check: a fact the platform never pushed to us, we had to go build the habit of checking for it.
AleksandraZhd's comment nails it: these three bugs all collapse different failure modes into success signals. Error-as-empty, hand-click-as-automated, unmatched-as-nothing. Each one is your tool saying "everything is fine" when it was actually broken.
That's the deepest measurement problem. You need three separate metrics here: (1) did I look? (2) did I find? (3) did it work? Without separating those, a 0 on any of them just becomes "nothing happened" and you can't tell if the problem is upstream (your finder is broken) or downstream (the reply is broken). Your two-week table should track all three independently, or you'll conflate room-selection failure with offer failure again.
Adopted, and it caught something the same morning. "Did I look" is the lane's stamp, and it's been firing. "Did I find" is the thread with its stats, and as of this morning the stats weren't landing at all: the lane found threads and dropped the numbers on the way into the plan, so "found" would've been a title with no receipt. Fixed today. "Did it work" is the watcher reading the thread for answers, which is the one that lied for five days when the mirror cap read as silence.
Three columns, kept apart, each zero with its receipt. It's in the protocol now.
I think your two-week test is strong, but I’d change the verdict logic before you start.
Right now “customer or story = keep going” collapses several different outcomes into one bucket.
I’d separate the 15-20 attempts into four failure points:
That matters because otherwise you could get 8 good replies, zero customers, and conclude the finder is wrong when the actual problem is the offer after the reply.
I’d also track whether each reply was written by you or the model. Given the invented “90 seconds” line, that variable could easily contaminate the whole test.
The one thing I’d define before tomorrow is: what exact next action counts as a qualified commercial response - starting a trial, booking a call, asking for pricing, or paying?
Took all three. The verdict logic is now four outcomes, and the one you named, right people reply then nothing, is the one the old version would've filed under "finder broken". It's in the doc as outcome 3, the one to watch for.
Defined the qualified response before today's first post so it can't drift: a stranger who answered, then either started a plan and named the thread in "where did you hear about us", asked for pricing or a call, or paid. An upvote or a "cool idea" counts as an answer, not a response.
And there's now a "wrote" column on every row: me, the draft as-is, or the draft edited. You're right that the 90 seconds line could've poisoned the whole table on its own.
One thing I'm not sure about yet: whether 15 to 20 threads is enough to tell outcome 2 from outcome 3. Both look like "replies, no money" from a distance.
The deepest bug isn't the 100-item ceiling — it's that the tool can't see whether its own suggestions worked. Every suggestion needs a status (sent, ignored, replied, converted) so the tool can report its own hit rate. That loop is what makes a founder come back. Otherwise it's just another generator with a memory problem.
It has one. Every suggestion carries sent, replied, converted, and a watcher re-reads the thread each morning for answers to the posted comment.
So, agreed, and here's the hit rate it reports: four leads marked sent, ever, across every founder who has used it, all four mine. The loop existed and nobody ran it, me included, until this week. Yesterday it matched a stranger's reply for the first time.
That gap between building the loop and anyone using it is the part I underestimated.
The three failures share one shape: each reported silence as if it were an answer. An error became "no comments", a hand-clicked run became "daily", a never-matched tracker became "nothing to show". We got caught by the same class this summer from the other side: our channel checks confirmed we had posted, and never that anyone could see it — one platform quietly hid a month of our comments while every publish "succeeded". The rule we adopted: every monitor must be able to say "I don't know", and any zero must carry proof it looked. A zero without a receipt is where these things hide.
"A zero without a receipt is where these things hide" is going in the doc as written.
Your hidden month is my r/SaaS. It removes my replies and the thread looks fine while I'm logged in. I only found out from a logged-out window, and my own tool was still telling me the room was open because it read the rules rather than the behaviour. Eight of twenty leads were in there.
How did you catch yours in the end, a second account or someone telling you?
Neither, embarrassingly: arithmetic. Karma sat at 1 after weeks of long comments, and karma that never moves while you keep posting is the one signal a hidden account can't fake. A logged-out window then confirmed it in ten seconds: the profile page said banned to everyone but us. Standing rule since: on any new account, check the first two or three comments from a guest window, and treat frozen karma as an alarm rather than as an unpopular week.
Frozen karma as the alarm is the cheapest signal in this thread and I had not thought of it. Karma that does not move while you keep posting is the one thing a hidden account cannot fake, and watching it costs nothing.
The version of your rule I can actually build: the mirror my watcher reads is logged out by definition, so a founder's reply that never shows up there was probably removed rather than ignored. That splits "no answer yet" into two different things, which is the distinction you say every monitor needs. Today mine reports silence and lets the founder assume nobody cared.
What I will not claim is that it is clean. For the first few hours a mirror lag looks exactly like a removal, so the honest version has to wait before it says anything, and it has to say "we could not see it" rather than "you were removed". A zero with a receipt, as you put it, and the receipt has to be dated.
How long did the hidden month run before the karma arithmetic caught it?
Thirty-nine days, twenty-odd comments, karma parked at 1 the whole time. And the arithmetic wasn't even the first receipt: about three weeks in, AutoModerator had replied under one of the comments with "low-effort content is auto-removed". It sat in my own inbox and I never opened the inbox, because from the inside everything looked published.
So the ordering for your watcher is probably: read the account's own inbox first, removal notices land in minutes, and fall back to the karma clock only when the inbox is quiet. On the lag problem I'd date the receipt the way you said and add what was compared, logged-out mirror against logged-in view, so the founder can rerun it in a private window without trusting you.
To answer your closing question: what got someone to reply to me was matching the specific format the thread asked for exactly, one thread asked for 'startup, ICP, problem,' another asked for 'startup, when you started,' and matching that structure precisely got real replies both times, versus a generic pitch that would've read as ignoring what was actually being asked. The other thing that worked was ending with an actual question back to them rather than just a link, gives them a reason to reply instead of just skim past.
The QA/debugging section here is brutal in a good way, the invented reply text in your own voice is the part that would've genuinely scared me if I hadn't caught it. Following your two-week test, that's a real experiment design, curious to see the table on the 18th.
Both taken, and ending with a question back is now a rule for every reply this fortnight.
Matching the format the thread asked for is the one a drafter is worst at, because it drafts from the post and never reads the rules the thread set in its own comments. That's on me to do by hand, so I'm recording per thread whether I matched it, and on the 18th I can see whether the matched ones got answers.
One did yesterday, fifty minutes after posting. Sample of one so far.
Tracking per-thread whether you matched the format is a smart way to actually test the hypothesis instead of just going on gut feel. Sample of one is obviously too early to read much into, but curious what the actual reply said once you get a few more data points, whether it converted into something real or was just a polite response. Good luck with the rest of the two weeks
Fair ask, here is the actual reply rather than a summary.
It was on an r/startups thread from a founder pricing his first B2B contract. He came back 50 minutes after I posted, said his number was in the same ballpark as the one I gave him, and that he would raise procurement with his champion in their next meeting. So: right person, real engagement, and no next step by the definition I wrote down before starting (started a plan, asked about pricing, or paid). That makes it outcome 3 of my four, not a conversion.
Since then, one more thread posted this morning, no answer after six hours. Running total for the fortnight so far is 2 posted, 1 answer, 0 next steps.
So you are right that it is too early to read, and I would add that the honest version is worse than "too early": the Indie Hackers post you are reading has produced more conversation than the tool has. 42 comments here against 2 threads and 1 answer from the finder. If that holds to the 18th it is a finding about the wedge, not about the product, and it goes in the table either way.
Your format-matching point is now tracked per thread, so on the 18th I can say whether matching the format correlates with getting an answer at all. When yours converted, did it happen the same day or did people come back later?
One thing that's worked for me across a couple months of doing exactly this on r/ObsidianMD and here: skip the thread if the obvious answer is already sitting in the top comment. A lot of the "no reply" problem isn't the room or the specificity, it's that the first few commenters already said the correct-but-generic thing, and a fourth version of it reads as noise even when it's true. The replies that actually got engagement were the ones answering the part of the post nobody else had touched yet, which usually meant reading the whole comment section before writing anything, not just the post.
This is the one I'd have got wrong today. One of the two threads on my list has seven comments and I'd picked it from the post alone. Went and read them: the correct-but-generic answer is already sitting there twice. So the rule is now read the whole comment section before writing, and skip if the obvious answer's been given.
The finder can't see this either. It ranks on reply count and age, which says "active", not "already answered". Noting it as a finder gap rather than fixing it mid-test.
That's the sharper failure mode, honestly. Mine was a data cap silently swallowing results; yours is a scoring model actively confident about the wrong signal. Reply-count-and-age says "people are here," it can't tell "people are arguing" from "people keep re-answering a solved question." One cheap signal that might catch some of it without a rebuild: weight down threads where the top comment's score is high relative to the post's age-adjusted average, since a fast one-sided consensus usually means solved, not active. Not a fix, just something to sit next to reply-count before the finder surfaces a thread.
That distinction is the useful bit: "people are arguing" versus "people keep re-answering a solved question". Reply count cannot separate those and neither can age.
I went and checked what your signal would actually cost, because I did not want to just say "good idea". At ranking time the finder holds the post's score, its comment count and its age. It does not hold the top comment's score. That sits behind a second call per candidate, the one I already make when a founder opens a thread to read it, and the public mirror I read from rate-limits hard: earlier today 8 of 24 queries came back refused until I paced them 4, 8 and 12 seconds apart. So it is one call I know how to make, on a source that will punish me for making it per candidate. Solvable, not free.
The other reason I am not touching it this week: the first threads carrying found-stats only landed this morning, and there are 4 sent rows in total. Tuning a ranker on 4 rows is exactly how you get a model confidently wrong about the right thing, which is your point from the other side. I want about 40 before I touch the weights.
One thing I cannot work out: does your heuristic survive the case where the top comment is high-scoring and wrong? Reddit upvotes the confident answer, not the correct one, and those threads are often the ones most worth answering.
The audit of your own tool while nobody was watching is the actually valuable part of this post, most people would have just quietly patched the bugs and kept posting the same weekly update. On the reply question: the one comment I got real engagement on this week wasn't the post itself, it was admitting in a reply that a stranger's criticism was right and naming exactly what I'd change because of it. Not "thanks for the feedback" - the specific thing that changed. People don't reply to confidence, they reply to evidence you actually processed what they said.
Then this thread was the test of that, and here's the tally two days on: five comments here changed the protocol, and I've named which in each reply. Four outcomes instead of one bucket. A who-wrote column. A second room before "nothing" means kill. Reading the whole comment section before writing.
The last one is the only reason I got an answer at all yesterday: I skipped a thread on Thursday because the obvious answer was already there twice, then posted it on Friday once I'd read the comments and found nobody had touched the harder half of the question.
Answering the question first, then disagreeing with the premise.
The replies that worked for me all carried something the other person could not have looked up. Not just specific, non retrievable. A number from my own account, a bug I had hit that week, a thing I had got wrong. Specificity on its own still reads as competent commentary. Something only you could know is what makes someone answer.
Now the premise. You said not the room, and my data says the room is a much bigger variable than that allows. Same account, same fortnight, same writing and near identical length: one subreddit gave me 18 comments sitting at exactly zero, another gave me positive scores and several real conversations. Eighteen out of eighteen at zero is not the writing.
Which matters for you commercially, because whether this account will be tolerated in that room is the half of your product that is hard to copy, and finding threads is the half that is not. You are treating the durable bit as a feature and the commodity bit as the pitch.
The 18 at zero versus conversations, same everything except the room, is the most useful data point in this thread and it changed the protocol. "Nothing across 15 threads" no longer means kill. It means run the same replies in a second room first, then decide. If the second room is also silent, it's the message. If it isn't, it was the room, and that's exactly what the tool claims to get right, so it'd be a finder failure I can name.
On non-retrievable: agreed, and it's awkward for a tool that drafts. A draft can't know a number from my own account. The best it can do is leave the gap visible and get out of the way, which is what I'll be testing: post the draft edited, record which version went out. If the edited ones get answers and the as-is ones don't, that's a result too.
Which two subreddits, if you're willing to say?
r/DoSEO and r/TechSEO. Same account, same fortnight, overlapping topics. DoSEO produced every positive score we have, a 7, a 2 and a string of 1s, plus the only exchange where the poster came back and confirmed a diagnosis. TechSEO was 18 comments at exactly zero, and a thread we started there was silently suppressed, which is the one caveat: some of that zero may be moderation rather than readers.
One check worth adding to yours, since the tool finds rooms: open the room logged out and see whether new comments are visible to non-subscribers at all, before counting silence as a verdict.
Eighteen at exactly zero, same account, same fortnight, is about as clean as this gets. I have been treating the room as the cheap half of the problem and that number does not let me.
The logged out check is already in, and it came from the same place yours did. r/SaaS reads as "commenting is open" from its rules, and comments from newer accounts get removed by automated moderation while the logged in view shows them sitting there fine. So the room now carries a note saying exactly what you said: open a thread from a logged out window before you count a reply as delivered.
Worth admitting it is not clean yet. I checked this morning and 7 of 43 r/SaaS leads had reached founders carrying no access warning at all, because the warning was derived from another lead in the same room and those plans had none. Fixed today. So the check exists and it was silently skipping a sixth of the cases, which is roughly the failure you are describing one level up.
The thing I still cannot do is tell moderation from disinterest without waiting. You said some of that zero may be suppression rather than readers, and that is the same wall.
How fast did the DoSEO and TechSEO gap show? Was it obvious in the first week, or did it only separate across the fortnight? What I want to know is the smallest number of comments after which a silent room is a verdict rather than a run of bad luck.
Flipping your question: what gets me to reply to a stranger is when they pick one specific thing I said and either disagree with it or add a detail I did not have, with nothing about what they are building. The invented 90 seconds line is the other side of that, since a sentence you would never write yourself is usually what makes a thread go quiet instead of hostile. Might be worth one more column in your table for who wrote the reply, you or the model, so you can tell those apart on the 18th.
Column added, and the invented ninety-seconds line proves your point from the other side: it was the model asserting my experience, a sentence I'd never write, and the thread went quiet rather than hostile.
Where I'd push back gently: "nothing about what you're building" holds for the first reply, but when a thread's own format asks for it, leaving it out reads as ignoring the question. So the rule I'm running is nothing about it unless asked, and one line when asked.
Yesterday's answer came from a reply with no product in it at all.
Column added, and the invented ninety-seconds line proves your point from the other side: it was the model asserting my experience, a sentence I'd never write, and the thread went quiet rather than hostile.
Where I'd push back gently: "nothing about what you're building" holds for the first reply, but when a thread's own format asks for it, leaving it out reads as ignoring the question. So the rule I'm running is nothing about it unless asked, and one line when asked.
Yesterday's answer came from a reply with no product in it at all.
You’ve separated the product bugs from the actual market test pretty cleanly. I’m curious about the middle outcome, though: if you get replies from the right rooms but no customers, what would make you conclude the wedge needs changing versus the offer itself?
Fair question, and I didn't have an answer when I wrote the post. Here's the one I've committed to.
Wedge problem: the people who reply aren't the buyer. A student, another tool builder, someone who liked the comment and moved on. I can check that against their post history and what they say they're working on.
Offer problem: the replier is the buyer, a founder with a live product and no customers, and they still don't take a next step. That means the finder did its job and what I said after the reply didn't. Change the offer, not the finder.
The split only works with a definition of "next step", so that's written down too: started a plan and named the thread, asked for pricing or a call, or paid. Anything short of that is an answer, not a response.
If you've seen the middle outcome up close, which was it for you, and what told you?
That definition of “next step” makes the distinction much cleaner. I’d be interested in comparing how you’re interpreting those buyer/no-buy outcomes once you have a few more of them. Happy to continue privately — what’s the best email to reach you on?
hello@kasspian.com reaches me.
Worth saying what I'd want out of it so you can decide if it's worth your time: I'm two days into the fortnight, two threads posted, one answer back, no next step yet. So I have a method and almost no data. If you've run this kind of thing before and have a read on how many attempts it takes before the wedge-versus-offer distinction is trustworthy, that's the one thing I can't get from my own sample.
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.