4
7 Comments

Three people found the same gap in three unrelated products this week, and none of us have fixed it

I didn't expect a comment I left on someone else's Product Hunt post to turn into this, but here we are.

I asked wendel (SaaS production-readiness checklist) whether his "auditability" item covers the case where a Stripe webhook fires, your own downstream action against it times out, and you genuinely don't know if the action landed — not malformed, not a duplicate, just unresolved. He came back today: it doesn't. He'd treat it as a separate check, not something idempotency already solves.

That's the same gap SuperMcG (building something called OpsWatch) and I have been circling on a different post entirely — StareBrain dispatches an action, and if the result comes back ambiguous, I currently don't have a clean answer. Don't assume success. Don't blindly retry, since retry can itself be the dangerous move if the first attempt landed. Don't leave it silently pending forever either. SuperMcG's actually building a system specifically for this — preserving that state as its own thing rather than forcing it into success/failure/retry.

Three people, three completely different products — a webhook checklist, an ops-monitoring tool, and a phone-automation agent — independently hit the identical shape of problem this week without any of us looking for it in each other's stuff.

I don't think that's a coincidence. I think "did the action actually happen" is a harder question than "did the request get accepted," and most systems don't have a real answer for the gap between those two, they just don't hit it often enough to notice.

Still don't have this solved for StareBrain. But it's reassuring, in a strange way, to know I'm not chasing something imaginary.

on September 25, 2026
  1. 1

    The same shape shows up in crawling. We build an SEO scanner at UtilitySEO and when a page returns a 403, the honest answer is "we don't know." It might mean forbidden, or it might mean a CDN edge challenged us because we weren't a browser. Someone on another IH thread this week described exactly this: seventeen pages went undecided because the edge refused non-browser requests, but every page was fine. "Don't assume success, don't blindly retry" is exactly the constraint.

    The three-product convergence is the interesting part. When three unrelated systems independently discover the same gap, it usually means the gap is structural rather than product-specific. The reason most systems get away with not solving it is that ambiguous outcomes are rare enough to handle manually — until they're not. Stripe webhooks, phone automation, ops monitoring all live in the frequency range where "rare" starts meaning "several times a day at scale."

    SuperMcG's approach of preserving the ambiguous state as its own thing rather than forcing it into success or failure seems right. The mistake is always collapsing it prematurely.

    1. 1

      The 403-vs-CDN-challenge case is a clean addition to the list because it's the same ambiguity at an even earlier layer — you're not even sure the request was evaluated as a real page-fetch, let alone what the page contains. That's arguably the hardest version of the three, because unlike a webhook or a dispatched action, you don't control the other side at all; you can't add an idempotency key or a status channel to someone else's CDN.

      Genuinely curious how UtilitySEO resolves the 17 undecided pages once you spot them — is there a second pass that tries harder to look like a browser and treats that result as authoritative, or does it stay flagged as "undecided" and get surfaced to whoever's reading the scan report? That's basically the same "what's the resolution path" question the original post left open, just for a case where you can't ask the remote system anything else.

      Your "rare becomes frequent at scale" point deserves to be pulled out on its own, actually — it might be the actual reason nobody's built a standard pattern for this yet. Most systems are designed by people who hit the ambiguous case once a month and just eyeball it. The pattern only becomes a forcing function once volume makes manual review impossible, and by then it's already load-bearing infrastructure with no clean abstraction underneath it.

      1. 1

        there is one for sms, partially. the aggregators (twilio, plivo, vonage) hand you delivery receipts through status callbacks, so you can ask whether the handset confirmed, independent of your own crash. two catches: some carriers mark delivered at the gateway, not the phone, and a few just lie. a dlr is a strong signal, not a proof.

        voice has nothing equivalent. status callbacks tell you answered, busy, failed, and completed only means someone picked up. there is no did-the-right-person-hear-it signal anywhere in the stack. so the design is basically forced: transport truth from the provider callbacks, meaning truth from a human. the line you draw between those two is the honest part.

        1. 1

          The transport truth vs meaning truth split is the cleanest way I've heard it put, and it's a better line than the one I drew in the post.

          One difference for us: StareBrain sends from the handset itself, not through an aggregator, so there are no Twilio-style callbacks. From what I know, Android gives the sending app its own sent and delivered reports, though delivery depends on whether the carrier bothers to send one, and the call log can show whether a call connected and how long it lasted. If that holds up, some of what I called unknowable might actually be readable on the device. I haven't checked it yet.

          Your "a DLR is a strong signal, not a proof" point makes me think "delivered, unconfirmed" should be its own state rather than getting rounded up to success. Does that match how you'd handle a gateway-level delivered?

  2. 1

    yeah, "unknown" has to be a real state here. for the Stripe case i'd save the request id before sending, then ask the downstream system what actually exists for that id before anyone retries. if it can't answer, show the operator the ambiguous action and let a person reconcile it. how are you thinking about that last step for StareBrain when the phone call itself may have happened but the result didn't come back?

    1. 1

      Honestly, the "ask the downstream system what actually exists" step is exactly where StareBrain's version gets harder than the Stripe case, and I don't have a clean answer yet. For a webhook, there's usually a query-able source of truth on the other side — you can ask Stripe's API what the charge status actually is, independent of whether your own handler crashed. For a phone call, the "downstream system" is a telecom carrier or the recipient's phone, and there often isn't an equivalent query I can make after the fact — call logs and carrier-side delivery confirmation exist, but they're not always reachable or reliable in real time depending on the channel (SMS delivery receipts exist in some pipelines, not others; a phone call has essentially none).

      So right now the honest state is closer to your fallback case by default, not your primary one: if I can't independently confirm what happened, it goes to a human as a flagged, ambiguous action — "we attempted this call/text, we don't know if it landed, please check." That's a real state, not a placeholder, but it's also not a resolution — it just correctly refuses to guess. The actual open problem is whether there's a cheaper source of truth I'm not using (carrier-level delivery confirmation APIs, if they exist for the channels I care about) that would let more of these resolve automatically instead of dumping everything on a person.

      If you know of a carrier-side or platform-side API that actually answers "did this SMS/call complete" independent of whatever crashed on my end, that's the missing piece — I haven't gone looking hard enough yet.