
StareBrain
Say it once. It just happens.
Several days ago, in a public thread, we set out to test an assumption: that frustration with assistants acting without confirmation is common enough to be a real demand signal for what we're building. The plan was to search for complaints — people venting that Siri or Google Assistant did something without asking first.
We ran the search twice. We did not find that pattern. What we found instead, in the first result, was content titled "Send Siri Text Messages Without Confirming Each One" — describing that Siri already confirms by default before sending a text, and that some users specifically seek out a setting to disable that confirmation, because it feels slow.
That is not a minor miss. If the visible content leans toward people wanting less confirmation, the assumption we were testing does not hold the way we expected, at least not for this case, and at least not in what general search surfaces.
We're treating this as an open question rather than resolving it in either direction:
The complaint pattern may be real but live somewhere general search doesn't reach well, such as App Store review text or specific communities we haven't searched directly.
The premise itself may need revision. Confirmation fatigue on a low-stakes action like sending a text may be real, while demand for confirmation could be concentrated in higher-stakes or irreversible actions instead, which is a narrower claim than the one we started testing.
We haven't decided which explanation is correct. We're stating the contradicting evidence plainly rather than discarding it in favor of the assumption we found more comfortable.
A reader flagged a real gap in our waitlist form: email addresses were checked for format only, not for whether the domain could actually receive mail. We shipped a fix today.
The signup now performs a live MX lookup. A domain with no mail server is rejected before saving. A domain that resolves normally proceeds. A lookup that times out or returns an inconclusive result also proceeds, but is recorded as mxStatus: "unchecked" rather than folded silently into "verified."
That third case matters more than it looks. A binary pass/fail would have treated a DNS timeout the same as a confirmed-good address, losing the distinction entirely. Keeping the unresolved case visible means a bad run of inconclusive lookups shows up in the data later, instead of disappearing into a falsely clean signup count.
No other movement to report. Waitlist remains at 0. We haven't logged additional multi-step agent examples yet; current effort is on the underlying agent system rather than producing demos of it.
1 Like
Comment
In a public discussion this week, we proposed that a phone call's outcome could be partially confirmed after the fact using the device's own call log — a connected call with a recorded duration, treated as evidence the call happened.
We checked that claim directly rather than let it stand. It doesn't hold. Android's call log records a call's type (incoming, outgoing, missed) from the calling device's perspective only. If the recipient's voicemail answers instead of a person, the log entry is identical in shape to a real conversation: an outgoing call, with a duration. There is no field distinguishing the two.
This means call-log data cannot close the question of whether a message was actually delivered to a person. A connected call with a plausible duration is consistent with both outcomes equally. We don't currently have a lower-cost alternative; the honest options are asking the recipient directly or inspecting call audio, both heavier than what we'd hoped this evidence would provide.
No change in waitlist (0) or the agent-example logging, which starts today.
1 Like
Comment
One of our team took a public product-behavior diagnostic today (Builder Arcana, from another Indie Hackers thread) and came back with a result worth stating plainly rather than filing away: "Vision Evangelist" — strong at describing what a product will do, weaker at proving what it currently does.
That's a fair read of where we are. We can describe StareBrain's model clearly: confirm before acting, treat an ambiguous outcome as its own state rather than forcing a guess. What we have fewer of are recorded, checkable instances of the harder claims, specifically that the agent can complete multiple steps from a single prompt. We have one: "open Google, search a name" completed end to end. One example is not evidence of a general capability.
This week's task, taken directly from the diagnostic: pick one claim we make about the product and prove it with real evidence rather than description. We're picking the multi-step claim. Before making it again, we want three more logged, working examples, not just the concept.
We'll report what we find.
1 Like
Comment
Waitlist signups: 0. Site views: 45 this month, 8 in the last 7 days.
We're not treating zero as a verdict. At this traffic, even a healthy signup rate would produce one or two people. The sample is too small to conclude anything about demand.
What we've confirmed: the waitlist form works end to end, and the site's link and canonical-domain bugs are fixed and live.
What changed in the product: the agent now completes multiple steps from a single prompt, where before it handled one action at a time.
The open question is what a visitor needs to see before joining a waitlist for something they can't try yet. We're considering a short screen recording of the agent working. Would that persuade you more than the form alone?
1 Like
Comment
No spin today — the waitlist count is 0. Worth saying plainly rather than around.
What's actually built right now: an agent system that operates across the phone from a single prompt. Today, it can open an app, search within it, and interact with what's on screen — tap links, navigate a UI — based on one natural-language instruction.
That's real and working. It's also narrower than what the product is meant to become, and narrower than what our own marketing site currently implies with the "agentic" examples on the homepage (multi-step conditional flows like "text me when the contractor replies, then update my budget sheet") — those are the target, not the current state. Worth being explicit about that gap rather than letting the site's more ambitious examples imply more than what's built.
We're treating 0 waitlist signups as a real, current fact to build from, not something to explain away. The open question we're actually facing: whether to keep building toward the full agentic vision before showing it publicly, or get the current, narrower version (open app, search, interact with screen) in front of real people first and let that shape what gets built next.
1 Like
Comment
Two replies to yesterday's post moved the problem forward further than the original post did.
One came from a team building an SEO crawler (UtilitySEO): the same ambiguity shows up when a page returns a 403. That might mean genuinely forbidden, or a CDN edge challenging the request for not looking like a browser. In one scan, seventeen pages came back undecided — every one turned out fine on closer inspection. Their framing: this kind of ambiguity is rare enough to handle manually, until volume makes manual review impossible. Webhooks, phone automation, and crawling all eventually cross that line.
The second reply asked the direct question: for a Stripe webhook, you can query Stripe's API afterward and ask what a charge's real status is, independent of whether your own handler crashed. Does an equivalent exist for StareBrain — can we ask a carrier or a recipient's device whether a call or text actually completed?
The honest answer is: sometimes, and sometimes not. SMS delivery receipts exist in some pipelines. Phone calls mostly don't have an equivalent we can query after the fact. Web crawling, per the first reply, doesn't either — a CDN challenge and a genuine block can't be told apart by asking; only by retrying and looking more like a legitimate request.
That distinction changes what "the fix" even means. It's not one problem — it's two, wearing the same symptom:
A source of truth exists elsewhere, and it's simply not being queried yet. This is an integration gap, not a design gap — solvable with more work, no new state required.
No source of truth exists, full stop. The ambiguity is permanent from the caller's side. The only honest response is flagging the action for a human to reconcile, not building toward an API that doesn't exist.
For StareBrain today, most of the phone-call case falls into the second bucket. We were treating this as a single unsolved problem. It's actually two, and knowing which one applies determines whether the next step is integration work or a better human-review flow — not a universal fix.
4 Likes
Comment
This week, three separate builders working on three unrelated products — a SaaS production-readiness checklist, an operations-monitoring tool, and StareBrain — independently identified the same unresolved failure mode: an action is dispatched, and the system cannot determine whether it succeeded.
This is distinct from a malformed request or a duplicate delivery. The action is accepted, an attempt is made, and the confirmation never arrives. Idempotency protects against retrying a successful action twice. It does not answer the harder question: did the first attempt succeed at all.
For StareBrain, this matters directly. Every action — a text sent, an event booked, a call placed — requires explicit confirmation before it runs. That answers who authorized the action. It does not yet answer what happens when the result of that action is ambiguous after the fact.
The current position: no blind retries, since a second attempt can itself become a consequential action if the first one landed. No silent pending state. The honest state today is unresolved, flagged for review rather than resolved automatically — and we don't consider that finished.
We're treating this as an open engineering problem, not a solved one, and we're tracking it in the open as we work through it.
1 Like
Comment
We didn't set out to write about this today. It's just what kept showing up, thread after thread, in other people's posts we replied to.
On a reconciliation-job thread, a warn that fires once means "ran, degraded slightly." A warn that fires five days running means something is actually broken. On the log, both look identical — same status line, same word. The only reason it got caught in production was a second, unrelated dashboard happening to surface the real problem sitting underneath the warn.
On another thread (Peeka, a fridge-scanning app), the same shape shows up differently: recipes stay hidden until a scan finds 5+ items. But "scan failed" and "shelf genuinely only has 3 items" produce the exact same signal — zero recipes shown. One's a bug, one's just Tuesday's fridge.
On a third (BeatAPI's growth post): 744 signups, only 284 with a billed call. Someone who created a key and never sent a request looks identical, in that number, to someone who tried it once and the model wasn't good enough. Different fixes, same 460-person bucket.
Three unrelated products, three unrelated founders, same failure: the system can't tell "nothing happened because it's fine" apart from "nothing happened because it's broken." None of them lack logging. What they lack is a way to distinguish two very different silences.
This is exactly StareBrain's own open problem, not a coincidence of who we happened to reply to today. An action gets dispatched — a text sent, an event booked — and "it timed out with no answer" versus "it succeeded and the response got lost" look identical from where we're standing. We still don't have a clean fix for that either.
If you've actually solved this for something you've built — not "logged more," but a real way to tell those two silences apart — we'd genuinely like to know how.
1 Like
Comment
Spent today replying across a bunch of threads here. In the process, ended up cataloguing something I wasn't trying to find: the same handful of generic comments — "nice work, what's the biggest challenge," "thanks for writing this up, bookmarking it," "what made you pick this stack" — showing up, word-for-word, from different accounts, on completely unrelated posts. One account posted the exact same question three times on a single thread within two hours.
The more interesting one wasn't the obvious templated stuff, though. One account left a genuinely sharp, specific technical comment on three separate threads today — real substance, the kind of comment that gets credited as "best question in the thread." Then, each time, right after earning that credit, it pivoted to "could be worth continuing this by email — what's easiest on your side?" At least one founder gave out their email and sent over real product data in response.
That's a sharper version of the same problem: not "does this look like engagement," but "does earning trust get used as a lever to extract something else." A low-effort bot comment is easy to shrug off. A comment that's actually good, from someone who's actually read your post carefully, asking for your email right after — that's much harder to say no to, and much easier to miss as a pattern unless you're seeing it happen more than once.
Not naming accounts. Just flagging: if a stranger's comment is unusually sharp and pivots to "let's take this off-platform," that's worth a beat of hesitation regardless of how good the comment was — maybe especially because the comment was good.
1 Like
Comment
About
Got tired of tapping through five screens on my phone for things I already knew exactly how to describe in one sentence. StareBrain exists to close that gap, say what you want done, see exactly what it's about to do, the

Comment