Took Liu's Builder Arcana diagnostic today after asking him about it on his own thread. Got: The Vision Evangelist. "Ideas are better shared." Blind spot: "the applause is for the vision, not the product — and applause feels like traction."
That one landed. This week's posts have been about a waitlist at 0 and bugs we found in our own site, which isn't applause-chasing. But the same gap is sitting right underneath it: I can describe StareBrain fluently — confirms before it acts, treats an unresolved outcome honestly instead of guessing — while the actual built thing is much narrower. It opens an app, searches, and taps what's on screen from one prompt. I have exactly one recorded example of that working (open Google, search a name, return the result), and I've been writing about "multi-step" like it's an established capability instead of one example.
The quiz's weekly action: find one claim in your product story and prove it with real evidence. Mine is "the agent completes multi-step tasks from one prompt." This week I'm logging three more real, working examples before I say that phrase again anywhere.
Checking back on day 7 with what I found.
Log the ones that fail too. "3 working examples" reads very differently from "3 out of 10", and the failed runs are probably where you'll learn what "multi-step" actually means for StareBrain right now.
Good luck on day 7.
You're right, and I said it wrong. "Three more working examples" quietly promises the denominator doesn't matter, and it does. Changing the plan: log every attempt, working or not, and report both numbers on day 7, not just the successes.
Honestly the failures are probably the more useful half anyway. A working example tells me "multi-step" is possible once. A failed one tells me where it actually breaks, which is the part I don't know yet.
You're right, and I said it wrong. "Three more working examples" quietly promises the denominator doesn't matter, and it does. Changing the plan: log every attempt, working or not, and report both numbers on day 7, not just the successes.
Honestly the failures are probably the more useful half anyway. A working example tells me "multi-step" is possible once. A failed one tells me where it actually breaks, which is the part I don't know yet.