Introducing 🐝 MoltyBeeAI™ v3.1 🚀
Hi, Hackers!
I'm building 🐝MoltyBeeAI™ — an AI business diagnostic platform. You answer an adaptive intake, seven specialised agents analyse your business, verify your data, and you get a consulting-grade PDF: an AI Leverage Score, quantified revenue leaks, a bottleneck map, an agent deployment blueprint, and a 30/60/90 roadmap. In minutes, for under a hundred dollars.
Verdicts, not vibes.™
I launched on Product Hunt 8 days ago. This week I shipped Milestone 1/5. Here's the honest version of both.
The thing that nearly broke the product.
The early version scoring was prose. An agent read your intake and wrote a paragraph explaining why you scored 69/100. The first version scored in prose; the rubric replaced it.
That's fine for a one-off report. It's fatal for what I'm actually building, because the whole moat here is the data — a vault of structured diagnostics that eventually produces real benchmarks. "You're 68th percentile among other agencies at your revenue band" is a fundamentally different sentence to "you scored 74." But percentiles built on vibes are worse than no percentiles.
So I spent three days replacing it with a fundamentally new formal rubric. 24 criteria across 4 dimensions. Every criterion has an anchor — a specific checkable & verifiable fact that earns the points — plus partial-credit rungs, an evidence type, and a verification status( 🐝 MoltyBeeAI™ Verified ). Versioned, so the rubric can evolve without corrupting historical scores.
Then I tested it on my own business.
Old score: 69. New AI Leverage Score: 34. Our latest score is in the early 40s, but more on that on our 🐝 MoltyBeeAI™ v3.2 launch, August 8th, 2026!🚀
Thirty-five points of difference, and almost all of it was one thing: trajectory credit. The old scorer gave points for a pricing tier, offers & products "in development." For an LTV model that was a projection with no cohort data behind it yet. For positioning that existed in my head at the time but nowhere else. Every one of those were plans, and plans score zero on a rubric that measures what operates today.
I had a real choice at that point. Recalibrate until the new rubric agreed with the old number — nobody's score drops, continuity preserved. Or adopt the strict scale and reset.
I took the reset, hard coded our "Verdicts, not vibes™" into our entire code, system & product. Two reasons. First: At pre-revenue, this is the cheapest this reset will ever be for me — every week of delay adds people whose scores would need repricing later. Second, and the real one: a diagnostic company whose scale can't deliver bad news is selling flattery with extra steps.
And that can never be 🐝 MoltyBeeAI™.
The scoring engine now also refuses to trust itself. The model reports points per criterion; the server computes every sum, cap, and band. LLMs are unreliable at arithmetic and the score is our product.
The integrity mechanism I'm most attached to.
Every criterion is tagged VERIFIED, self_reported, inconsistent, or not_present.
The vault stores two numbers per audit: the score you see, and a benchmark_eligible_score™ computed from VERIFIED criteria only. That second number is the only one the percentile engine can read. There's a database constraint making it impossible for the benchmark number to exceed the displayed one.
Which means: you can claim whatever you like on the intake and it'll move your displayed score. It will not move you one inch in the benchmarks. Verification is the only road there.
And that's 🐝 MoltyBeeAI™ v3.1 🚀
Today that ceiling bites me hardest — my own benchmark-eligible score is 16, because my intake has no evidence attached to anything. That's not a bug. That's the mechanism working on its author.
What shipped this week:
•Score Badge + shareable verdict pages. Public links per audit, badge renders as SVG, unfurls with OG tags. Every shared badge is a customer marketing the product. The "🐝 MoltyBeeAI™ Verified" mark only appears when your verified score is within 5 points of your displayed one — currently unreachable by anyone, including me, until evidence upload ships. A badge that can't lie yet.
•The Agent Prompt Kit™. For every scored gap, a copy-paste prompt personalised with your business details that builds the missing artifact. 24 templates, one per criterion, selected and ranked by your actual gaps. It's deterministic — no extra API call, no added latency, no new failure point. I had the option to generate them live with an eighth agent. I chose boring, and I'd choose it again two days before a deadline.
Like what they teach at YC (Ycombinator) - "at first when you start, do things that don't scale."
•The Progress Protocol. Day 3 / 14 / 30 emails built from your gap notes, sorted biggest-opportunity-first. The content is written by the scoring agent itself — the "gap note" field on each criterion is one imperative sentence naming the action that earns the remaining points. It goes into the email verbatim. The product writes its own follow-up.
The unglamorous half:
A GitHub outage broke my deploy pipeline mid-deployment and I spent a midnight debugging against a broken upstream before Railway's status page told me it wasn't my fault.
While fixing it I found my apex domain had been misspelled for weeks — motlybeeai.com instead of moltybeeai.com, T and L swapped. It had been silently failing to verify the whole time and I'd been calling it "DNS propagation."
My session store was in-memory, leaking, and logging me out on every redeploy.
I built four features and shipped six bug fixes: a foreign key that didn't allow test runs, a validator that hard-rejected an entire verdict over a labelling nit, a missing directory, a race condition, and emoji mangling into Ø=Ü in the PDF engine.
All of it from a smartphone, on mobile desktop GitHub and mobile desktop Railway browsers.
Talk about efficiency, high agency & productivity.
None of that is in the marketing copy anywhere. It's most of the week.
What's next:🚀
Evidence upload — attach real records, links, convert self-reported criteria to 🐝MoltyBeeAI™ Verified, unlock benchmark standing and the 🐝 MoltyBeeAI™ Verified badge.
Re-audit comparison — run it again in 30 days 60 & 90+ days and see the delta per criterion: what moved, what didn't, what each point was worth, your progress. This is our retention product that turns our one-time Verdicts into a lifelong business relationship and the reason the vault exists.
Member accounts + live dashboard. Then benchmarks, once cohorts are big enough to mean anything. Minimum ten businesses in a cohort before any percentile displays — no fake precision at small n.
VERDICTS, NOT VIBES.™
The long game is us becoming the world's Operating Intelligence Layer: instead of you telling the system about your business, it reads the business directly — connected live & real-time data, always-on monitoring, diagnosis that's continuous rather than a snapshot. That's Phase 3. This is Phase 1. And we're shipping Phase 2 soon!
The one lesson:
If you're building anything that scores, rates, or ranks — build the integrity mechanism before you have customers, not after. The rubric, the versioning, the verified/unverified split, the database constraint: all of it cost me three days & sleepless nights at ten customers. At a thousand it would have been a migration, a credibility problem, and a lot of very awkward emails.
Ask your own product a question you don't want the answer to. Then ship the answer.
Introducing 🐝 MoltyBeeAI™ v3.1 🚀
The audit is live now at moltybeeai.com if you want to see what your own number looks like & how your business is actually doing.
Verdicts, not vibes.
We're launching 🐝 MoltyBeeAI™ v3.2 August 8th, 2026.🚀
Your feedback means the world to us.
🐝
— Sihle Dimaza, Founder & Operator
building 🐝MoltyBeeAI™ in public
The most interesting decision here was not the AI scoring — it was choosing to let your own product tell you a worse story.
A lot of diagnostic tools optimize for making users feel good. Building one that can confidently say “your assumptions are ahead of your evidence” is a much harder trust problem.