3
13 Comments

I scored my own startup with my own product. It gave me a 34/100 AI Leverage Score™. I shipped the 34.

Introducing 🐝 MoltyBeeAI™ v3.1 🚀

Hi, Hackers!

I'm building 🐝MoltyBeeAI™ — an AI business diagnostic platform. You answer an adaptive intake, seven specialised agents analyse your business, verify your data, and you get a consulting-grade PDF: an AI Leverage Score, quantified revenue leaks, a bottleneck map, an agent deployment blueprint, and a 30/60/90 roadmap. In minutes, for under a hundred dollars.

Verdicts, not vibes.™

I launched on Product Hunt 8 days ago. This week I shipped Milestone 1/5. Here's the honest version of both.

The thing that nearly broke the product.

The early version scoring was prose. An agent read your intake and wrote a paragraph explaining why you scored 69/100. The first version scored in prose; the rubric replaced it.

That's fine for a one-off report. It's fatal for what I'm actually building, because the whole moat here is the data — a vault of structured diagnostics that eventually produces real benchmarks. "You're 68th percentile among other agencies at your revenue band" is a fundamentally different sentence to "you scored 74." But percentiles built on vibes are worse than no percentiles.

So I spent three days replacing it with a fundamentally new formal rubric. 24 criteria across 4 dimensions. Every criterion has an anchor — a specific checkable & verifiable fact that earns the points — plus partial-credit rungs, an evidence type, and a verification status( 🐝 MoltyBeeAI™ Verified ). Versioned, so the rubric can evolve without corrupting historical scores.

Then I tested it on my own business.

Old score: 69. New AI Leverage Score: 34. Our latest score is in the early 40s, but more on that on our 🐝 MoltyBeeAI™ v3.2 launch, August 8th, 2026!🚀

Thirty-five points of difference, and almost all of it was one thing: trajectory credit. The old scorer gave points for a pricing tier, offers & products "in development." For an LTV model that was a projection with no cohort data behind it yet. For positioning that existed in my head at the time but nowhere else. Every one of those were plans, and plans score zero on a rubric that measures what operates today.

I had a real choice at that point. Recalibrate until the new rubric agreed with the old number — nobody's score drops, continuity preserved. Or adopt the strict scale and reset.

I took the reset, hard coded our "Verdicts, not vibes™" into our entire code, system & product. Two reasons. First: At pre-revenue, this is the cheapest this reset will ever be for me — every week of delay adds people whose scores would need repricing later. Second, and the real one: a diagnostic company whose scale can't deliver bad news is selling flattery with extra steps.

And that can never be 🐝 MoltyBeeAI™.

The scoring engine now also refuses to trust itself. The model reports points per criterion; the server computes every sum, cap, and band. LLMs are unreliable at arithmetic and the score is our product.

The integrity mechanism I'm most attached to.

Every criterion is tagged VERIFIED, self_reported, inconsistent, or not_present.

The vault stores two numbers per audit: the score you see, and a benchmark_eligible_score™ computed from VERIFIED criteria only. That second number is the only one the percentile engine can read. There's a database constraint making it impossible for the benchmark number to exceed the displayed one.

Which means: you can claim whatever you like on the intake and it'll move your displayed score. It will not move you one inch in the benchmarks. Verification is the only road there.

And that's 🐝 MoltyBeeAI™ v3.1 🚀

Today that ceiling bites me hardest — my own benchmark-eligible score is 16, because my intake has no evidence attached to anything. That's not a bug. That's the mechanism working on its author.

What shipped this week:

•Score Badge + shareable verdict pages. Public links per audit, badge renders as SVG, unfurls with OG tags. Every shared badge is a customer marketing the product. The "🐝 MoltyBeeAI™ Verified" mark only appears when your verified score is within 5 points of your displayed one — currently unreachable by anyone, including me, until evidence upload ships. A badge that can't lie yet.

•The Agent Prompt Kit™. For every scored gap, a copy-paste prompt personalised with your business details that builds the missing artifact. 24 templates, one per criterion, selected and ranked by your actual gaps. It's deterministic — no extra API call, no added latency, no new failure point. I had the option to generate them live with an eighth agent. I chose boring, and I'd choose it again two days before a deadline.

Like what they teach at YC (Ycombinator) - "at first when you start, do things that don't scale."

•The Progress Protocol. Day 3 / 14 / 30 emails built from your gap notes, sorted biggest-opportunity-first. The content is written by the scoring agent itself — the "gap note" field on each criterion is one imperative sentence naming the action that earns the remaining points. It goes into the email verbatim. The product writes its own follow-up.

The unglamorous half:

A GitHub outage broke my deploy pipeline mid-deployment and I spent a midnight debugging against a broken upstream before Railway's status page told me it wasn't my fault.

While fixing it I found my apex domain had been misspelled for weeks — motlybeeai.com instead of moltybeeai.com, T and L swapped. It had been silently failing to verify the whole time and I'd been calling it "DNS propagation."

My session store was in-memory, leaking, and logging me out on every redeploy.

I built four features and shipped six bug fixes: a foreign key that didn't allow test runs, a validator that hard-rejected an entire verdict over a labelling nit, a missing directory, a race condition, and emoji mangling into Ø=Ü in the PDF engine.

All of it from a smartphone, on mobile desktop GitHub and mobile desktop Railway browsers.

Talk about efficiency, high agency & productivity.

None of that is in the marketing copy anywhere. It's most of the week.

What's next:🚀

Evidence upload — attach real records, links, convert self-reported criteria to 🐝MoltyBeeAI™ Verified, unlock benchmark standing and the 🐝 MoltyBeeAI™ Verified badge.

Re-audit comparison — run it again in 30 days 60 & 90+ days and see the delta per criterion: what moved, what didn't, what each point was worth, your progress. This is our retention product that turns our one-time Verdicts into a lifelong business relationship and the reason the vault exists.

Member accounts + live dashboard. Then benchmarks, once cohorts are big enough to mean anything. Minimum ten businesses in a cohort before any percentile displays — no fake precision at small n.

VERDICTS, NOT VIBES.™

The long game is us becoming the world's Operating Intelligence Layer: instead of you telling the system about your business, it reads the business directly — connected live & real-time data, always-on monitoring, diagnosis that's continuous rather than a snapshot. That's Phase 3. This is Phase 1. And we're shipping Phase 2 soon!

The one lesson:

If you're building anything that scores, rates, or ranks — build the integrity mechanism before you have customers, not after. The rubric, the versioning, the verified/unverified split, the database constraint: all of it cost me three days & sleepless nights at ten customers. At a thousand it would have been a migration, a credibility problem, and a lot of very awkward emails.

Ask your own product a question you don't want the answer to. Then ship the answer.

Introducing 🐝 MoltyBeeAI™ v3.1 🚀

The audit is live now at moltybeeai.com if you want to see what your own number looks like & how your business is actually doing.

Verdicts, not vibes.

We're launching 🐝 MoltyBeeAI™ v3.2 August 8th, 2026.🚀

Your feedback means the world to us.

🐝

— Sihle Dimaza, Founder & Operator
building 🐝MoltyBeeAI™ in public

on July 28, 2026
  1. 2

    The circularity of scoring yourself with your own scoring tool and then shipping that exact version number is pretty satisfying. Did the score itself influence what you actually shipped, or was it more of a validation that you were on the right track?

    1. 1

      Well, the score was more of an internal audit, using our own system & model, to gain more clarity on what works, and what needs our attention the most. While we work on shipping our Milestones.

      So it was more of a buil, ship, test, get feedback, iterate & improve. Then repeat the cycle. 📈🐝

      And with each iteration or improvement, the system gets better overall, the business gets better overall & the startup grows.

      That's why we conducted the audit, and continue to use the model on our own business, startup & entire system or operations.

  2. 2

    The part I’d pressure-test next is comparability across rubric versions.

    Versioning preserves the historical record technically, but a founder running the audit again in 30 days may interpret every score change as business progress, even when part of the movement came from a revised rubric.

    I’d consider showing two deltas separately: improvement under the same rubric, and change caused by the scoring model itself.

    Are you planning to freeze each audit against its original rubric, or also rescore historical audits when the rubric changes?

    1. 1

      We're built a defensive or foolproof safety mechanism against that. That's exactly the flaw our system is built to guard against.

      Every claim gets verified & checked, to avoid & protect against self inflated scores.🐝

      So rest assured that what you get as a score, will be verdicts based on evidence provided.

      1. 2

        Sorry, I may not have explained that clearly.

        I wasn’t referring to founders inflating their claims. I meant score changes caused by updates to the rubric itself.

        For example, the same verified evidence might score 34 today and 41 next month because the criteria or weighting changed. A founder could read that +7 as business progress even though nothing in the startup had actually changed.

        Your verification process protects against inflated inputs, which makes sense. I was wondering about the separate issue of rubric-version comparability: does each audit keep the exact rubric version it was scored against?

        1. 1

          The older ones were scored by the old version. New audits are scored by the new version that requires you to upload evidence to support your claims.

          Then the system analyzes everything, scores you according.

          The score only grows when the business does. Not with better claims or different word & intake inputs. But with evidence to support that, as the business grows.🐝

          1. 2

            Thanks — that clarifies the intent behind the evidence requirement.

            I’m asking about a narrower reproducibility case:

            If the exact same evidence bundle that received 34 today were rescored after the rubric or weighting changed, is it guaranteed to still receive 34?

            If not, does each audit display the rubric version it was scored against, so users can distinguish business progress from a scoring-system change?

            That version label is what would make comparisons over time trustworthy.

            1. 1

              I see. Awesome, my friend. Thank you and honestly, I love your questions. They're definitely one of the sharpest ones I've ever been asked.

              Two separate guarantees in there, and they land differently.

              Version labelling: yes. Every audit stores and displays the rubric version it was scored under. Benchmarks are computed within version cohorts only — a v1.0.2 score is never ranked against v1.1.0. Re-audit comparisons carry a versions_comparable flag; if the rubric moved between runs, the comparison says so rather than pretending the delta is all business progress.

              Rescoring determinism: partially, and I'll be precise. The rubric is deterministic — fixed criteria, fixed anchors, and every sum, cap and band computed server-side rather than by the model. What isn't perfectly reproducible is the judgement layer, since an LLM reads the intake and assigns per-criterion points. Empirically that variance is small: three consecutive runs on the same business gave 34, 34, 36, based on our front-end top of the acquisition funnel being mostly manual as opposed to our fully automated backend & delivery system, so forth. A fourth & recent run gave us 45, and that one was real — the business had shipped more than three new features between runs, and introduced a new client/customer acquisition channel, with live campaigns. Which was a real business improvement.

              But more on that on our 🐝MoltyBeeAI™ v3.2 launch, August 8th. (2026)

              Under a changed rubric, no, and deliberately so. v1.0.1 → v1.0.2 of the rubric fixed a double-penalty bug where a criterion was docked for missing evidence even when its anchor was met. Same evidence, higher score, correctly. That's exactly why versions are pinned to records: so a jump like that reads as "the ruler was fixed," not "the business improved."

              The next release adds server-side evidence verification, which shrinks the judgement surface further — verification eligibility is decided by deterministic content checks before the model sees anything.🐝

              I'm shipping this one tonight. And there's more rolling out.

              Verdicts, not vibes.™

  3. 2

    Respect for posting the real number instead of burying it. Kind of reminds me of resumes honestly, they're all self reported, nobody's verifying most of it. I'm building an AI resume tool (Applio) and that's the same problem, people trust what's written more than what's actually true.

    1. 1

      I love this, and I know your startup or model will be a hit because now more than ever, clarity & grounding verifiable facts are the currency. In a whole where anyone can fake just about anything.

      So I love your approach, my friend. And one day, I look forward to teaming up with you/y'all and creating lasting impact in the business & job world, with models that actually matter. And solve real problems. 🐝

  4. 2

    The most interesting decision here was not the AI scoring — it was choosing to let your own product tell you a worse story.

    A lot of diagnostic tools optimize for making users feel good. Building one that can confidently say “your assumptions are ahead of your evidence” is a much harder trust problem.

    1. 1

      I meant every word of, "building 🐝MoltyBeeAI™ public, together with you," and most founders would've buried this insight and share glamorous results instead.

      But I choose to grow with you guys, to build something real & much larger, from the tranches. To wherever we want to build & take 🐝MoltyBeeAI™ together!

      That's my vision & true reason behind this.

      You don’t realize how much your feedback means to me.

      Thank you.

      1. 1

        Appreciate that, Sihle.

        I respect the willingness to let the product challenge its own assumptions. That’s not an easy decision to make, especially when the easier path is protecting the original narrative.

        I’m curious — after going through this reset, what feels like the biggest unresolved challenge for MoltyBeeAI right now?

        Is it still proving the scoring model itself, getting users to trust the verdicts, defining the strongest use case, or something else?

Trending on Indie Hackers
Stop losing deals in the gap between "sounds good" and getting paid User Avatar 69 comments We scanned 50,000 domains. Your cold email list is really four systems. User Avatar 66 comments How to rank #1 on ChatGPT? User Avatar 54 comments I Tested Agenmatic for Finding Customers in Communities — Here’s What I Learned User Avatar 42 comments Building in public: a chat assistant that runs your server so you don't have to live in the terminal User Avatar 31 comments Just got invited to Web Summit Lisbon. Now I need 5 more clients in 13 days. User Avatar 28 comments