1
1 Comment

Your AI Prompt Works... Until It Doesn’t" – Building PromptPerf to fix that.

TL;DR -> I’m building https://PromptPerf.dev, a tool to test AI prompt consistency across models.
Because I got tired of shipping prompts that silently break.

#The Problem I Kept Running Into:

I'm working on AI apps where LLMs are core to the product; think GPT/Claude/Gemini APIs baked into backend logic.

At first, I treated AI prompts like static code:

Write once. Deploy. Move on.

But one day, I ran the same prompt through GPT-4o (temp 0.1) ten times.

Got five different outputs.

Same model. Same input. Completely inconsistent behavior.

I realized something obvious in hindsight:

LLMs are non-deterministic.
But our apps need deterministic results.

Then the Real Pain Started
Soon after, news came that Gemini 1.0 will be deprecated.
That meant all my working prompts for Gemini were now useless.

So now I’m stuck in this painful loop:

Re-test all prompts

Re-tune each for the new model

Revalidate edge cases

Hope nothing breaks downstream

I was losing hours doing this every time a model changed.
And it’s only going to get worse as OpenAI/Gemini/Claude roll out updates.

So I Built https://PromptPerf.dev
It’s a tool to test prompt consistency before I ship.

#Here’s what it does:

Run the same prompt across multiple models (GPT-4, Claude, Gemini etc)

Adjust temperature and repeat multiple runs

Compare outputs and measure consistency

Why I’m Posting Here
I’m in early days. Building it alone. No VC. No ads. Just coding at night and learning from real-world pain.

Here’s what I’m hoping for:

Feedback from devs/founders building AI-powered apps

Suggestions on how you currently handle prompt testing (if at all)

Anyone interested in testing early or giving product feedback

It’s not polished yet but it works, Ive done proof of concept.
You can join the waitlist here: https://promptperf.dev

Would love to hear if this is something you’d actually use.

Let’s make AI development feel a little less like gambling 🎯

🧠 P.S.
If this resonates with anyone working on LLM apps, I’d love to hear:

How often do you re-test prompts when models change?

Have you been burned by quiet model updates before?

Is prompt QA part of your workflow yet?

on April 15, 2025
  1. 1

    The quiet deprecation days got me worse than the big announced ones. Same prompt, “same” model family, suddenly a different shape — and I’d only clock it after a client-facing run looked off. What actually helped was a tiny cold-run checklist next to each keeper (one good case, one edge) and a hard rule: after any bump, those two runs happen before I trust the old wording. If both fail, I fork and rewrite from intent instead of patching in place. When a prompt dies quietly, do you catch it in a harness, or only after something ships weird?