1
0 Comments

Need feedback from people building with LLMs.

I'm building PromptProbe to help people test whether their prompts are actually reliable.

The idea is simple: Run the same prompt multiple times, compare the outputs, and see how consistent it really is.

My question is:

If you work with LLMs, what do you struggle to measure today that existing AI tools don't show you?

I'm trying to avoid building features nobody needs, so I'd love brutally honest feedback.

Try here - https://www.promptprobe.tech/

posted toAvatar for product PromptProbe
PromptProbe