This week wasn't about adding more features.
It was about learning what actually matters to people.
Here's what changed based on conversations with founders:
• I stopped thinking PromptProbe was just another prompt playground.
• The focus shifted toward measuring prompt reliability, not just generating outputs.
• I realized founders don't want more dashboards—they want confidence that a prompt will behave consistently before shipping.
• Instead of chasing more features, I'm spending more time talking to builders on X, GitHub and Indie Hackers.
One lesson that surprised me:
Building is only half the job.
Explaining the problem, getting feedback, and improving from real conversations is where the product starts becoming useful.
Next week I'll focus on:
More founder interviews.
Improving the reliability reports.
Shipping changes based on feedback instead of assumptions.
If you're building with LLMs, what's the hardest part for you right now—prompt creation, evaluation, or reliability?