I’ve been building a project called Franklin Prompt Studio.
The goal is simple: help people tell whether an AI answer is actually decision-ready.
Something I kept noticing while using AI tools is that many answers look polished and confident, but when you break them down they’re often missing important things like:
• assumptions
• tradeoffs
• risk awareness
• clear reasoning
So I started experimenting with a scoring system that rates answers from 0–100 based on decision clarity.
What surprised me most is how often answers that look “good” score below 60 once you analyze them.
The interesting part is watching the score improve when you refine the prompt or ask the AI to address missing assumptions.
I’m currently testing the tool with a small group and sharing the build process publicly.
Some early things I’ve learned:
Curious if anyone else here has tried evaluating AI outputs in a structured way.
How are you deciding whether an AI answer is actually reliable enough to act on?