I got tired of the guesswork, so I built an AI Evaluator inside my tool, PromptKelp. It’s a "Senior Prompt Engineer" agent that roasts my system prompts for logic errors, hygiene, and self-consistency.
The most satisfying part? I’m using PromptKelp to manage the prompts that power PromptKelp.
I just pushed a new feature today: Debug Logs Analysis. It allows me to pull in production logs and let the evaluator compare the "intended" prompt logic against the "actual" model behavior. It’s already caught two major contradictions in my health-agent logic that I would have missed.