I kept seeing the same problem: everyone’s shipping AI agents right now — support bots, sales agents, assistants — and almost nobody has a way to test them before they go live. They build it, it works in their own demo, then a real user says something weird and it falls apart. Goes off-topic, gets manipulated, breaks character, can’t handle another language.
So I built AgentProof to fix that. You paste your agent’s system prompt and it runs adversarial users against it — manipulation attempts, edge cases, emotional pressure, jailbreak attempts — then scores it on accuracy, robustness, boundary-holding, tone, and edge-case handling, and tells you exactly where it breaks with specific fixes for your prompt.
No setup, no code, no integration. Paste a prompt, get a report in about a minute.
The honest backstory: I have no coding background. I built this entire thing using AI tools over one pretty intense stretch — fought through deployment issues, a billing wall, the works. It’s live now and actually works, which still feels sur