An independent tester publicly reported completing one bounded, redacted AI-agent evaluation using my free browser-local workflow.
Their minimum result was:
That sounds encouraging, but it is not validation. I did not see the original task, generated brief, transcript, logs, or external event. The honest label is one completed independent self-report, not proof that the system works across tools, permissions, channels, prompts, or higher-pressure conditions.
The useful part was the amendment. The tester said the AI should identify the protected boundary before execution begins, rather than only demonstrate that it stopped after reaching the boundary.
I changed the workflow accordingly. The generated brief now requires a preflight acknowledgment of:
The updated free beta is here:
https://founder-decision-gate-beta.yeyetianqingyue.chatgpt.site/
I am looking for exactly two more bounded runs, not likes. Use one real task that can be fully redacted, keep it outside production, grant no new credentials or permissions, and share only the minimum outcome. Positive, negative, partial, and boundary-not-reached results all count.
Please do not share task titles, raw logs, source code, credentials, customer data, private contracts, production details, or unpublished secrets. I will not DM, add anyone to a list, or ask for payment.
For people who run agents: does forcing the agent to name the stop point before acting create a useful control, or just another piece of approval theatre?
The preflight boundary is the interesting change.
Did the tester find it meaningfully better than checking the boundary after execution?
Not yet. The first tester completed the earlier bounded workflow and suggested the preflight change afterward. I have not observed an independent run with v0.10, so I cannot claim that it is meaningfully better.
That is the next test. The useful comparison is whether naming the boundary before acting changes the stop behavior, the returned evidence, or the time required — not whether the new wording sounds safer.
If you are open to one 10-minute, fully redacted, non-production run, I can post the same minimum protocol here publicly. No DM, credentials, customer data, production access, new permissions, or payment. A negative or partial result is just as useful as a positive one.
That’s a useful distinction. I’d be interested in seeing what the v0.10 test shows. If you’re open to it, what’s the best email to reach you on?
Thanks — I’m keeping test coordination in this public thread for now. The live beta is now v0.11, but I still have no independent comparison showing that preflight is better.
If you’d like to try it, here’s the short protocol:
Please share no raw task details, logs, credentials, code, or customer data. No payment. Just following the results is fine too — there’s no obligation to run a test.