3
1 Comment

Anyone with payment or auth code want to try a merge-blocking AI review? I'll set it up for free

Quick backstory: I built a GitHub Action called Verdict. When a PR touches payment, auth, or agent code, it gets reviewed by two AI models from different companies plus Semgrep, and if they disagree (or anything serious comes up) the merge is blocked until a human overrides it.

The honest state of things: it has only ever run on my own repos. On my payment gateway it caught three real bugs in a fix I was about to merge, which felt good. But nobody outside my own repos has tried it, and I can't fake that part.

So here's my ask. If you've got a repo with payment, auth, or agent-execution code, I'll write the config and workflow file for you (about 15 minutes on my end). You add them yourself, so I never need access to your repo. It runs on your own Anthropic/OpenAI keys, and in my runs the API cost has been small. You can pull it out whenever you want.

All I want back is honest feedback: what's annoying, what's wrong, what it missed. If it flags nothing useful, tell me that too, it's just as helpful.

Repo and the write-up of the three bugs: https://github.com/xKazeex/verdict-action

on September 28, 2026
  1. 1

    Blocking on disagreement makes the false-alarm rate the number to watch, since a gate that cries wolf gets overridden by habit. One data point for picking the second model: a 7-model test (33 runs each) found Muse Spark 1.2 the fastest at about 18 s typical, and it caught a bug the other reviewers missed, but it approved false alarms when asked to judge findings. So it may fit the finder role better than the tiebreaker: https://shipwithmuse.live/builds/muse-spark-1-2-as-a-code-reviewer-vs-6-models (I help curate it)