
Assay
Adversarial AI code verification
I've been building with AI coding tools for months. And I kept hitting the same problem: everything compiles, everything looks great, but features are quietly broken.
So I built a tool to catch it. Then I ran it on popular open-source projects to see how bad the problem really is.
The results
We ran Assay (adversarial AI code verification) on 5 popular projects:
LiteLLM (18K stars): 1,381 claims verified, 185 bugs found (30 critical), score 78/100
Chatbot UI (28K stars): 476 claims, 41 bugs (12 critical), score 91/100
LobeChat (50K stars): 205 claims, 14 bugs (1 critical), score 87/100
Open Interpreter (55K stars): 12 claims, 4 bugs (2 critical), score 60/100
Total: 2,400+ claims verified. 250 bugs found.
Every finding is in an interactive dashboard with file paths, line numbers, and code evidence.
How it works
Assay extracts every testable claim from a codebase ("this validates auth tokens," "this handles null input," "this query prevents injection") and uses an adversarial AI pass to verify each one. Think red team for code, not code review.
The benchmark journey ($638 total)
HumanEval (164 coding tasks, $220): Baseline 86.6% pass rate. Assay reached 100% at pass@5. Self-refine was useless at 87.2%.
SWE-bench (300 real GitHub bugs, $246): Baseline 18.3% resolved. Assay: 30.3% resolved (+65.5%).
Open-source scans (~$100): 5 projects, 2,400+ claims, 250 bugs.
What I learned
1. The biggest projects have the most bugs. LiteLLM (52 routes) had 185. Smaller projects scored higher.
2. Critical bugs hide in plain sight. These projects have thousands of stars, active communities, and regular releases. The bugs aren't in obscure corners.
3. Traditional tools don't catch semantic bugs. Linters check syntax. Type checkers check types. Nothing checks if the code does what it claims.
Try it yourself
npx tryassay assess /path/to/your/project
Free, open source. Uses the Anthropic API (~$2-3 for a small project, ~$30-50 for a large one).
GitHub: https://github.com/gtsbahamas/hallucination-reversing-system
Live dashboards: https://tryassay.ai
Free offer: Reply with a repo link and I'll run Assay on it and share the dashboard. No charge, no strings. I want the data.
Have you caught bugs in AI-generated code that traditional tools missed?

Comment