Assay

Adversarial AI code verification

Visit Website
February 19, 2026 We found 185 bugs in LiteLLM and 41 in Chatbot UI using adversarial AI verification

I've been building with AI coding tools for months. And I kept hitting the same problem: everything compiles, everything looks great, but features are quietly broken.

So I built a tool to catch it. Then I ran it on popular open-source projects to see how bad the problem really is.

The results

We ran Assay (adversarial AI code verification) on 5 popular projects:

LiteLLM (18K stars): 1,381 claims verified, 185 bugs found (30 critical), score 78/100

Chatbot UI (28K stars): 476 claims, 41 bugs (12 critical), score 91/100

LobeChat (50K stars): 205 claims, 14 bugs (1 critical), score 87/100

Open Interpreter (55K stars): 12 claims, 4 bugs (2 critical), score 60/100

Total: 2,400+ claims verified. 250 bugs found.

Every finding is in an interactive dashboard with file paths, line numbers, and code evidence.

How it works

Assay extracts every testable claim from a codebase ("this validates auth tokens," "this handles null input," "this query prevents injection") and uses an adversarial AI pass to verify each one. Think red team for code, not code review.

The benchmark journey ($638 total)

HumanEval (164 coding tasks, $220): Baseline 86.6% pass rate. Assay reached 100% at pass@5. Self-refine was useless at 87.2%.

SWE-bench (300 real GitHub bugs, $246): Baseline 18.3% resolved. Assay: 30.3% resolved (+65.5%).

Open-source scans (~$100): 5 projects, 2,400+ claims, 250 bugs.

What I learned

1. The biggest projects have the most bugs. LiteLLM (52 routes) had 185. Smaller projects scored higher.

2. Critical bugs hide in plain sight. These projects have thousands of stars, active communities, and regular releases. The bugs aren't in obscure corners.

3. Traditional tools don't catch semantic bugs. Linters check syntax. Type checkers check types. Nothing checks if the code does what it claims.

Try it yourself

npx tryassay assess /path/to/your/project

Free, open source. Uses the Anthropic API (~$2-3 for a small project, ~$30-50 for a large one).

GitHub: https://github.com/gtsbahamas/hallucination-reversing-system

Live dashboards: https://tryassay.ai

Free offer: Reply with a repo link and I'll run Assay on it and share the dashboard. No charge, no strings. I want the data.

Have you caught bugs in AI-generated code that traditional tools missed?

Comment