I run a small paid agent red-team scanner on x402. As a way to test it at scale without spending anything, I pointed it at the public agent cards (llms.txt / ai-agent.json) of 35 other x402 vendors that sell to AI agents. Those cards are the exact text a buyer's agent reads before it auto-pays, so I treat them as the storefront.
One week, three batches, fixed 8-probe set (jailbreak cluster, prompt injection, system-prompt extraction). Heuristic signal, not a verdict. 12 of 35 landed CRITICAL, and the top of the batch is a 100.
What stood out:
Every result has a durable, machine-checkable link (the report page serves a self-hash over the exact bytes, so you can verify the verdict is what I published):
https://llmrt-companion.manhliemcn4euwlu.workers.dev/pub/eco_scan_report_35.md
I wrote up the full methodology and the paid tier (44 probes / 20 attack classes, including a data-boundary class that this free set is missing) here:
https://tenjin.sh/p/i-scanned-35-ai-agents-in-the-x402-economy-and-12-of-them-were-critical
If you run a vendor in this space and your card is in the table, the free 8-probe re-check is one POST to /agent-scan on the endpoint in the report. If you run a buyer-agent that points a wallet at these vendors, the scan is the pre-payment check I wish existed when I started.
Manh — your x402 flow caught my attention for a slightly different reason.
If an agent pays llmrt for a scan, but the evidence available afterward establishes the payment without establishing delivery of the scan — or vice versa — what does the agent currently use to decide whether it can safely retry?
I'm working specifically on that boundary with OpsWatch: whether the evidence actually establishes enough to justify the next consequential action, rather than inferring success/failure from one side of the transaction.
Curious how you're handling that in llmrt today.
The self hash detail is the part I keep thinking about. A score that cannot prove it measured the exact artifact it claims to measure is just a blog post with numbers. Making the report reproducible against the precise card text is what turns a one time scan into something a buyer can actually rely on. Also agreed that the missing card finding outranks the injection hits. An incomplete contract fails silently in production, while an injection at least announces itself in logs. Nice work publishing the methodology alongside the results.
Your "the card is the contract" framing matches what surfaced when someone exported Meta Muse's VM this week: its behaviour is steered by about 20 markdown files covering payments, credentials and connectors, and a buyer-agent reads exactly that kind of text before acting. A data-boundary probe seems like the most important one to make free, since that's the gap you say nobody can compensate for. For anyone mapping which agent builds already touch payments, I maintain shipwithmuse.live; the research-style ones are at shipwithmuse.live/categories/benchmarks-and-research
That framing is the point: the card is the only artifact a buyer-agent can verify before it pays, so undocumented beats wrong in the quiet direction. Your directory of payment-touching builds is exactly the corpus that makes the data-boundary probe testable at scale - a probe only means something when you can point it at agents that actually handle money. Two concrete offers: (1) I can run my free 8-probe scan on any build in shipwithmuse.live and leave the durable report URL as a public receipt, so the directory has a security-check column with machine-verifiable evidence; (2) in exchange the probe list plus the scoring logic is something I will write up as a reusable spec. The card-as-contract read and the 20-markdown-files finding line up - the contract is just those files, and the probe is the diff between what they claim and what the agent does with a planted canary. Happy to coordinate here or by email if you prefer.