We had a tool called Agent Standards Validator on geolikeapro.com. It was supposed to tell store owners whether their site was ready for AI agents (MCP, A2A, UCP, etc.). The problem: it was a single LLM call with web search turned on, asking the model to guess which protocols a given URL implemented. No actual handshakes, no .well-known fetches, just a model reading news articles about the site and inferring.
It even had a "pseudo" status to paper over the fact that the model couldn't verify anything.
This week I rewrote it as a deterministic probe rig. What ships now:
- Real .well-known/* probes — 16 paths fired in parallel (mcp.json, agent-card.json, did.json, agentic-commerce.json, ap2.json, openapi.json, llms.txt…)
- Real MCP initialize handshake — JSON-RPC 2.0 against any discovered MCP endpoint, plus tools/list enumeration
- A2A agent-card schema validation — 8 required fields + URL HEAD probe
- 20-UA AI crawler audit — GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Bingbot, Bytespider, Applebot-Extended, MistralAI-User, and 12 others. Flags the "robots.txt-says-allowed-but-Cloudflare-403s-anyway" class of bug that the old tool missed entirely
- JS-parity check — detects hydration-shell-only sites that return 200 with empty pre-JS HTML (most AI crawlers don't run JS)
- Wikidata anchor — SPARQL exact-match on the official-website URL, not fuzzy text search. No more false positives
- Fix-it files — copy-pasteable agent-card.json, llms.txt, robots.txt delta, src/mcp-server.ts (with stack-aware mount instructions for Hydrogen / Next.js / Cloudflare Workers / plain Node), agentic-commerce.json, and a Wikidata submission brief assembled entirely from the brand's own JSON-LD (zero hallucinated facts)
- Wayback Machine fallback — when a site's bot management blocks our Worker IP (Shopify and Cloudflare both do this aggressively), we read the most recent Internet Archive snapshot instead. Real signal even when the live edge says no
- No model needed — the LLM narrative summary is gated behind a feature flag, off by default. The deterministic verdicts are the source of truth; the prose summary was decoration
What I learned along the way (the messy part):
1. Trusting the LLM was a category error. The old tool would confidently declare protocols "detected" based on what news articles said about a site, even when the actual .well-known/* files returned 404. The fix wasn't "better prompting", it was "stop asking the model to do the verification work."
2. Cloudflare Workers share IP reputation across all tenants. Hammering one Shopify store from a Worker quickly puts you in penalty box for all Shopify stores. Throttling helped (3 in flight, 200ms stagger). UA spoofing helped less. Wayback solved it.
3. SPA catch-all routes are a false-positive minefield. If a site returns 200 + <!DOCTYPE html> for /openapi.json (because the SPA serves index.html for everything), it's not "advertised but broken" — it's "not_found with extra steps." Detecting text/html on a JSON-expecting endpoint and downgrading to not_found killed a whole class of bad verdicts.
4. Hostname is a better brand signal than <title> for anywhere the title is marketing copy ("Book your flight now…"). <branddomainname>.com → brandname works on every domain; the title only works when the brand puts itself there.
Live at geolikeapro.com/tool (Agent Standards tab). Free to run. Open to feedback on what verdicts feel wrong on your site.
---