I built skilleval because I kept installing community Claude SKILL.md files and having no idea if they were actually helping or just adding noise to my system prompt.
The problem: skills are just markdown instructions injected into Claude's context. There's no standard way to know if a skill measurably improves output quality, and when Anthropic ships a model update, a skill that worked yesterday might silently degrade today. Eyeballing a few outputs isn't a benchmark.
What it does: skilleval runs each of your test tasks through Claude twice, once with your skill active and once without, then sends both outputs to a blind LLM judge (order randomized to kill position bias) that scores which one is better and by how much, on a 0 to 3 scale. You get a per-task breakdown plus an overall effectiveness score.
Why Claude runs but Gemini judges: skills are Claude-specific by design, so the runner has to be Claude. But self-judging is a known bias source, so the judge is a different model family entirely, Gemini Flash by default, though Anthropic and OpenAI are also supported as judge providers. Gemini Flash is also about 10x cheaper than Claude Haiku for structured JSON scoring.
Here's a real, unedited run against one of the sample skills in the repo:
text

I'm including this instead of a cherry-picked perfect score because that's the actual point of the tool. It caught that this particular skill only helps about half the time. That's useful information you don't get from vibes.
Cost is low enough to run routinely: about $0.01 to $0.03 per 5-task eval in the default config, or under a cent in fast/dev mode.
v0.1, MIT licensed, works via npx @dileeppandiya/skilleval ./your-skill --tasks ./tasks.yaml. There's a CI example in the README for gating PRs on skill regressions.
Genuine question for Devs: for those of you building or maintaining SKILL.md files at scale, is a single blind-judge diff score enough signal, or would you want multi-run averaging (run each task 3 times, report variance) before trusting this in CI?
That's the next thing on my roadmap and I'd rather build it based on real usage than guess.