ToolCairn

Tool intelligence for AI coding agents

Visit Website
May 14, 2026 Why your AI agent keeps picking the wrong libraries

Hi IH,

We shipped ToolCairn last week on npm and on the official MCP Registry. It's an MCP server that helps AI agents pick the right tools — and the right stack of tools — for a build task, drawing from a live view of 35+ open-source registries (npm, PyPI, Cargo, Maven, Go, RubyGems, NuGet, and more).

The problem we kept hitting

Your AI agent picks tools from training data, not from this week's reality. Often that's fine. Sometimes the tool was the right answer two years ago and isn't now. Sometimes the version it reaches for has a peer-dep conflict three layers deep. Sometimes the library was abandoned and nobody told the model.

You don't find out at pick time. You find out an hour into the implementation, deep into a prompt chain, when something quietly breaks. Then you back out, switch libraries, re-prompt, and burn another loop. Every wrong pick is a refactor — and a refactor is the most expensive thing you can ask an agent to do.

What ToolCairn does

ToolCairn sits inside the agent's MCP setup as a context layer the agent consults before committing to a pick. The agent describes the build task. ToolCairn returns the right pick — or the right full stack — drawn from a live graph of thousands of tools across those 35+ registries, with version metadata, usage context, and composes-with relationships baked into the scoring.

Picks aren't ranked by what was popular in the model's training cut. They're ranked on what's maintained today, what composes with the rest of the agent's picks, and which versions resolve together at install time.

The agent then writes against picks that won't blow up at install or three iterations into the implementation. The hour stays yours. The tokens stay yours.

After use, ToolCairn closes the loop — the graph learns from real outcomes, which sharpens the next pick.

What we haven't solved yet

This is v1.0. Three honest gaps:

1. Trust signals are thinner than we want. We expose license, last-updated, and peer-dep coherence today. Supply-chain signals, maintainer activity, and breaking-change frequency are on the list — none shipped.

2. Client coverage is one. Claude Code is the only fully tested host today. The MCP protocol means Cursor, Cline, Windsurf, and Codex should all work, but per-client install + tool-call validation is still in progress.

3. Registry coverage isn't uniform. The graph is deepest in registries with the most active maintainer ecosystems and thinner elsewhere. Edge-case picks in less-active registries will surface as wrong picks first.

Try it

- https://toolcairn.neurynae.com

- Install (Claude Code): claude mcp add toolcairn -- npx @neurynae/toolcairn-mcp

- GitHub (MIT): https://github.com/neurynae/toolcairn-mcp

What we'd love help thinking through

1. When you've watched an agent pick a wrong tool, where did it go wrong — framework, library, or version level?

2. What trust signal would actually shift your willingness to let an agent act on a recommendation without reviewing every line?

3. Which client integration matters most to you next (Cursor, Codex, Windsurf, Cline, VS Code AI)?

4. If you've shipped anything that spans multiple package registries, where did your indexing strategy fall apart first?

Around all day. Comments here, GitHub issues, replies — all get read. Shipping visible fixes within 72 hours during launch week.

— The NEURYNAE team

Comment

About

MCP solved how agents call tools, not which tool to pick. Selection — not access — is the bottleneck. ToolCairn fills that gap: ranked picks, version-aware, neutral.