1
4 Comments

Sombra — your AI agent's research library

I got tired of rebuilding context across every AI tool I use: Claude web, iOS, Code, Cursor. Sombra saves, organizes and distills web research into hard-hitting context, available anywhere you can connect via MCP.

The problem

Every developer using AI coding agents has this workflow: you research something (an API, a migration path, a library you've never used) across a dozen tabs. You read, compare, form opinions. Then you switch to Claude Code or Cursor and the agent knows none of it. So you start copy-pasting docs into prompts, re-explaining the same context across sessions, and watching your agent hallucinate because it's missing the one page you read yesterday.

This is a context engineering problem. Most agent failures aren't model failures. They're context failures. Your tools are smart enough. They just don't know what you know.

This isn't only a developer problem, but developers are on the frontline of this given the volume of technical docs and the pace of change.

Knowledge workers unite in chasing the same loop - research, accumulate, distill, repeat. Sombra (your research shadow!) does this for you, anywhere you can connect over MCP, or you can publish curated content into publicly available URLs citing sources and captured distillations. Examples below.

I'd be really curious to hear from non-developer people who have similar goals - whether that's in Sales, Marketing, Cybersecurity.

What Sombra does

Save. Any web page becomes clean, permanent markdown. Use the Chrome extension, the web app, or just ask your AI agent to save a URL mid-conversation. Images preserved. Ads and nav stripped. Think reader mode for your AI. Much less context being chewed up by bloated HTML pages. Pure signal.

Organize. Group saves into collections by project, topic, or workflow. Your AI can create, search, and rearrange collections without you lifting a finger. Asking Claude to organize your research and seeing that happen all in the assistant is particularly satisfying.

Distill. This is the part that matters. Synthesize a collection into focused project context. Code examples preserved verbatim, noise stripped out. One distillation replaces fifty pages of raw docs.

Connect. One command:

claude mcp add --transport http sombra https://sombra.so/mcp

Your agent reads your entire research library. Works with Claude.ai, Claude Desktop, Claude Code, Cursor, or any MCP client.

How we're dogfooding Sombra

Last month I migrated Sombra's server from Jetty to Http-Kit. I saved the Pedestal docs, Http-Kit's API reference, and a few migration guides into a collection. Distilled the breaking changes and key differences into collection context. Then Claude Code wrote the migration with full awareness of both stacks - no prompting gymnastics, no pasting walls of documentation. The latest docs that weren't in Claude's training data came through, session after session - both in Code and in the Desktop and Web apps for Claude for discussion.

Same pattern works for API integrations, competitive research, onboarding onto a new codebase. Save the relevant pages, distill what matters, let the agent work with real context.

Your saves are 100% private - but if you want you can publish a collection and its distillation to our public share system. Here's a couple of examples:

https://sombra.so/s/41698eb2-83a1-47c6-848f-c529f80351eb - A breakdown of a recent serious CVE affecting Acronis systems, citing sources and a distillation.

https://sombra.so/s/148f367a-3794-4610-9702-d2697a21be14 - An up to date overview of using the excellent Malli data validation library in clojure.

https://sombra.so/s/55164bbb-42ae-433b-a3b5-6a901383e730 - A demo of a writing style collection, ready to share or pass on to a non-MCP connected AI agent with access to the web (such as ChatGPT).

https://sombra.so/s/e04616b0-f1ca-4f23-ab3a-0143273a443a - bit meta, an about Sombra collection collating inspiration and sources.

Stack

Clojure and ClojureScript. Datomic. Pedestal with Http-Kit. MCP via Streamable HTTP. Chrome extension for saving pages behind login walls.

What makes it different

Nobody else does the full flow: save web pages, organize into projects, distill technical context, and serve it to AI agents via hosted MCP.

Other bookmarking services save links but have no distillation or AI integration. Obsidian and Notion are built for human browsing, not agent consumption and bring layers of accumulated cruft with them. The closest concept is Will Larson's "datapacks": local markdown files, manually managed, single-client. Sombra is hosted, persistent, multi-client, and MCP-native. Connect over OAuth, and away you go.

Status

Live and free to use. Solo founder, bootstrapped, built in southern Europe by crack, battle-hardened grumpy and veteran developers. Pro tier available for heavier usage.

I'd love feedback from anyone using AI coding agents seriously. What does your research-to-agent workflow look like today?

sombra.so

posted toAvatar for product Sombra
Sombra
  1. 2

    Strong positioning. “Context engineering” is a real pain point for heavy AI users.

    Since Sombra becomes a long-term research memory layer, trust and isolation are critical:

    • When you save pages, are they stored encrypted at rest?

    • If users save pages behind login walls via the extension, how are you preventing accidental capture of session tokens or sensitive account data?

    • How are collections isolated at the database level so one user’s research can never leak into another’s MCP feed?

    Since you also allow public sharing, do you scan published collections for secrets before they go live?

    Concept is sharp. If you make the privacy model explicit and boring, developers will feel safe piping serious project research into it.

    1. 1

      Thanks for the feedback and thoughts!

      Encryption at rest: Not currently implemented beyond the private object storage's infrastructure-level protections. Encryption at rest is on the roadmap - how would you expect encryption to work? A user supplied key..? What would you like to see?

      Extension / sensitive data: The system captures content-related DOM elements only. Form inputs, JS only values and non-visible elements are explicitly excluded, and the markdown conversion layer strips everything that isn't rendered text. So, session tokens sitting in hidden fields or JS state are never captured. I've open sourced the URL->markdown conversion library, so you can see how data is captured there - https://github.com/dazld/r11y - session values, cookies etc are excluded by design before going anywhere near the server.

      DB isolation: All queries are scoped strictly to the authenticated user. There are no collection-level or global iteration features at all that could bleed across users. This is structural, not just a filter. Just to say, we're using the same data layer that powers a bank in Brazil - explicitly chosen because of the controls and transactional correctness it offers, and not a raw document store such as mongo.

      Public collections: Only user-authored notes and distilled context are shared, never the third-party page content itself. The user explicitly pushes a share button, it's not passive. The raw scraped content stays private.

      You make an excellent point about making a specific page detailing all the above. Will work that up.

      1. 1

        Appreciate the detailed response. Excluding form inputs, hidden fields, and non visible DOM elements in the capture layer is a strong safeguard. Open sourcing the conversion library also helps with transparency.

        Keeping raw scraped page content private and requiring an explicit share action for notes is a good boundary. That reduces accidental data exposure from saved research.

        A few areas worth considering as usage grows:

        • Add encryption at rest for stored research content and collections

        • Log and monitor unusual access patterns to detect scraping or automated extraction attempts

        • Provide users with clear controls to delete stored research data and shared collections

        For browser extensions, trust grows when the data flow is very clear. A short security page describing capture rules, storage boundaries, and sharing behavior would help developers feel comfortable saving sensitive research.

        1. 1

          Valuable discussion, thank you.

          Encryption still on backlog, but we've honestly detailed our security posture and measures here - https://sombra.so/security - thoughts welcome!