Sombra

Save web research into AI-ready context via App and MCP.

Visit Website
March 5, 2026 Tracking How Your Research Evolves: History in Sombra

We're going to dive into a feature on Sombra that I think is quietly one of its most important — tracking the history of your web scrapes, notes, and distilled context as your collections evolve over time.

The final output looks like this, for those of you who want to skip ahead, a live synced public link: Somatic Mosaicism & Clonal Evolution

Sharing can also be a snapshot of a specific point in time - you can choose when you create a public link. Collections are always private without this.

Still here? Let's see where that came from... 🦉

It started with a question from a 10-year-old

I'd started a chat with Claude, prompted by a question from my son. He wanted to know about life that doesn't originate from a seed. Kids ask amazing questions - the kind that send you tumbling down rabbit holes you'd never find on your own.

Talking with AIs reminds me of the old Wikipedia deep-dives. You start with one question and twenty minutes later you're in a completely different universe. This was no exception.

I know less than nothing about biology, and Claude was understandably already slapping my wrists about some of my assumptions - but that's to be expected. I was curious, and curiosity doesn't require credentials.

Down the rabbit hole: Pando

Claude pointed me at something entirely new to me. I'd never heard of Pando before - a single clonal organism of quaking aspen in Utah, estimated to be thousands of years old. What looks like a forest of individual trees is actually one genetically near-identical organism connected by a shared root system. Truly astonishing.

For those of you who are seasoned Wikipedia delvers, this is all familiar territory - you start digging into a topic and suddenly you're learning about somatic mosaicism, clonal hematopoiesis, and how mutation rates scale with lifespan across mammals.

Engage superpowers: I had Sombra connected

While I was going down this rabbit hole, Sombra was right there with me. I was saving pages, adding notes, and building up a collection as I went - turning a random, wandering exploration into something structured and reusable.

At this point, my random exploration was concrete. I had a collection with links, content, and notes, plus a cheatsheet tying it all together.

Distilling the context

I wanted this collection to be more than just a pile of bookmarks and snippets. So I asked Claude - via a specialised prompt from Sombra - to distill everything into a coherent context and save it back to the collection.

The cover context you saw on the link above was distilled and added to the collection, along with the sources and notes contributing to it. From here, when I come back to this in the future, I can look briefly at the distillation, or look at the saved scrapes, and get my brain back into the same place it was when we created the collection.

Notes, sources, and context all in one place. And synced to my Dropbox and Google Drive.

Switching AIs, same collection

Later, I wanted a more focused context specifically about Pando. ChatGPT also has access to the same collection via Sombra's MCP integration, and it updated the context for me. I didn't need to get it up to speed - it just read the same docs and sources we had already saved.

This is the part that I think matters most in practice. Your research isn't locked into whichever AI you happened to be using when you did the work. The collection is the source of truth, not the chat.

History: seeing what changed

Directly in the UI, I can see exactly what changed by clicking on the history icon. Every save — whether it's a web scrape, a note, or a distilled context update — is versioned and saved.

AIs can go off on tangents. Being able to trust them to write and update data for you without a safety net is a big ask. The history feature, across all save types, gives you the confidence to let AIs update your collections freely, knowing you can always see how your research has evolved and go back to what was there before.

It's a simple idea: version everything, show the diffs, let you roll back. But when you're building a research workflow where multiple AIs are reading and writing to the same knowledge base, it becomes essential.

How would you use this?

I'd love to hear how you'd put this to work. Research projects? Competitive analysis? Learning journals? The combination of structured collections, AI-powered distillation, and full history tracking opens up some interesting workflows.

Thanks for reading, and happy research! 📚

Comment

March 5, 2026 Google Drive sync support

Sombra has a pretty great story on collaboratively getting your research into AI assistants via MCP - but we're well aware that sometimes you want to get your data into other tools too, or even just have a backup.

We already had Dropbox support, but we've just added Google Drive too!

All your notes, contexts and clippings are automatically synced into Google Drive as pure markdown, into a specific folder just for Sombra. If you have both Dropbox and Google Drive connected, don't worry - your data will pop up in both.

Try chatting with ChatGPT or Claude with Sombra hooked up, and you can watch all your data syncing as you chat and work.

Magical!

Comment

February 22, 2026 Sombra — your AI agent's research library

I got tired of rebuilding context across every AI tool I use: Claude web, iOS, Code, Cursor. Sombra saves, organizes and distills web research into hard-hitting context, available anywhere you can connect via MCP.

The problem

Every developer using AI coding agents has this workflow: you research something (an API, a migration path, a library you've never used) across a dozen tabs. You read, compare, form opinions. Then you switch to Claude Code or Cursor and the agent knows none of it. So you start copy-pasting docs into prompts, re-explaining the same context across sessions, and watching your agent hallucinate because it's missing the one page you read yesterday.

This is a context engineering problem. Most agent failures aren't model failures. They're context failures. Your tools are smart enough. They just don't know what you know.

This isn't only a developer problem, but developers are on the frontline of this given the volume of technical docs and the pace of change.

Knowledge workers unite in chasing the same loop - research, accumulate, distill, repeat. Sombra (your research shadow!) does this for you, anywhere you can connect over MCP, or you can publish curated content into publicly available URLs citing sources and captured distillations. Examples below.

I'd be really curious to hear from non-developer people who have similar goals - whether that's in Sales, Marketing, Cybersecurity.

What Sombra does

Save. Any web page becomes clean, permanent markdown. Use the Chrome extension, the web app, or just ask your AI agent to save a URL mid-conversation. Images preserved. Ads and nav stripped. Think reader mode for your AI. Much less context being chewed up by bloated HTML pages. Pure signal.

Organize. Group saves into collections by project, topic, or workflow. Your AI can create, search, and rearrange collections without you lifting a finger. Asking Claude to organize your research and seeing that happen all in the assistant is particularly satisfying.

Distill. This is the part that matters. Synthesize a collection into focused project context. Code examples preserved verbatim, noise stripped out. One distillation replaces fifty pages of raw docs.

Connect. One command:

claude mcp add --transport http sombra https://sombra.so/mcp

Your agent reads your entire research library. Works with Claude.ai, Claude Desktop, Claude Code, Cursor, or any MCP client.

How we're dogfooding Sombra

Last month I migrated Sombra's server from Jetty to Http-Kit. I saved the Pedestal docs, Http-Kit's API reference, and a few migration guides into a collection. Distilled the breaking changes and key differences into collection context. Then Claude Code wrote the migration with full awareness of both stacks - no prompting gymnastics, no pasting walls of documentation. The latest docs that weren't in Claude's training data came through, session after session - both in Code and in the Desktop and Web apps for Claude for discussion.

Same pattern works for API integrations, competitive research, onboarding onto a new codebase. Save the relevant pages, distill what matters, let the agent work with real context.

Your saves are 100% private - but if you want you can publish a collection and its distillation to our public share system. Here's a couple of examples:

https://sombra.so/s/41698eb2-83a1-47c6-848f-c529f80351eb - A breakdown of a recent serious CVE affecting Acronis systems, citing sources and a distillation.

https://sombra.so/s/148f367a-3794-4610-9702-d2697a21be14 - An up to date overview of using the excellent Malli data validation library in clojure.

https://sombra.so/s/55164bbb-42ae-433b-a3b5-6a901383e730 - A demo of a writing style collection, ready to share or pass on to a non-MCP connected AI agent with access to the web (such as ChatGPT).

https://sombra.so/s/e04616b0-f1ca-4f23-ab3a-0143273a443a - bit meta, an about Sombra collection collating inspiration and sources.

Stack

Clojure and ClojureScript. Datomic. Pedestal with Http-Kit. MCP via Streamable HTTP. Chrome extension for saving pages behind login walls.

What makes it different

Nobody else does the full flow: save web pages, organize into projects, distill technical context, and serve it to AI agents via hosted MCP.

Other bookmarking services save links but have no distillation or AI integration. Obsidian and Notion are built for human browsing, not agent consumption and bring layers of accumulated cruft with them. The closest concept is Will Larson's "datapacks": local markdown files, manually managed, single-client. Sombra is hosted, persistent, multi-client, and MCP-native. Connect over OAuth, and away you go.

Status

Live and free to use. Solo founder, bootstrapped, built in southern Europe by crack, battle-hardened grumpy and veteran developers. Pro tier available for heavier usage.

I'd love feedback from anyone using AI coding agents seriously. What does your research-to-agent workflow look like today?

sombra.so

4 Comments

  1. 2

    Strong positioning. “Context engineering” is a real pain point for heavy AI users.

    Since Sombra becomes a long-term research memory layer, trust and isolation are critical:

    • When you save pages, are they stored encrypted at rest?

    • If users save pages behind login walls via the extension, how are you preventing accidental capture of session tokens or sensitive account data?

    • How are collections isolated at the database level so one user’s research can never leak into another’s MCP feed?

    Since you also allow public sharing, do you scan published collections for secrets before they go live?

    Concept is sharp. If you make the privacy model explicit and boring, developers will feel safe piping serious project research into it.

    1. 1

      Thanks for the feedback and thoughts!

      Encryption at rest: Not currently implemented beyond the private object storage's infrastructure-level protections. Encryption at rest is on the roadmap - how would you expect encryption to work? A user supplied key..? What would you like to see?

      Extension / sensitive data: The system captures content-related DOM elements only. Form inputs, JS only values and non-visible elements are explicitly excluded, and the markdown conversion layer strips everything that isn't rendered text. So, session tokens sitting in hidden fields or JS state are never captured. I've open sourced the URL->markdown conversion library, so you can see how data is captured there - https://github.com/dazld/r11y - session values, cookies etc are excluded by design before going anywhere near the server.

      DB isolation: All queries are scoped strictly to the authenticated user. There are no collection-level or global iteration features at all that could bleed across users. This is structural, not just a filter. Just to say, we're using the same data layer that powers a bank in Brazil - explicitly chosen because of the controls and transactional correctness it offers, and not a raw document store such as mongo.

      Public collections: Only user-authored notes and distilled context are shared, never the third-party page content itself. The user explicitly pushes a share button, it's not passive. The raw scraped content stays private.

      You make an excellent point about making a specific page detailing all the above. Will work that up.

      1. 1

        Appreciate the detailed response. Excluding form inputs, hidden fields, and non visible DOM elements in the capture layer is a strong safeguard. Open sourcing the conversion library also helps with transparency.

        Keeping raw scraped page content private and requiring an explicit share action for notes is a good boundary. That reduces accidental data exposure from saved research.

        A few areas worth considering as usage grows:

        • Add encryption at rest for stored research content and collections

        • Log and monitor unusual access patterns to detect scraping or automated extraction attempts

        • Provide users with clear controls to delete stored research data and shared collections

        For browser extensions, trust grows when the data flow is very clear. A short security page describing capture rules, storage boundaries, and sharing behavior would help developers feel comfortable saving sensitive research.

        1. 1

          Valuable discussion, thank you.

          Encryption still on backlog, but we've honestly detailed our security posture and measures here - https://sombra.so/security - thoughts welcome!

About

I got tired of rebuilding context and info across every AI tool - Claude web, iOS, Code, Cursor. Sombra saves, organizes and distills web research into hard-hitting context, available anywhere you can connect via MCP.