Qarinah

Evidence-linked project memory for coding agents

Visit Website
August 9, 2026 Agent memory should return evidence, not just a story

Most coding-agent workflows eventually face the same continuity question: what should the next session receive?

Replaying every prior message preserves history, but it also sends a large amount of material that the current task may not need. A short rolling summary is smaller, but it can hide where a decision or tool result came from.

I built Qarinah to explore a third approach: keep the authoritative project record local, then compile a small context pack for the current task. Selected items retain their event IDs and content hashes, so the receiving agent can trace a claim back to its source.

The architecture has four practical parts:

1. An opt-in, append-only JSONL event chain for permitted project events and explicit decisions.

2. Rebuildable graph, SQLite, Markdown, and Open Knowledge Format views.

3. A bounded retrieval step that can require evidence coverage or abstain.

4. Cross-agent handoffs for Codex, Claude Code, Cursor, the CLI, and compatible MCP clients.

The six-task fixture

I wanted the context-volume claim to be reproducible, so the repository includes the evaluator and machine-readable result. The fixture retains 240 project-history records and runs six software scenarios: React editing, database migration, TypeScript refactoring, web research, production debugging, and governed release work.

Both paths receive the same current-task sources. The baseline also receives full retained history; the Qarinah path receives a cited context pack.

- Full-history path: 442,113 estimated input-context tokens

- Qarinah path: 5,682 estimated input-context tokens

- Reduction: 98.7148%, or 77.81:1 context compression

- Coverage: every required target was directly covered in the top five

- Summarization: no model-written summary items

- Estimator: ceil(characters / 4)

That percentage describes this committed repeated-context fixture. It is not a provider bill, latency measurement, or universal task-quality guarantee.

The current public npm release is Qarinah 0.1.6. It is Apache-2.0 licensed and created by Ajnas N B.

Website: https://qarinah.io

Source and reproducible artifacts: https://github.com/AjnasNB/qarinah

npm: https://www.npmjs.com/package/qarinah

White paper v1.3: https://doi.org/10.5281/zenodo.21843240

Zenodo concept record: https://doi.org/10.5281/zenodo.21547684

Benchmark method: https://qarinah.io/benchmarks/

If you work with coding agents, I would value feedback on the handoff that matters most in your workflow: a decision, a failed approach, a source citation, a tool outcome, or a permission boundary.

4 Comments

  1. 1

    That's what I am also trying to build at olivergraph.

    I guess some challenges we are seeing are:
    - sometimes coding agents work fine without fetching history in which case this adds additional token cost
    - knowing if we should be fetching prompts in additional turns
    - when codebase becomes more convulted, then this is where noise obviously happens

    Good luck tho

    1. 1
      Those are exactly the failure modes I was trying to avoid with Qarinah. The idea isn’t that the agent should always fetch memory. Retrieval is bounded around the current task rather than automatically replaying history, so when historical evidence isn’t useful the context pack should stay minimal instead of adding the whole memory layer as overhead. For later turns, the retrieval can be run again against the new task/intent rather than assuming the first context pack is still the right one. And the convoluted-codebase problem is probably where this matters most. Instead of giving the agent a larger pile of “memory”, Qarinah tries to return a small set of directly relevant evidence with event IDs/hashes, and it can abstain when the required evidence isn’t covered. So I think we’re attacking a very similar problem — would be interesting to compare Qarinah’s evidence-pack approach with what you’re doing at OliverGraph.
  2. 1

    The distinction between “remembering context” and being able to trace a claim back to evidence is really interesting.

    I think that's an important problem with agent memory: a summary can preserve what was decided while losing why it was decided or where that information came from.

    The idea of letting retrieval abstain when evidence coverage isn't sufficient is especially interesting to me. I'd rather have an agent say “I don't have enough evidence for this” than confidently reconstruct something that only sounds right.

    I wonder if, over time, the real value of agent memory will be less about remembering more and more about knowing which remembered information is actually trustworthy.

    1. 1
      Exactly -that distinction is basically the reason I built Qarinah. I don’t think an agent remembering “the answer” is enough. For important project state, it should ideally be able to return something closer to: “Here is the decision, here is the event it came from, here is the source/tool outcome behind it, and here is the hash you can use to verify that record.” That’s also why I’m interested in abstention. If Qarinah can’t retrieve enough evidence for a required claim, I’d rather the next agent know that the evidence is missing than receive a plausible reconstructed summary. So yes -I think trustworthy memory may end up being less about maximizing recall and more about provenance, evidence coverage, and knowing when not to claim something.