I've been building Codapult Guard for a while around a problem I kept running into with coding agents: a change can work perfectly and still make the project worse.
TypeScript passes. Tests pass. The build is green. But the new code bypasses a boundary, adds another way to reach persistence, duplicates an existing service, or changes a dependency that was supposed to stay isolated.
I didn't really want another AI reviewer for that.
Guard looks at the repository itself and keeps project memory around things like imports, AST relationships, routes, capabilities, Git history and change impact. The team can then turn the relevant observations into explicit policy and have Guard check future changes against it.
One thing I care about is the baseline. Real projects already have old violations, so failing on the entire existing codebase isn't very useful. Guard can baseline what is already there and focus the fast gate on new findings.
I just released 0.5.0 recently. The recent releases added impact-aware guardrails, explainable findings, protected policy approval, run provenance, architecture budgets, concurrency-safe state, tool adapters and run observability.
It's local-first, model-agnostic, open source and doesn't require Codapult. The core checks don't call an LLM.
GitHub: https://github.com/codapult/codapult-guard
I'm mainly interested in feedback from people actually running Claude Code, Codex, Cursor or other coding agents on non-trivial projects.
The detail that matters most here is buried in your last paragraph: the core checks do not call an LLM.
That is not just a latency or cost footnote, it is the reason this class of check is trustworthy. An AI reviewer asked whether a change respects a boundary gives you a probabilistic answer at a per-token price, and the answer drifts when the model version changes. A deterministic check gives you the same verdict today as in six months, for nothing per run. Anything you can express as a rule should not be rented back from a model. Save the generation budget for the parts that genuinely need judgment.
The baseline decision is the other half. A gate that fails on day-one legacy debt gets disabled within a week, and then you have no gate at all. Baselining old violations while failing hard on new ones is what makes the difference between a tool that is kept and a tool that is trialled.
One thing I would watch as you add adapters: the moment Guard starts summarising repository state into a prompt for any advisory step, it joins the same context budget it was meant to protect. Keeping the gate out of the context window is a feature worth defending explicitly.
Full disclosure, I am the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I benefit when people stop optimising token counts, so discount my enthusiasm for the deterministic approach accordingly.
Does the guard run at Stop on the whole diff, or can it check incrementally as files change?
Guarding architecture before the agent writes more code is the right layer. The failure mode I see is soft warnings that still let the PR land, so I would make new boundary breaks fail the run while baselining old debt. Do you fail the run hard on a rule break, or only leave a report for a human to ignore later?
Guard fails hard on new violations when they're covered by an error-level rule or contract. Existing findings can be baselined, while warnings stay advisory. For architecture boundaries I'd normally use error; warnings are more for cases where the result needs a human look rather than an automatic failure.
Baselining existing violations is what makes this usable; a gate that fails on day-one legacy code gets switched off within a week. On tool adapters: Meta's Muse Code has lifecycle hooks (SessionStart, PreToolUse, PermissionRequest, Stop and more) registered under ~/.config/muse/hooks/, so Guard's fast gate could run at Stop, before a change is handed back. This build shows a third-party tool wiring into them without a fork: https://shipwithmuse.live/builds/herdr-muse-lifecycle-hooks (I help curate it)
Yeah, Stop is probably the cleanest place for that kind of check. You want the agent to finish its work first, but still catch the obvious stuff before the result gets handed back. I hadn't looked closely at Muse's hooks yet, but that integration point makes a lot of sense.
I built Chatform.in - conversational forms that people actually finish ( launching on product hunt this Thursday. https://www.producthunt.com/products/chatform-3?utm_source=twitter&utm_medium=social )
And recently shipped one more product Please drop an upvote or review, it will be really helpful https://www.producthunt.com/products/shipwithmuse?utm_source=other&utm_medium=social
Have teams actually changed their agent workflow after seeing Guard catch a new architectural violation, or is adoption still mainly driven by interest in the idea?
Not really yet. It's still early and I'm mostly trying to get Guard in front of people actually using coding agents on real projects. The workflow change is the part I'm most interested in seeing once people start using it on larger codebases.
The larger-codebase test is probably where the real signal shows up. Could be useful to compare notes as that develops by email sometime.
Yeah, definitely. Once I have a few larger codebases running through it, I'll have something more concrete to compare than the current early results.
That larger-codebase test should give you much better evidence than the current sample. If you’re open to it, what’s the best email to reach you on?
The baseline-plus-new-violations approach makes this practical for real codebases. One addition I’d love is time-bounded exceptions with an owner and reason, so legacy debt is tolerated without quietly becoming permanent architecture policy.
Yeah, I agree with this. A baseline shouldn't quietly become a permanent exception list. Having an owner, reason and some kind of expiry would make it much harder to forget why an exception exists in the first place.