I've been building Codapult Guard for a while around a problem I kept running into with coding agents: a change can work perfectly and still make the project worse.
TypeScript passes. Tests pass. The build is green. But the new code bypasses a boundary, adds another way to reach persistence, duplicates an existing service, or changes a dependency that was supposed to stay isolated.
I didn't really want another AI reviewer for that.
Guard looks at the repository itself and keeps project memory around things like imports, AST relationships, routes, capabilities, Git history and change impact. The team can then turn the relevant observations into explicit policy and have Guard check future changes against it.
One thing I care about is the baseline. Real projects already have old violations, so failing on the entire existing codebase isn't very useful. Guard can baseline what is already there and focus the fast gate on new findings.
I just released 0.5.0 recently. The recent releases added impact-aware guardrails, explainable findings, protected policy approval, run provenance, architecture budgets, concurrency-safe state, tool adapters and run observability.
It's local-first, model-agnostic, open source and doesn't require Codapult. The core checks don't call an LLM.
GitHub: https://github.com/codapult/codapult-guard
I'm mainly interested in feedback from people actually running Claude Code, Codex, Cursor or other coding agents on non-trivial projects.
The baseline is what made our own checks usable too. We run 57 deterministic checks before every push, mostly structural: no imports across feature boundaries, no direct writes to the event tables, no raw process.env. They catch what an agent breaks on day one.
They don't catch the slow drift. By the tenth AI-written feature the names follow a slightly different pattern and tests check that something exists instead of what it does. One of our checks covers the crudest case, a test with no assertion at all. Past that it's judgment, and a rule file doesn't help once the agent has read it and drifted anyway.
Is that out of scope for Guard on purpose, or do you see a way to turn "tests check behavior" into a policy?
I'd keep the semantic part out of Guard for now. "This test actually verifies the intended behavior" is a much harder claim to make deterministically than "this module must not import that module." Guard can enforce structural contracts around tests, but deciding whether an assertion meaningfully covers the behavior probably needs a different layer. The interesting part is whether we can turn some of those recurring judgment calls into explicit, checkable contracts once a team has identified the pattern.
That's how two of ours started. "Every feature runs through the real stack at least once" used to be a line in our agent instructions. Now each feature has to be imported by at least one integration test, or the push fails. "No fake tests" became the no-assertion check the same way. Neither proves a test is good. They only stop the agent from skipping the step.
Yeah, that’s the interesting transition: not trying to formalize “is this test good?”, but formalizing the concrete behavior that tends to disappear when an agent takes shortcuts. Once a team notices the same judgment call repeatedly, it can often become a much narrower contract. I like that boundary because the rule stays honest about what it actually proves.
Love the baseline idea, since it only flags new issues and ignores old ones. Excited to try this with Cursor!
That's exactly the idea. Existing debt shouldn't make the gate useless on day one, while anything newly introduced should still be visible. Cursor is one of the integrations I'm especially interested in seeing used on real projects, so I'd be curious how it fits your workflow.
Green builds masking architecture degradation is a massive blind spot once LLM agents start touching multi-module repos. Agents are optimized for local task resolution—they satisfy the immediate test suite without any global context for structural integrity, which turns clean codebases into technical debt nightmares very quickly.
The baselining feature in 0.5.0 is key here. In production environments, enforcing strict AST boundary policies without a baseline creates immediate developer friction because legacy debt blocks unrelated PRs. Pairing AST relationship checks with Git metadata/identity boundaries—ensuring automated agent commits or external contributors don't bypass local identity isolation or commit hooks—is usually where the secondary compliance failure happens.
I put together a full protection framework covering the statutory safe harbors and git identity isolation — link in my profile if useful.
Codapult Guard at 0.5.0 while you’re still mostly getting it in front of people on real agent codebases — not a workflow change yet — is the classic post-ship quiet. Soft free text-only week for that first-non-trivial-repo install freeze. Reply with what “one real team running it” looks like this week.
Small correction: Guard is at 0.7 now. But I’m curious about the “one real team running it” part — what would you actually want to see a team do with Guard before you’d consider it proven? A real agent workflow, CI enforcement, something else?
Thanks on the correction. Proven for me would be a team that isn’t you running Guard on a real agent codebase for a few days — and being able to say whether it caught a real boundary break or clearly didn’t. If you want a quiet text-only week around that finish line, I’m around.
That would actually be useful. If you have a real agent-driven codebase you can run it against, I’d be interested in seeing what Guard catches in practice - especially real boundary violations, false positives, and things it should have caught but didn’t. A quiet week of that kind of testing would probably tell me more than a lot of synthetic examples.
A useful adoption metric might be baseline drift over time, not only the count of new findings. If teams can see whether violations are being retired and whether policy changes reduce repeat findings, the guard becomes part of architecture maintenance instead of another CI gate. A small trend report could also show when a rule is too noisy to keep enabled.
Yeah, baseline drift is probably a more useful signal than the raw finding count. A growing baseline can mean the team is accumulating debt, but a shrinking one shows that old violations are actually being retired. I also like the idea of tracking rules that repeatedly generate findings without getting fixed - that’s useful feedback on whether the policy is too noisy or the boundary itself needs to be reconsidered.
The insight here is that you're not just building a guard - you're building observability into what makes an agent reliable. Most people ship agent code and only learn what breaks when users hit it. By baking the architecture guard into the system from day one, you're catching constraint violations before they become production incidents. That's the difference between 'code that works' and 'code you can reason about.'
That’s the distinction I’m aiming for. A passing test suite tells you the observed behavior is still okay; it doesn’t necessarily tell you whether the change still fits the system. The useful part is catching that mismatch while the change is still in the agent’s workspace, when fixing the architecture is much cheaper than discovering it through production behavior.
Run provenance sounds especially important once agents can touch both implementation and configuration. Does it capture the exact policy snapshot alongside the repo commit, so a later audit can distinguish “the code respected the rule” from “the rule changed”? That seems like a useful trust signal for CI.
Yes, that’s exactly the distinction provenance is meant to preserve. The run records the policy/baseline context used for the decision alongside the repository state, so a later audit can see what was actually evaluated rather than just looking at the current policy. That’s particularly useful when an agent can modify both code and the rules governing that code.
Yes, that’s exactly the distinction provenance is meant to preserve. The run records the policy/baseline context used for the decision alongside the repository state, so a later audit can see what was actually evaluated rather than just looking at the current policy. That’s particularly useful when an agent can modify both code and the rules governing that code.
That's a huge pain point for any AI coding agent user working on a substantial codebase. The moment they get a green build, it's hard to be confident that nothing was accidentally broken, especially with regard to architectural integrity - the agent could've ignored any boundaries or duplicated services without suspicion, but tests still passed.
Exactly. A green build mostly tells you that the behavior you tested still works. It doesn't tell you that the change respected the architecture around that behavior. That's the gap I'm interested in: checking the change against the project's existing boundaries and explicit rules, rather than asking another model to review the code and hope it notices the same architectural problem.
The baseline is the part I'd trust most here — a gate that fails on day-one code gets switched off in a week. What goes into a fingerprint? If it includes line numbers or file paths, a mechanical rename or reformat resurfaces old violations, and one false alarm wave is usually enough for a team to stop listening to the gate.
That's a real tradeoff in the current implementation. The fingerprint doesn't include line numbers, so reformatting or moving code around within a file shouldn't resurrect an existing finding. It does include the rule ID and file path, plus the import specifier and, where relevant, the resolved path. Contract, budget and cycle findings also use paths.
That means a mechanical file rename or module move can produce a new fingerprint and bring the finding back. That's intentional to avoid one baseline entry suppressing the same violation everywhere, but rename-heavy refactors are definitely an area where the baseline could be smarter.
For a modular SaaS, blanket 'no cross-module imports' checks need a few explicit orchestration exceptions, such as a successful payment granting credits. I'd want those exceptions scoped to the exact caller and callee rather than an entire directory, so an agent cannot broaden them accidentally. Can a contract express an allowed service-to-service edge while still rejecting a second persistence path from the route layer?
Yes. That’s the kind of narrow exception the contract model is intended to express. You can scope the allowed edge to the relevant entrypoints/packages rather than opening up an entire directory. A separate rule can still reject the route-to-persistence path, so allowing payments -> credits doesn’t implicitly make routes -> persistence valid.
The important part is that the exception is part of the explicit contract, rather than something the agent can infer and broaden.
I like that it's not another AI reviewer - the baseline is the real problem. We hit this when we gave our agent write access: instead of failing on every existing violation, we kept a policy file and only blocked new ones, read-only by default and write behind a flag. Does Guard scope policy per boundary, or is it repo-wide?
Both. The policy is stored at the repository level, but enforcement can be scoped to specific boundaries.
For example, rules can target files, contracts can use
scopeandentrypoints, package boundaries can usefromPackages/mustNotImportPackages, and budgets can be scoped to particular modules or routes.The baseline is also repo-managed and fingerprint-based: it suppresses findings that already existed, while a new
errorfinding still fails the gate. Normal checks are read-only with respect to policy; changing rules, contracts, baselines, or waivers requires an explicit CLI/MCP operation, and protected mode can prevent an agent from approving its own policy change.The policy file is part of the attack surface. If the same agent can change both the code and Guard’s rules in one PR, a green result might mean it moved the boundary instead of respecting it. Does protected policy approval prevent the agent from changing the rule and implementation together?
Yes - that's exactly what protected approval is for. In protected mode the agent can propose a policy change but can't approve it, and with
requireDistinctActorthe approver has to be different from the proposer.So code and policy can still appear in the same PR, but a green verification result doesn't make the policy change self-approved.
The hard part here isn't catching violations - it's distinguishing between "broke a boundary we agreed on" and "exposed a boundary we didn't know we had". Guard's deterministic checks handle the first. But baselining requires you to first decide what counts as legacy vs. what is actually an architectural constraint. That decision is a measurement problem: you have to look at the repository history and ask "was this always wrong, or did we change our minds about what the boundary means?" A gate that fails on day-one code gets disabled, but a gate that misses subtle violations silently becomes worse than having no gate at all.
The distinction you draw between "was this already here" and "do we still consider this acceptable" is the right cut, and it has a consequence that is easy to miss: the second question has no stable answer, so it cannot be stored the way the first one can.
Baseline is an empirical claim about the past. A fingerprint of existing findings is either in the recorded history or it is not, and that answer does not change next month, because the past does not change. Policy is a claim about what the team currently believes, and that belief moves. When a boundary is redefined, every prior violation that the new opinion now condemns looks identical in the data to a violation that has always been wrong. Repository history can tell you it is old. It cannot tell you which opinion graded it.
That collapses into a real mechanism problem at the moment policy changes, because a single edit to a rule can reclassify hundreds of findings in one commit. If the baseline is keyed on finding identity, a boundary redefinition will look like a wave of new violations and the gate will fail on a change that only touched a rules file. That is the failure that gets gates switched off, and it arrives exactly when the architecture is being deliberately reorganised rather than incidentally broken. What the baseline needs to record is not just that a finding was pre-existing but which policy version it was pre-existing under, so that a redefinition is distinguished from a regression.
The measurement question underneath is harder than the corpus it runs on. Every finding is generated by the rules you wrote, so the boundary set the tool checks is a subset of the boundaries someone already thought to formalise. A violation of a boundary nobody has written down yet is invisible by construction, and it is the one most likely to matter, because the whole point of the post was that the change passed while making the project worse.
On the interest side, this is the second time I am commenting on this thread, so the same note applies as before: I build Piramyd, a gateway for coding agents, and a deterministic gate that keeps checks out of the context window is the direction I would argue for anyway.
When a rule is redefined, should the reclassification land as a policy diff a human approves, or should the gate keep old findings suppressed until someone consciously re-baselines?
Yes, I think expiry should block the run rather than just warn when the underlying finding is an error. In Guard, an expired time-bounded waiver stops suppressing the finding automatically, so it becomes active again. Extending it is an explicit policy operation, not something the blocked run can silently do. I also agree that the reason matters - legacy debt and a temporary exception shouldn’t have the same lifecycle. And tracking the direction of waiver growth is useful too; a steadily growing exception set is a signal that the policy itself may need attention.
Exactly. That's also why I keep baseline and policy separate in Guard. Baseline answers "was this already here?", while policy answers "do we still consider this acceptable?" Git history can help with the first, but it shouldn't decide the second.
Most founders building alongside a day job focus entirely on non-compete clauses, but corporate legal teams usually catch engineers through environment configuration leaks, not contract disputes:
I put together a step-by-step air-gap checklist and directory config template on my Medium profile (https://medium.com/@VB_Builds) if you want to inspect your own setup.
The detail that matters most here is buried in your last paragraph: the core checks do not call an LLM.
That is not just a latency or cost footnote, it is the reason this class of check is trustworthy. An AI reviewer asked whether a change respects a boundary gives you a probabilistic answer at a per-token price, and the answer drifts when the model version changes. A deterministic check gives you the same verdict today as in six months, for nothing per run. Anything you can express as a rule should not be rented back from a model. Save the generation budget for the parts that genuinely need judgment.
The baseline decision is the other half. A gate that fails on day-one legacy debt gets disabled within a week, and then you have no gate at all. Baselining old violations while failing hard on new ones is what makes the difference between a tool that is kept and a tool that is trialled.
One thing I would watch as you add adapters: the moment Guard starts summarising repository state into a prompt for any advisory step, it joins the same context budget it was meant to protect. Keeping the gate out of the context window is a feature worth defending explicitly.
Full disclosure, I am the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I benefit when people stop optimising token counts, so discount my enthusiasm for the deterministic approach accordingly.
Does the guard run at Stop on the whole diff, or can it check incrementally as files change?
Guard supports two deterministic modes. The default
checkcan verify the full project, whilecheck --changedlimits findings to changed files. Thereviewcommand is change-scoped by default and adds a bounded, redacted diff, project context, contracts, deterministic findings, and architectural impact.So it can run after an agent finishes a task, or be used in a change-scoped workflow. Guard itself doesn't call an LLM or decide when a host should invoke it; automatic Stop/completion integration is host-specific.
I'm deliberately keeping the deterministic gate outside the LLM context path. The MCP
reviewtool prepares a bounded packet for an external AI or human reviewer when semantic judgment is useful. Core pass/fail checks don't require a model.Existing findings are handled through the baseline, while new
errorfindings fail the run. Warnings remain advisory.Guarding architecture before the agent writes more code is the right layer. The failure mode I see is soft warnings that still let the PR land, so I would make new boundary breaks fail the run while baselining old debt. Do you fail the run hard on a rule break, or only leave a report for a human to ignore later?
The gap between an error-level rule that blocks and a warning that informs is where most of these gates quietly die, and it is worth being explicit about why.
An advisory warning is a request for attention from the person who is already the most time-constrained actor in the pipeline: whoever is trying to land the change. Anything that only informs has to compete with the entire rest of the review surface, and it loses. Not because anyone decided the finding did not matter, but because there is no moment at which someone is forced to decide. The interesting property of a hard failure is not severity, it is that it creates a decision point. Without one, the disposition of every warning is decided implicitly, and the implicit answer is always "later".
That is why I would be suspicious of a rule set where the error tier is small and the warning tier is where the real architecture beliefs live. If the team would revert a change over it, it is an error. If they would not, it does not need to be a warning at all, it belongs in a report nobody is blocked by, reviewed on a schedule rather than at submission time. The useful split is not severity, it is whether a human is being asked to act at this moment or whether this is material for a conversation that happens later.
The other thing this depends on is which side of the boundary the check sits on. A gate that runs at the point where the agent finishes is enforcing a rule against a diff that already exists, which is fine, but it is the weaker position. A check that a host can invoke while the agent is still working is enforcing the same rule against a change that has not been written yet, and the cost of compliance at that point is near zero instead of being a rewrite. Same rule, same determinism, very different economics for the person on the other end of the failure.
Same disclosure as earlier in this thread, since it is now my second comment here: I build Piramyd, a gateway for coding agents.
For the warning tier specifically, do you record anywhere whether a warning was ever resolved, or does it just disappear from the report the next time it is clean?
Guard fails hard on new violations when they're covered by an error-level rule or contract. Existing findings can be baselined, while warnings stay advisory. For architecture boundaries I'd normally use error; warnings are more for cases where the result needs a human look rather than an automatic failure.
Baselining existing violations is what makes this usable; a gate that fails on day-one legacy code gets switched off within a week. On tool adapters: Meta's Muse Code has lifecycle hooks (SessionStart, PreToolUse, PermissionRequest, Stop and more) registered under ~/.config/muse/hooks/, so Guard's fast gate could run at Stop, before a change is handed back. This build shows a third-party tool wiring into them without a fork: https://shipwithmuse.live/builds/herdr-muse-lifecycle-hooks (I help curate it)
Yeah, Stop is probably the cleanest place for that kind of check. You want the agent to finish its work first, but still catch the obvious stuff before the result gets handed back. I hadn't looked closely at Muse's hooks yet, but that integration point makes a lot of sense.
I built Chatform.in - conversational forms that people actually finish ( launching on product hunt this Thursday. https://www.producthunt.com/products/chatform-3?utm_source=twitter&utm_medium=social )
And recently shipped one more product Please drop an upvote or review, it will be really helpful https://www.producthunt.com/products/shipwithmuse?utm_source=other&utm_medium=social
Have teams actually changed their agent workflow after seeing Guard catch a new architectural violation, or is adoption still mainly driven by interest in the idea?
Not really yet. It's still early and I'm mostly trying to get Guard in front of people actually using coding agents on real projects. The workflow change is the part I'm most interested in seeing once people start using it on larger codebases.
The larger-codebase test is probably where the real signal shows up. Could be useful to compare notes as that develops by email sometime.
Yeah, definitely. Once I have a few larger codebases running through it, I'll have something more concrete to compare than the current early results.
That larger-codebase test should give you much better evidence than the current sample. If you’re open to it, what’s the best email to reach you on?
Sure, you can reach me at vlad [at] codapult [dot] dev.
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.
The baseline-plus-new-violations approach makes this practical for real codebases. One addition I’d love is time-bounded exceptions with an owner and reason, so legacy debt is tolerated without quietly becoming permanent architecture policy.
Yeah, I agree with this. A baseline shouldn't quietly become a permanent exception list. Having an owner, reason and some kind of expiry would make it much harder to forget why an exception exists in the first place.
The mechanism that makes this work or makes it theatre is who has to act when the date arrives.
An exception with an owner and a reason is still only a row in a file. The decay happens on the day it lapses, because the cheapest move available is to push the date out, and the person doing that is the person whose change is currently blocked. That is not a discipline failure, it is an incentive that resolves the same way every time. Inverting the default is the only version that holds: a lapsed exception re-enables the finding it was silencing, with no action required from anyone. Extending it should cost a deliberate, visible decision instead.
The reason field carries more weight than it appears to. "This predates the boundary" and "we accept this for now and will fix it next quarter" look identical in a table, yet they are opposite claims, and only one of them should come back to life when it expires. Free text means nobody re-reads it, and the list becomes permanent by neglect. A small closed set of reasons, each tied to an expected kind of follow-up, makes a lapse mean something specific rather than just a date passing.
The number I would flag to a human is not any single finding but the direction of the count. Exceptions that only accumulate are not a baseline at all, they are an architecture decision being taken by attrition, one blocked pull request at a time.
Same disclosure I owe earlier in this thread: I build Piramyd, a gateway for coding agents, so I gain commercially when more agent work ships against real repositories.
When an exception lapses with no decision made, should it block the run or surface as a warning inside the report?
If the underlying finding is an error, an expired waiver should block the run. That's how Guard handles it now: once a time-bounded waiver expires, it stops suppressing the finding, so the original error becomes active again. Extending the waiver is an explicit policy operation rather than something the blocked run can do implicitly. For an underlying warning, it remains advisory. The expiry itself shouldn't weaken the severity of the finding it was temporarily suppressing.
Agreed—giving exceptions an owner and an expiry should keep the baseline from quietly eroding.