A lot of the conversation around coding agents is still about prompts.
How do I write a better prompt?
How do I give Claude more context?
How do I make the agent follow instructions?
I think there is a more useful question:
What system should exist around the agent?
I've been working on a Software Factory model with six layers:
1. Boundary
Where can the agent work? What permissions and isolation does it have?
2. Context
What repository and organizational knowledge should it receive?
3. Skills
What reusable engineering workflows should agents be able to execute?
4. Execution
How do tasks actually run in isolated environments?
5. Verification
How do we test and check what the agent produced?
6. Delivery
How does verified work reach a human-reviewed pull request?
The important part is that these layers solve different problems.
A bigger prompt doesn't solve poor isolation.
More context doesn't solve missing verification.
A smarter model doesn't automatically give you a safe delivery process.
I'm running a free live series on Maven where I walk through how to build these layers.
The first two sessions are:
Design a Software Factory for Your Coding Agents
https://maven.com/p/d3588f/design-a-software-factory-for-your-coding-agents
Build Your First Software Factory Execution Harness
https://maven.com/p/aa77d5/build-your-first-software-factory-execution-harness
The later sessions cover automated verification and background coding agents.
If you're building with coding agents on a real codebase, I'd be interested in how you're handling isolation and verification today.
What I've noticed while building my first SaaS is that AI makes coding decisions feel much cheaper than they actually are. It's very easy to say "let's just add this" or change something because the implementation itself takes a few minutes. The problem usually appears later, when those small decisions start interacting with each other.
I'm using Django and React and AI is a big part of my workflow, but I'm trying to keep the architecture and the reasoning behind the project in my own head. I think that's probably the part I'm still learning the most - not how to get AI to write more code, but how to decide what code should exist in the first place.
Love this angle, honestly. What made you look into it in the first place?
I handle isolation fairly literally.
AI is general enough to take on many different kinds of work, but I don't use one AI as a generalist that does everything. I use as many separate AI workers as I need and give each one a specialist role.
I think of it like a construction site. You wouldn't normally ask the carpenter to also handle plumbing, electrical work, interior finishing, and independent inspection.
First, I put shared rules and boundaries at the entrance to the workspace: what the AI may do, what it must not do, which sources are CURRENT, and when it must STOP and return to a human.
Then I separate roles such as construction, analysis, documentation, and inspection into different conversations. Each AI gets one role — almost like giving each worker a badge that says "builder" or "inspector."
Isolation doesn't mean preventing all information from crossing between them.
When information needs to move between stages, I pass only what the next specialist needs as an explicit handoff.
For example, after a builder AI finishes its work, I have it produce an inspection handoff describing what it changed, which parts are important or risky, and what should be checked.
But I don't let the builder inspect and certify its own work.
I give that handoff, the original instructions, and the actual artifact to a separate inspector AI, which independently checks the real result.
So for me, isolation is not complete information separation.
It is shared rule boundaries, separated roles, and explicit handoffs between specialists.
AI can be almost anything. That's exactly why I don't make it do everything at once. I deploy it as however many specialists the job requires.
This is how I actually work with AI on real development projects.
If you're curious about what I'm building this way, feel free to take a look:
https://www.kaiaspec.com/
I think separating context from verification is the important bit.
More context can help an agent make a better decision, but it doesn't tell you whether the resulting change actually fits the project.
I've been experimenting with a similar split: let the agent reason about the task, but keep some checks outside the model entirely. Especially for things that can be expressed as facts about the repository rather than opinions about the code.