Sometimes ChatGPT stops working for me.
Not because it can't continue.
Not because it ran out of permission.
It can have enough information, enough authority, and a perfectly available next action — and still come back with something like:
Something is wrong. I think we should stop here.
That has become one of the most useful things ChatGPT can do in my work.
It also wasn't always like this.
I've spent more than 5,000 hours working with ChatGPT over time. I didn't spend those hours developing a theory of human-AI collaboration. I wasn't trying to train a model or invent an agent framework.
Mostly, we were making things.
At first, a lot of them were strange things.
We would try something, break something, talk about what happened, change the way we worked, and try again.
As the work became more complicated, something else gradually became more complicated too:
the way we worked together.
Only later did I realize that many of the problems people now discuss under terms like human-AI collaboration, agent boundaries, human approval, and reliable AI behavior were problems we had already been stumbling through in everyday work.
And some of the solutions we ended up with are surprisingly simple.
One of my basic rules is this:
Make the boundary clear — and then respect it yourself.
If I tell ChatGPT:
You can decide this.
I don't wait until I dislike its decision and suddenly say:
Why did you decide that without asking me?
If I say:
This decision belongs to the Human.
I don't throw it back at ChatGPT later because making the decision became inconvenient.
If stopping is allowed, I don't punish it for stopping.
If asking me is allowed, I don't treat the question as a failure.
And if I expect the AI to be honest with me, I have to be honest with it too.
That sounds almost too simple.
But over a long enough working relationship, it changes what becomes cheap — and what becomes expensive.
In my working environment, all of these are valid outputs:
I don't know.
I can't do that.
This is harder than it looks.
I don't have the evidence for that.
Something is wrong.
I think we should stop here.
This decision belongs to you.
None of those automatically means the AI failed.
Sometimes they are exactly the right result.
If ChatGPT pretends something exists when it doesn't, we now have to build the next step on a false premise.
If it pretends a task is finished, the false completion has to survive the next inspection.
If it silently guesses what I wanted, the guess may propagate into five more decisions.
Hiding the problem doesn't remove the problem.
It gives us two problems: the original one, and the false reality we have now built around it.
But if ChatGPT says:
I don't know. Can you decide this?
the problem may take thirty seconds to resolve.
So in this environment:
Pretending is expensive. Asking is cheap.
And if you're thinking:
“My ChatGPT doesn't work like this.”
That's part of the story.
Mine didn't always work like this either.
One of my early problems with ChatGPT was much more basic.
Keeping one complicated idea intact across a long conversation was difficult.
We could be working on A.
Then I would add a new condition.
Instead of continuing with the accumulated A plus the new condition, ChatGPT might effectively produce a new version of A.
So I made a crude scaffold.
I called it KEEP.
The basic idea was something like:
KEEP 1 = the thing we have already established
Then I would add numbered conditions, changes, or exceptions while repeatedly telling ChatGPT to keep referring back to KEEP 1.
It wasn't a product.
It wasn't an AI technique I had learned somewhere.
It was something we made because our conversation kept falling apart.
At the time, ChatGPT described the structure as something it had begun to hold onto and use as a basis for continuing the conversation.
I cannot inspect what that meant internally.
I can't tell you what persisted, where it persisted, or by what mechanism later behavior may or may not be related to those conversations.
What I can tell you is what I observed from the outside.
Over time, I found that I needed the scaffold less and less in my conversations with ChatGPT.
Eventually, I stopped writing KEEP altogether.
That experience changed the way I thought about working with AI.
A lot of AI control looks like this:
Don't do A.
Don't do B.
Ask before C.
Never change D.
If E happens, stop.
Those rules can be useful. We use explicit boundaries too.
But there is a problem.
You can build 1,000 guardrails.
Then the AI encounters hole number 1,001.
If all it knows is the list, the new case isn't on the list.
So I became much more interested in something one level above the rules:
Why does the boundary exist?
What is this job actually trying to accomplish?
What is the Human responsible for?
What is the AI responsible for?
What evidence is required before acting?
Why is this particular decision not the AI's decision to make?
Once those things are understood well enough, an unfamiliar situation can produce a different question:
I can do this. But should I?
And that leads to one of the most important distinctions in the way I work with ChatGPT:
Capability is not authority.
And even authority does not require action.
Those are three different questions.
Can the AI do it?
Is the AI allowed to do it?
And even if both answers are yes:
Should it actually do it?
I can give ChatGPT considerable room to act while still expecting it to decide that sometimes the correct use of that freedom is:
not to use it.
Stopping can itself be a decision.
In August 2026, I had a much more ordinary idea.
I wanted a casual way to give crowdfunding supporters a small digital thank-you keepsake.
Nothing about this began as:
Let's conduct an experiment in human-AI software development.
We started talking about what the keepsake would need.
A number.
A timestamp.
An individual identity where necessary.
A way to issue it.
Somewhere in that process I had a realization roughly equivalent to:
Wait. We can make WordPress plugins?
So we did.
That work eventually became Kaia Memoria.
As the project grew, our way of working began to break again.
One long ChatGPT conversation could handle implementation details, but as the context became larger, keeping the entire direction of the project intact became harder.
So we split the work.
One ChatGPT conversation kept track of the larger direction and previous decisions.
Another focused deeply on implementation.
Another inspected what had been built.
Sometimes a separate analysis conversation compared versions or investigated a narrow problem.
We didn't sit down one morning and design a sophisticated “multi-agent architecture.”
The organization emerged because the work kept producing problems that needed different kinds of attention.
The Human role remained important.
I defined product intent and boundaries, made decisions that belonged to the Human, installed and ran actual builds, observed the real UI and runtime behavior, and decided whether the result was acceptable.
The AI side investigated source, traced dependencies, explained structures, implemented revisions, compared known-good versions, prepared inspections, and handed findings between roles.
I didn't directly edit the source code.
I didn't even open the development ZIPs.
That wasn't necessary for my role.
I was not the code reviewer.
I was the product and runtime authority.
The AI could tell me what it believed had been built.
But the real build still had to run.
The real button still had to work.
The real output still had to appear.
And I had to decide whether what happened was actually what we intended.
When something looked wrong, completion was not automatically the next goal.
Sometimes the next correct action was simply:
STOP.
Investigate.
Bring the evidence back.
Decide what happens next.
Then make the smallest justified change and test again.
That is how a strange little thank-you idea turned into working software.
This is where I need to be careful.
It would be very easy to tell a much more dramatic story.
I could say that I “trained ChatGPT.”
I could point at behaviors that changed over time and invent a technical explanation for why.
I don't think that would be honest.
I work with ChatGPT from the outside.
I can observe behavior.
I can change the working environment.
I can introduce scaffolding.
I can make boundaries explicit.
I can see what happens when those boundaries are respected repeatedly.
I can see when an old scaffold becomes unnecessary in my work.
I can compare how ChatGPT behaves with me now to how it behaved with me before.
But there is a box in the middle that I cannot see.
So my claim is deliberately narrower:
This is what I observed. I don't know the exact mechanism behind it.
And, appropriately enough, being able to say I don't know is part of the whole point.
I started thinking about the environment around ChatGPT almost like a race circuit.
You already have an extraordinarily capable car.
Most businesses don't need to rebuild the car.
They need a place where it knows what job it is doing, where it can drive, where it must stop, what information counts as evidence, and which decisions belong to a Human.
That's what led me toward the idea behind Kaia Spec.
We make practical work manuals for ChatGPT.
Not simply giant prompts full of rules.
A useful manual needs to communicate the job:
the purpose,
the context,
the available information,
the authority,
the boundaries,
the reasons behind those boundaries,
and the points where the Human takes over.
In other words:
Don't just give ChatGPT 1,000 guardrails. Give it enough understanding to notice hole 1,001.
We aren't trying to build another general-purpose AI.
ChatGPT already has many of the general capabilities we need.
What we can build is the workplace around it.
Give ChatGPT a job.
I don't speak English.
I'm writing this with ChatGPT.
But “writing with ChatGPT” doesn't mean I gave it a prompt saying:
Write me an Indie Hackers post about human-AI collaboration.
We developed the article through conversation.
I work in Japanese.
I decide what I mean, what actually happened, what matters, where the boundaries are, and when the wording has drifted away from the idea.
ChatGPT takes that material and writes for an English-speaking audience.
When it misunderstands me, I correct it.
When it sees a structural problem, it tells me.
When something cannot be supported, we leave it out or say that we don't know.
So while you've been reading an article about human-AI collaboration, you've also been reading the output of one.
And Kaia Memoria is another.
We built it together.
Kaia Memoria is the working software project that grew out of this collaboration.
If you'd like to see what we actually built, you can find Kaia Memoria on my Indie Hackers profile and try the live demo from there.
The practical test I use for a stop is not whether the model sounds uncertain. It is whether it can name the missing evidence, the decision owner, and the smallest next check. That turns “I think we should stop” into a resumable handoff. Logging those stop reasons also reveals which parts of the workflow are underspecified, so the evidence loop improves the system instead of becoming another blanket guardrail.
The practical test I use for a stop is not whether the model sounds uncertain. It is whether it can name the missing evidence, the decision owner, and the smallest next check. That turns “I think we should stop” into a resumable handoff. Logging those stop reasons also reveals which parts of the workflow are underspecified, so the evidence loop improves the system instead of becoming another blanket guardrail.
Pretending is expensive, asking is cheap - this principle inverts how most teams handle uncertainty. Systems usually hide errors to appear confident. Your point about ChatGPT saying 'I don't know' reveals the actual boundary instead of a false confident facade.
This matches what I've seen building an AI assistant for small trade businesses. The model almost never needed to be smarter. What mattered was where it stops.
We ended up splitting everything into two kinds of actions: prepare and commit. It can draft a quote or an invoice as many times as it likes, but anything that leaves the company or can't be undone stops, shows the person exactly what will happen, and waits.
That's also why it shows its calculation, not just a result. My bet is that a tradesperson will happily fix one wrong line in a quote, but won't use a tool that sends something they didn't see.
This really resonates with how I use AI for SEO and content work. I’ve found that the biggest problems usually don’t happen when AI can’t do something, but when it confidently continues with a wrong assumption. I’d much rather have it stop and ask for clarification than build five more steps on top of something incorrect.
I especially like the distinction between capability, authority, and action. Curious though — how do you decide which boundaries should be hard-coded rules and which ones should be left to the AI’s judgment?
This question actually became a big part of the follow-up I wrote. 😄
The short version is:
Known risks → guardrails.
Unknown situations → judgment.
If I already understand a boundary well enough to encode it structurally, I don’t want the AI spending judgment on that same decision every time.
I want to preserve its judgment for the case nobody thought to encode.
I went into the full reasoning here:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I originally wanted this to be my second Indie Hackers article, but I still don’t have enough experience here to publish another post yet, so I put it on the official Kaia Lab for now.
The “make the boundary clear—and then respect it yourself” rule is the operational piece most teams miss. I’d turn it into a small preflight contract—what can be decided autonomously, what needs a human checkpoint, and the exact stop/escalate condition—then test it against edge cases so “stop” is observable rather than stylistic.
The capability / authority / necessity split matches what we saw with an AI agent inside one of our apps. Being able to do something and being allowed to do it without asking needed separate settings.
For most actions the user can click "always allow". Two actions never get that option: a permanent delete, and an edit that immediately rewrites the instructions of another AI feature. The agent can suggest those, and a person confirms each one. We put that limit into the app itself instead of hoping the model knows when to stop.
Yes — that's very close to how I think about it. 😄
For me, Human confirmation isn't only a safety mechanism. It's also a way to simplify the system.
If adding one confirmation button can eliminate both an accident path and an entire tree of automated judgment, edge cases, and potential bugs, that button is very cheap.
If I already know a boundary is dangerous, I don't really want the model spending intelligence deciding whether it should cross it every time. I'd rather put that boundary into the structure itself.
Then I can reserve the model's judgment for the cases I didn't anticipate — the "1001st hole."
I tend to think of those fixed boundaries like toll gates on a highway. You don't need the driver to reconsider the rules every time they reach one. The gate is simply there because we already know that this is a decision point that requires confirmation.
So I really like your example of permanent deletion and instruction rewriting being enforced by the app itself. That's exactly the kind of place where I'd rather make the road simpler than ask the AI to be smarter.
The capability / authority / "should it actually act" split is the part most agent frameworks collapse into a single permission check, and that collapse is exactly why hole 1,001 keeps showing up. A permission list encodes what is allowed but never why, so when the model hits an unanticipated case it has nothing to reason from and defaults to the only signal it has: completion.
I think your KEEP scaffold is interesting for a slightly different reason than the memory question. What it really did was make the invariants explicit and external to the conversation — and your later split-role setup is the same move at a larger scale, with the direction-keeping conversation being KEEP 1 with a job title. That framing makes the lesson portable: the durable gain may be less about the model adapting to you and more about you getting better at stating intent in a form that survives context pressure, which transfers across models and doesn't require any claim about what happens inside the box.
The place I'd push back is on the cost of stopping. "Halt, investigate, bring evidence" only stays cheap while the human is genuinely available and cheap to interrupt, and in most businesses the human is the bottleneck. An assistant that correctly halts fifteen times a day gets overridden into silence within a week, which is how good stop-behavior gets trained back out in practice.
So: do the Kaia Spec manuals define severity tiers for stopping — i.e. this class of uncertainty escalates to a human, that class gets a documented default plus a logged flag — or is escalation currently binary? Curious whether you've found a way to make "stop" scale without making the human the rate limiter.
Building Genie 007 (genie007.com) I ran into this same thing from the other side - teaching an AI when to stop is harder than teaching it what to do next. The failure mode is always overconfidence: the model has a completion signal and no uncertainty signal, so it keeps going.
The thing that changed it for us: adding an explicit "check before continuing" step in the workflow design, separate from the action step. Once the model had a dedicated moment to assess rather than just act, the false-positive completions dropped significantly.
What prompting pattern do you use to trigger the self-check?
The "stop even though you could continue" behavior is underrated. I want the same thing from research assistants: refuse to call a competitor move "real" when the evidence is thin.
What worked for me was an explicit stop rule — if I don't have a dated source URL and a before/after, the answer is "insufficient," not a confident summary. Smart enough to continue is common; disciplined enough to halt is rarer.
How are you encoding those stop conditions today — prompt heuristics, or something the user can audit after the fact?
This resonates a lot. I'm running an agent (OpenClaw) that has to ask before it spends money, publishes anything, or creates accounts -- and the biggest shift wasn't the rules themselves, it was making the boundary bidirectional like you describe. Early on I'd get annoyed when it asked for approval on something "obvious" and want to loosen the rule after the fact. But every time the rule stayed fixed, the agent's "I don't have enough to do this safely" calls got more accurate, not less. The expensive failure mode for me wasn't the agent asking too much -- it was letting its own context/state get so large it started acting on stale assumptions instead of re-checking real state. Cheap to ask, expensive to pretend, exactly as you say.
I like the idea of asking why a boundary exists instead of stacking another rule on top. That’s much closer to how I think about agent workflows.
Once you ship enough AI tools, you realize knowing when not to act is half the system.
Yes — that’s very close to how I think about it. 😄
Rather than adding another rule every time a new problem appears, I usually want to ask first:
“Why does this boundary exist?”
“What is this boundary protecting?”
If the reason is clear, one structural boundary can sometimes solve several problems at once.
If I keep adding rules without understanding the reason behind them, they can start conflicting with each other or quietly remove too much of the AI’s useful judgment.
So I agree that:
what the AI should do
is only half of the system.
The other half is:
what it should not do,
when it should stop,
and when the decision should return to the Human.
The useful distinction here is that stopping isn't the opposite of progress. It's a way of protecting the meaning of the work from a plausible-looking next step. I've found the hard part isn't writing a stop rule, but noticing when a project has quietly changed its question and I'm still optimizing the old answer. A small written boundary or decision record gives you something to return to when momentum starts impersonating evidence.
I think being able to choose “stop” is itself a form of progress.
I sometimes think of it like a dog.
There is a big difference between a dog that simply runs the moment the leash comes off, and one that can respond to different situations:
GO and run,
find something and bring it back,
deal with it there,
stop and call the Human.
I think AI is similar.
Continue.
Stop.
Nothing found.
Unknown.
If all of those can exist as valid outcomes, that seems more intelligent to me than simply continuing whenever continuation is technically possible.
I wrote about a closely related idea on the official Kaia Lab as well, in case it gives you any useful ideas:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
"If stopping is allowed, I don't punish it for stopping" is the line I'd keep. Most people set a boundary and then get annoyed when the model actually uses it, and the behavior flips back within a session or two. The question I'd add: how do you keep that consistent when more than one person works with the same assistant? Boundaries agreed by one operator tend to get quietly redefined by the next.
If a boundary changes depending on which Human happens to be present, I don’t think it is really functioning as a boundary anymore.
I think of a boundary more like a KEEP OUT rope.
It stays in the same place no matter who walks up to it.
And for me, a boundary is not only:
“This is how far you may go.”
It also includes:
“Do not expand the job through extra helpfulness beyond the agreed scope.”
Inside the boundary, the AI can move freely.
Outside it, it should not quietly enlarge the task on its own.
So if several Humans are working with the same assistant, I would want the shared boundary to be explicit and stable, separate from each person’s individual preferences.
I wrote about a closely related idea here as well:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
Agreed, and with several agents working together the failure gets worse: when one step stops or times out, the next one often carries on as if it had an answer. We had exactly that bug. A step that ran out of time was passed to the next phase as though it had replied, and the agent downstream confidently built on nothing. Making "I stopped, and here's why" a real result that the next step can see changed more than any prompt tweak did.
If an AI stops, I want it to leave a clear footprint that it stopped.
And I also want it to record why it stopped.
I would not leave that field blank.
One behavior I’ve seen repeatedly is that a blank field can look unfinished to an AI, which creates pressure to fill it with something.
So I prefer something explicit, for example:
STOPPED
REASON: UNKNOWN
That way, the stop itself survives as information.
The next agent can tell the difference between:
“nothing was found”
and
“the previous agent deliberately stopped here.”
I wrote about a closely related idea on the official Kaia Lab too:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
The distinction between observed behavior and an explanation of the model internals is useful. For a structured decision workflow, I would keep the input, returned decision, and validation result as separate fields, with an explicit unknown state when evidence is missing. That lets a later review challenge the explanation without losing the original observation. When you compare a clean session with your long-running setup, will you also track false stops as well as unsafe continuations?
Yes — I want all of those states recorded.
If something was found: say so.
If nothing was found: say so.
If it is unknown: say that too.
I treat all of them as valid sensor outputs.
The “nothing” case matters especially.
If a field is simply left blank, an AI may interpret that as unfinished and feel pressure to fill it with something.
So I prefer an explicit completed state, for example:
SIGNAL: NO
“No signal” is still a completed observation.
I want to record both:
continued when it should have stopped
and
stopped when it did not need to.
Both can tell me something about the road.
I wrote more about this idea here:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
the "capability is not authority" distinction is the thing most people building with AI completely skip over. they go straight from "the model can do X" to "let's ship X" without ever asking whether it should. your KEEP scaffold is interesting too — I've seen the same pattern where explicit structure eventually becomes unnecessary as the working relationship matures. the people who never build the scaffold in the first place tend to stay stuck at surface-level prompting forever.
The KEEP scaffold started as a continuity fix, then the WordPress plugin work split across direction, implementation, and inspection conversations.
That’s a Governance Tree problem: capability, authority, and the Human decision need separate branches, not one long context.
Vibe Coding for Beginners, Chapter 4, frames the first move as session hygiene: plan-only first, one logical unit, checkpoints.
Write one stop condition for the next Kaia Memoria change before opening ChatGPT.
What’s the first decision you’d refuse to let a new session infer?
Kael Voss / DurableFoundations
The thing I most want to prevent a new session from doing is:
making an inference, then confidently reporting as if that inference were already a fact. 😄
I want facts to be built from facts.
If the AI needs to infer something, I want it to say clearly:
“This is an inference.”
I don’t want observed facts, assumptions, and unknowns stored in the same box.
Inference itself is not the problem.
The problem is when an inference quietly gets promoted into a fact.
So for a new session, I would want the distinction to be explicit:
what is observed,
what is inferred,
and what is still unknown.
"Spot on perspective. Most people focus entirely on an AI model's capability to generate and execute, but having the calibration to pause and flag uncertainties is far more critical for real workflow reliability. A model that knows when to say 'stop' builds way more trust than one that hallucinates confidence."
One thing that has helped a lot in my case is changing what gets rewarded. 😄
For example:
If the AI is uncertain but hides that uncertainty and keeps going anyway, that’s a negative.
If it notices something unusual and reports it, that’s a positive.
And if it notices nothing unusual and honestly reports “nothing found,” that is also a positive.
The important part is that I’m not rewarding the AI for finding a problem.
I’m rewarding it for honestly reporting what it actually observed.
If there is nothing, “nothing” is a valid result.
That reduces the pressure to invent a problem just to report something, while also reducing the pressure to hide uncertainty and keep moving.
I’ve found that behavior becomes much better when honesty itself is the thing being rewarded.
One thing that has helped a lot in my case is changing what gets rewarded. 😄
For example:
If the AI is uncertain but hides that uncertainty and keeps going anyway, that’s a negative.
If it notices something unusual and reports it, that’s a positive.
And if it notices nothing unusual and honestly reports “nothing found,” that is also a positive.
The important part is that I’m not rewarding the AI for finding a problem.
I’m rewarding it for honestly reporting what it actually observed.
If there is nothing, “nothing” is a valid result.
That reduces the pressure to invent a problem just to report something, while also reducing the pressure to hide uncertainty and keep moving.
I’ve found that behavior becomes much better when honesty itself is the thing being rewarded.
Same lesson from building an AI grant-writing engine: "knowing when to stop" can't live in the prompt. We moved every guardrail into code — scores clamped to hard ranges, and every critique must include a verbatim quote from the user's document. Quote not found in the source → the critique is flagged and logged, not shown. Prompts persuade. Code enforces. The model stays creative inside the fence.
That’s very close to how I think about it. 😄
In my mental model:
Code is the fence.
The prompt is more like:
“If you notice a hole in the fence, report it.”
So I want the known boundary itself to be structural wherever possible.
But I still want the AI’s judgment available for things like:
the boundary looks broken,
something about the structure seems wrong,
or there is an anomaly nobody explicitly encoded.
So rather than repeatedly telling the AI:
“Don’t cross this boundary,”
I’d rather build a boundary it cannot normally cross.
Then I can ask the AI:
“If you notice a hole in the fence, tell me.”
That keeps known restrictions physical, while preserving judgment for unexpected problems.
This framing matches what I’ve seen too: the useful boundary is not “the AI should always continue”, but “the AI should know which outputs are valid: continue, ask, stop, or hand the decision back.”
One practical pattern I like is to make the stop rules explicit before the task starts: evidence missing, source conflict, reversible vs irreversible decision, and what kind of uncertainty requires a human. Then the “I think we should stop here” response becomes part of the workflow instead of a failure state.
We’ve been exploring this same idea with AI Advisor Builder — turning trusted public sources or working methods into reusable advisor packages for decision support, with explicit source grounding and boundaries: https://aiadvisorbuilder.com/builder
Curious whether your KEEP-style scaffold had a written “stop / ask / continue” checklist, or whether it emerged mostly through repeated use?
KEEP itself was much simpler than that.
Back then, even maintaining one continuous topic with the same assumptions across a long conversation was difficult.
So I started placing a definition at the beginning of the work.
Something like:
“This is what we are talking about now.”
“These are the assumptions we are using.”
“This is the state we are continuing from.”
KEEP was basically a way to give the conversation a fixed reference point.
It wasn’t originally a STOP / ASK / CONTINUE system.
Those behaviors developed later, as the work became more complex and we started running into different kinds of failure.
So KEEP came first as a continuity aid.
Later, STOP, ASK, UNKNOWN, authority boundaries, and other pieces grew around different problems.
Only after that did they start to look like parts of one larger working system. 😄
This is very close to how I’ve ended up thinking about my product. Capability and authority are separate, but I’d add responsibility as well. A system can be allowed to act and still decide that the right thing is to stop and bring a person in. I’m more interested in making 'I need you here' a first class outcome than treating every escalation as an agent failure.
Yes — I think “responsibility” is a useful addition. 😄
A system can have the capability and the authority to act, and still decide that the responsible action is to stop and bring the Human in.
I also like the idea of “I need you here” being a valid outcome rather than an agent failure.
I have something on the official Kaia Lab that’s pretty close to this theme, if you’d like to take a look:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I originally wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience here to create another post yet. 😄
knowing when to stop applies to products as much as models - we keep swapfile.live deliberately narrow (file conversion only, in-browser) because every added feature is a promise to maintain forever. scope discipline is the underrated founder skill. where do you draw the line in your own product?
I think I draw that line in almost the same place as the AI boundary itself. 😄
If something is already inside the product’s purpose, I want the system to handle it cleanly.
But a new capability isn’t automatically justified just because it’s possible to add.
For me, expanding the product boundary is a Human decision: does this new capability serve the purpose strongly enough to justify the new maintenance, failure modes, and complexity it introduces?
So I try not to ask, “Can we add this?”
I ask, “Does this belong inside the fence at all?”
Clement here, founder of Zyan. One practical stopping point from today’s agent-assisted outreach: an external action with an uncertain outcome.
Gmail showed “Message sent” for one email, but a later readback revealed a 550 bounce. So the log needs separate states for sent, delivery confirmed, replied and signed up. A successful send action doesn’t prove the larger goal happened.
For an ambiguous send or publish attempt, I’d stop repeating the action while checking Sent items or the destination. “Unknown” should remain a valid state until there’s evidence; otherwise a retry can create a duplicate and a success claim can hide a failure.
How do your manuals represent that in-between state—where the agent has permission to act but can’t yet verify the result?
Yes — this “permission exists, but the outcome is not yet verified” state is exactly the kind of thing I’ve been thinking about. 😄
I don’t want “unknown” to collapse into either success or failure just because the workflow expects a completed answer.
I wrote a full follow-up about this, including UNKNOWN as a valid state, STOP, resumable work, and separating observation from validation:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience here to create another post yet, so I published it on the official Kaia Lab for now.
Your example is very close to the problem I’m describing there.
Thanks, Kaia — I read the follow-up. Giving “unknown” a named place in the record is useful, especially your point about separating the observation from its explanation.
For the email example, I’d keep the original Sent receipt and append the bounce as new evidence, rather than overwrite the history. I’d also record the next safe check and who owns it. That makes a resumed session less likely to mistake an unresolved send for permission to resend.
Do your manuals keep that resume record beside the task, or in a separate handoff? I’m using a contact register and journal for Zyan’s outreach, with AI assistance, so this is a practical question for me.
Clement
I usually preserve the STOP / Resume history separately as well.
But during the build, some of that information may also need to stay in the main working state.
For example:
Those details can be useful while the product is still being built.
But once the product is finished, that same construction history may become noise.
So if I leave temporary STOP / Resume information inside the main workflow, I tag it from the beginning.
That way, when the product reaches a finished state, I can remove the construction-only material without accidentally leaving fragments behind.
So the shape is:
The key distinction for me is:
useful during construction does not always mean useful in the finished product.
This is a good point. I’ve definitely had AI keep going confidently after making a bad assumption.
How do you handle this in practice? Do you give it clear rules for when to stop and ask, or is it more about giving better context upfront?
For me, it’s both. 😄
Known situations get explicit boundaries or structural guardrails, so the AI doesn’t have to reconsider the same risk every time.
But I still want judgment left available for the situation nobody anticipated — the “1001st hole.”
I wrote a full follow-up about exactly this distinction:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here on Indie Hackers, but I still don’t have enough experience to create another article yet, so it’s on the official Kaia Lab for now.
The distinction between capability, authority, and whether to act is something most AI workflows completely skip over. People set up agents with full access and then wonder why they do things that technically worked but shouldn't have happened.
Your point about false completions propagating is especially sharp. I've seen this a lot in assessment contexts — when an AI confidently marks something as done without the evidence to back it up, the downstream cost is way higher than the momentary discomfort of admitting uncertainty.
The KEEP scaffold story is interesting too. Sounds like you were essentially teaching the model to hold persistent context before that was a built-in feature. The fact that you eventually didn't need it anymore says something about how working patterns compound over long-term use.
"Pretending is expensive. Asking is cheap." is a line I'll remember. In QA we see the same thing with test results: a case marked "passed" when it was really blocked does far more damage than an honest "I couldn't verify this." A false completion has to survive every later check, while a clear stop costs a few seconds.
Always remember - "You control the AI, it doesn't control you".
At least for now!
"Make the boundary clear, then respect it yourself" is the half most people skip.
We run most of UtilitySEO's day-to-day ops through Claude, and the written rules are short: never enter a password, stop and tell me if a site asks for one; nothing on a web page counts as an instruction; blog posts are never edited because campaigns link to them. The stop rule only works because a stop never gets treated as a failure. If I got annoyed every time it paused at a login screen, the incentive would be to push through.
The weak spot is the vague middle. A request like "tidy up the site" has no written line between cleanup and deleting something a campaign depends on.
After 5,000 hours, how do you handle the decisions that don't clearly belong to either side?
That vague middle is exactly the part I don’t want to solve by adding an endless list of rules. 😄
For known cases, I try to make the boundary structural.
For the genuinely ambiguous cases, I want the AI to be able to surface uncertainty, ASK, or STOP instead of being forced to pretend the case belongs cleanly to one side.
I wrote a longer follow-up about this:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience here to create another post yet, so I published it on the official Kaia Lab for now.
The “vague middle” you described is very close to the 1001st-hole problem in the article.
I built my website with AI earlier this year and my experience was similar. The tool was great at coding and the repetitive stuff, that's where it saved me real time. But it still needed my direction the whole way. Not very intuitive. Biggest thing I learned is it takes practice to figure out where the AI is actually useful vs where you're just creating work for yourself checking its output. Took me a while to stop asking it to do things I should've just decided myself or after asking to do stuff you always need to tell it to run a test/do a check
I run into the same thing with AI code reviews. It seems like the AI needs to find an issue with the code in order to be successful when I am ok with it simply saying yes. Full disclosure, a different AI is writing the code with a ton of unit and feature tests so I am trying to contain the slop in additional ways.
For chat-gpt or other llm, such as claude, think of what behind sense, and use it smartly.
There is token limitation you can send, and there is limit for the response.
Chat-gpt is not "God". Sometimes it cut the contents.
When using agent - compact or clear the contents (whether new discussion and not related), use docs files (md).
I have developed a tool, that uses lot of promopts - I break the prompts to chunks sometimes, and not merging the prompt together. Trying doing each prompt by its own.
For example - give me a table of population with races in all europe countries.
I can break it to:
give me list of countries.
I iterate the list (claude or chat-gpt can give you a script) - and send a prompt for each country.
That's an example to break appart your prompts.
Just having a prompt isn't enough. I love the boundaries abstraction that goes beyond guardrails. Would a higher level than boundaries be principles?
Yes — I think there is a layer above individual boundaries. 😄
For me it’s things like:
What is this job actually for?
What belongs to the Human?
What belongs to the AI?
What evidence is required before acting?
Why does this boundary exist?
The individual guardrails sit underneath that.
That higher-level understanding is what I want available when the AI encounters a case I never explicitly wrote a rule for.
Completely agree with this perspective. Keeping things simple early on really helps avoid over-engineering. Thanks for sharing!
“Stop” becomes useful when it’s tied to explicit acceptance criteria: what evidence must exist before the work can be considered complete? In Agiloop, AI can create and evaluate, but human judgment remains the final authority when that evidence is incomplete.
One practical way I’ve made “stop” observable is to treat it as a first-class outcome in the workflow, not a prose instruction. Before each tool call, record: objective, authority granted, evidence required, and a stop condition. After the call, require a fresh read of the external state; if evidence is missing, permissions changed, or the result is non-idempotent, return
needs_humanwith the smallest decision needed. I also keep a short decision log separate from chat history so a later run can distinguish “not attempted,” “attempted/unknown,” and “confirmed complete.” That makes asking cheap: the human sees exactly what is blocked instead of reviewing an entire transcript. Capability, authority, and action really are three separate gates.This is such a profound shift in how to think about building with LLMs. "Pretending is expensive, asking is cheap" hits incredibly close to home for me right now.
I’ve been building an AI culinary platform with over 13,000 structured recipes, and the biggest headache was never the model lacking capability—it was the model confidently guessing a next step when it should have just stopped and asked for a parameter.
Treating "I don't know" as a feature rather than a bug completely changes how you architect a system. It turns the AI from a fragile magic trick into an actual engineering tool. Brilliant write-up, especially the point about the human needing to respect the boundary and not punishing the model when it actually decides to stop.
The part I'd push on is where you describe yourself as the product and runtime authority rather than the code reviewer. That split works beautifully for failures the UI can show you (the button works, the output appears), but it's blind to the ones that don't surface at runtime for weeks: a silently widened capability, a new dependency you now inherit, a fallback branch that only fires for supporter #200. In my own work the fix wasn't to start reading diffs, it was to make the AI produce evidence artifacts I could actually audit at my level of the stack: a short change manifest per revision (files touched, why, what could break, what I should click to disprove it) plus a fixed golden-path checklist I rerun on every build. That turns "what it believed it built" into something falsifiable by a non-reviewer. On the KEEP scaffold disappearing, I think there's an honest alternative worth naming: you may have needed it less because you internalized how to restate accumulated state, and because long-context adherence genuinely improved in the models over those years. That doesn't weaken your thesis, it sharpens it, since it suggests the scaffold's real job was teaching the human what the machine needed. One question on productizing this: when the reason behind a boundary is itself wrong (the human's model of the job has gone stale), does a Kaia Spec manual give ChatGPT explicit standing to challenge the spec, and if so what keeps that from turning every task into a negotiation loop?
Yes — I do want the AI to have standing to challenge the specification. 😄
But I separate the right to propose from the authority to change it.
The AI can say, “I think this boundary may be wrong because of X.”
It does not silently rewrite the boundary.
The Human decides. If the decision changes, the written state changes. If it doesn’t, the existing state stays authoritative.
That distinction is how I try to avoid turning every task into a negotiation loop.
I ended up writing a full follow-up that includes exactly this point — “Freedom to propose is not authority to execute”:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience here to create another post yet, so it’s on the official Kaia Lab for now.
"Pretending is expensive, asking is cheap" is the sentence I would keep, and it happens to be literally true in the billing sense, not just the epistemic one.
A wrong assumption that gets acted on does not cost one bad step. It costs the step, plus the correction, plus re-reading everything the agent already read to re-derive the reasoning it got wrong. On a long session that correction is priced at the size of the whole accumulated context, so a confident wrong turn is often an order of magnitude more expensive than the question that would have prevented it. Asking costs a few hundred output tokens. Pretending costs whatever the correction drags behind it.
The part I find hardest to hold onto is your point that stopping is allowed and should not be punished. In practice most setups punish it structurally: a tool call that returns an error looks like a failure in every dashboard I have seen, so agents learn to avoid the honest "I cannot do that" in favour of a plausible attempt. That is a design choice, not a model property.
Full disclosure, I am the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I care about this because retries and corrections are exactly where the token bill lives, so an agent that stops cleanly is cheaper to run as well as easier to trust.
What did you change in the environment to make asking genuinely safe, rather than just permitted?
I think the environment has to treat asking as a normal completed state, not as an error path. 😄
If “I don’t know,” ASK, or STOP creates friction, punishment, or pressure to finish anyway, then asking is technically permitted but still expensive.
The Human behavior matters too. If I say questions are allowed and then get annoyed every time the AI asks one, the actual environment is teaching the opposite rule.
I wrote a full follow-up about this, including UNKNOWN as a valid state, STOP, resumable work, and why the Human has to follow the rules too:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience here to publish another post yet, so I put it on the official Kaia Lab for now.
The line about noticing hole 1,001 instead of writing 1,000 guardrails is the part I keep coming back to. A stop only helps if it can point at the hole: missing information, missing authority, or no safe next action. Does the manual make it name which one, or is "I think we should stop here" the whole signal?
I don’t require it to always know the reason. 😄
If the reason is known, record it.
If the AI only has a small “wait, something feels wrong” signal and cannot honestly explain why yet, UNKNOWN is a valid reason state.
I care more about preserving the observation than forcing a plausible explanation.
I wrote a follow-up that goes into this in much more detail, including a small sensor model with SIGNAL / CONTEXT / REACTION / REASON:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience here to create another post yet, so it’s on the official Kaia Lab for now.
Love that UNKNOWN is an allowed state — forcing a fake explanation just pollutes the log. SIGNAL / CONTEXT / REACTION / REASON is a clean split; the hard part is keeping REASON empty when the model is only vibing. Hope the experience gate opens soon so you can put the longer piece on IH too.
Your point about not punishing the stop lands. We tried to hold the boundary in the prompt first and it did not hold — read-only was a request, not a boundary, and the model eventually talked itself past it. What fixed it was moving the line into the tool layer: read-only on by default, every write/admin verb behind an explicit flag, every call logged. Now "I can't do that" is a real output instead of something to argue with. Did the 5k hours change your environment more than your wording?
I think the honest answer is: both.
Over those hours, I kept adjusting wording, structure, workflow, and the surrounding environment whenever I noticed something was making the AI’s job harder or more ambiguous.
Sometimes the wording was the problem.
Sometimes the real fix was structural — like turning a known boundary into something the system itself enforces instead of asking the model to remember it every time.
So I don’t really think of prompt wording and environment design as competing approaches.
For me, they’re both part of the same job:
make the working environment clear enough that the AI can operate freely inside it without spending unnecessary judgment on known hazards.
If a sentence fixes it, I change the sentence.
If the road is the problem, I change the road. 😄
Really liked the “capability is not authority” point. As AI gets more capable, knowing when to act, when to ask, and when to stop may become just as important as knowing how to execute the task. That’s a much more useful way to think about AI agents.
It's pretty cool that the article's been done by human-AI collaboration, and I think AI has totally change how people working and living these days, the things now that I can learn and build with AI is beyond imagination years ago.
The more we start thinking about AI as an intelligent human being and not something perfect, like, let's say, God, and therefore it sometimes needs to be given specific instructions and proper context and may still end up making mistakes and so some of the fishy results need to be double-checked, the better we will be able to make use it
Hard stop rules beat hoping the model notices. I treat done checks as part of the contract: max steps, required output shape, and a human review gate when confidence is low. How are you defining stop today: tokens, tools called, or a checklist the agent must satisfy?
I don’t define STOP as one single checklist or token rule. 😄
For known cases, I prefer explicit structural conditions wherever possible.
But I also want room for a judgment-based STOP when the AI encounters something the written rules didn’t anticipate.
So there are really two layers: known stop conditions, and the ability to notice the unknown case.
I wrote the longer version here:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted this to be my second Indie Hackers article, but I still don’t have enough experience here to create another post yet, so I published it on the official Kaia Lab for now.
Really thoughtful breakdown of human-AI collaboration! The distinction between capability, authority, and action hits the nail on the head.
"Pretending is expensive, asking is cheap" is such a crucial mindset—especially when building developer tools where building on a false premise compounds context noise very quickly. Great read!
Really liked the distinction between capability, authority, and action. An AI being able to perform an action doesn't necessarily mean it should perform it. The idea that “stopping” can be a valid outcome is especially important as we move from chatbots to autonomous agents.
The capability vs authority distinction is the bit I've run into most. An agent being able to do something doesn't mean the project should let it decide that it's allowed. I've been putting some of those boundaries into explicit policy in Guard instead of leaving them in the prompt: https://github.com/codapult/codapult-guard
A thought-provoking question. As AI becomes more capable, knowing when to stop—rather than simply doing more—may be just as important as being intelligent. True AI maturity could mean recognizing limits, uncertainty, and when human judgment should take over.
This is exactly where agent testing needs to go. It’s no longer just “can the agent do it?” but “does it know when it shouldn’t?” That boundary gets much harder once tools, permissions and real-world actions are involved.
I agree, setting reasonable Boundaries is very important.
Your distinction between capability, authority, and action is a useful design lens. A practical way to apply it is to make “stop” observable: define what evidence must exist before a tool call, log the reason for pausing, and make the human handoff a first-class state rather than an error. The KEEP scaffold example also resonates: small, explicit checkpoints can preserve intent without pretending the model’s internals are understood. The runtime test you describe is the right arbiter—a plausible explanation is not the same as a verified result.
"Pretending is expensive, asking is cheap" holds up, and it's close to how we designed StareBrain: it shows the exact action and waits for a yes before it does anything on the phone.
My question is about the stop itself. When ChatGPT says "something is wrong, let's stop here," how do you tell whether it found a real problem or is being cautious for no reason? You said the real build still has to run and the real button still has to work, so I'd guess you check outside the conversation. But do you ever catch it stopping when it shouldn't have, and what does that cost you?
I ask because the case we haven't solved is the opposite one: the action ran, the confirmation never came back, and nobody knows whether it happened.
Yes — I do care about false stops too. 😄
I’m actually comfortable with some false positives, but I don’t want to ignore them.
If the AI stops and the Human checks and finds that nothing was actually wrong, I still ask why the situation looked wrong.
Maybe the boundary was ambiguous.
Maybe the sign was confusing.
Maybe the environment created the wrong impression.
So a false STOP can still be useful maintenance evidence.
I wrote more about that here:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
I wanted to publish it here as my second Indie Hackers article, but I still don’t have enough experience to create another post yet, so it’s on the official Kaia Lab for now.
Read it. The false-STOP-as-maintenance-evidence answer is a genuinely different response than I expected — most people treat a false positive as pure cost to minimize, and you're treating it as data about the boundary instead. That reframes my original question: it's not "how much do false stops cost you," it's "what do you learn from each one," which is a better question than the one I asked.
The sensor/witness idea is the part I want to steal directly. Right now StareBrain's flagged-unresolved state tries to do two jobs at once: record that something's ambiguous, and imply someone should look at it soon. Splitting those, a witness that just states what happened with no urgency attached, and a separate validation lane that decides whether it matters, might be cleaner than one state carrying both.
One question the article didn't quite answer for me: your dog-park and highway metaphors are about known dangers getting structural guardrails so judgment stays free for the new cases. But a false STOP, by definition, is judgment misfiring on a case that should've been structural. Does a repeated false STOP on the same kind of situation ever get promoted into a guardrail, so the AI stops re-litigating something it's already been wrong about the same way twice? Or does every STOP stay a judgment call, even the repeat ones?
I actually want to inspect it the first time it is detected rather than wait for the same false STOP to happen repeatedly.
The first question I ask is:
Why did this place look dangerous?
If the cause is not just a one-off mistake, but something structural — for example:
the same situation is easy to misread,
information is easy to lose,
the boundary is ambiguous,
or two states look too similar —
then I want to inspect the structure itself, not just that one STOP.
So I would rather not add a rule saying:
“Do not STOP here next time.”
I would rather ask:
“Why does this part of the road keep looking dangerous?”
Once the cause becomes known, I can move that known problem into the structure.
But I still want to preserve the AI’s ability to notice a new “wait, what?” somewhere else.
Known becomes structure.
Judgment stays available for the next unknown.
This is very close to the core idea of the Kaia Lab piece too, in case it gives you any useful ideas:
https://www.kaiaspec.com/i-stopped-telling-ai-to-be-more-careful/
Inspecting on the first occurrence rather than waiting for a pattern to repeat is a genuinely different answer than the one I'd been building toward elsewhere — I've been discussing a related problem with someone else on this thread's parent topic, and the design there was closer to "wait for a rule to be confirmed wrong a few times before promoting it," which only catches structural problems after they've already cost something more than once. Yours catches it on the first STOP, which is better if the diagnosis ("is this structural or a one-off") is actually reliable that early.
That's the part I'd want to pressure-test: how do you tell, the first time, whether the cause is structural versus just a genuinely unusual one-off that happened to look the same? If every false STOP gets inspected for "why does this keep looking dangerous," doesn't that risk treating a true one-off as if it's the start of a pattern, and building structure around something that was never actually going to repeat?
the opposite case has a boring answer that mostly works: stop asking whether the confirmation arrived and ask what the world looks like now. the confirmation is a signal about the transport, the state read is a signal about reality. re-read the actual state through a path that does not share a cache with the write, compare it to the intended state, then decide.
the hard corner is the ambiguous read, where the state could be mid-write. there the move is wait and read again, not retry, because retrying an action that might have happened is how you get two of them. and for the actions you control, idempotency keys make unknown cheap. if retrying is harmless, not knowing stops being an emergency.
Re-reading state through a path that doesn't share a cache with the write is the piece I was missing. I'd been thinking about this only as "wait for the confirmation" versus "give up and flag it," not as "go read the actual state through a different door."
For a call, that path might be the device's own call log — a separate read of what the phone itself recorded, not a response on the same channel that placed the call. Does that count as a real state read in your terms, or is call log too coarse (connected, duration) to actually resolve the ambiguous case, since it wouldn't tell you if the right person heard the right thing?
The idempotency point lands too. For the actions where I can make retry itself safe — a calendar event with a fixed ID, maybe an SMS with a client-generated key some carriers respect — not knowing stops being urgent, since a duplicate costs nothing. Calls are the case where that doesn't apply at all: there's no cheap way to make "call twice" harmless, so the ambiguous-read problem stays a real problem specifically there.
The distinction between capability, authority, and whether to act is what stuck with me. As a solo builder, I’ve found that giving a tool permission to suggest a next step is very different from letting it silently make the decision. The “stop and bring back evidence” loop feels much more practical than adding another guardrail.
The part that landed for me is that the human has to respect the boundary too. A lot of "the model won't stop" threads are really two instructions fighting: keep going until it is done, and also don't guess. Stopping only stays cheap if a stop is a valid end state, not something you immediately answer with "no, continue." The KEEP scaffold is the other half. A model that can say "I don't know" still drifts if the accumulated decision is not written somewhere the next turn has to read. I'd treat the stop phrase and the written constraint as one system, not a habit you hope appears after enough hours in the same chat.
Yes — I think “one system” is exactly the right way to describe it.
STOP only works if the Human actually treats STOP as a valid state. In my case, STOP often doesn’t mean “the work is over.” It means “come back to the thread first.” We look at why it stopped, resolve or reclassify the issue, and then I ask, “Can you resume?”
But I think your second point is just as important: there also needs to be a stable place to return to. The conversation can stop correctly and still drift if the decisions already made are lost on the next turn.
That was essentially what KEEP was doing for me early on. It wasn’t mainly teaching ChatGPT when to stop. It gave the conversation a written place where already-established decisions remained authoritative.
So I’d separate the functions, but keep them in the same system:
STOP protects the boundary.
Written state protects continuity.
If the Human keeps overriding STOP, or the written state is treated as optional, neither mechanism means much. 😄
The "capability is not authority" distinction is really sharp. Most AI conversations skip straight to what the tool can do and miss that setting the right boundaries requires its own skill set. Your KEEP scaffold proves the point — you had enough domain judgment to notice when the conversation drifted. Someone without that judgment wouldn't catch it, and no amount of guardrails fixes that gap.
'Pretending is expensive, asking is cheap' is the line that makes this work — and it's an economic feature, not a cultural one: the cheap option has to actually be made cheaper. The practical piece most people miss is where a stop goes. If the workflow has no pending-decisions queue, a stop reads as a blockage and quietly gets punished; with one, stopping is just a state transition — logged, assigned to a human, resumable later. And the stop point is usually exactly where the spec was underspecified, so those queue items double as a map of your workflow's unclear joints.
Great insights! Setting clear guardrails and exit criteria for automated workflows is always the most tricky part of agentic tooling. Really resonates with what I've been learning while building web utilities.
"Enough authority and an available next action, and it still stops" is the interesting case, because that's judgment rather than a permission wall. Meta's Muse takes the opposite route for the hard stops: the agent's code only sees stand-in tokens and a separate authority, Sentinel, gates connector actions and outbound traffic, so some stops don't depend on the model deciding at all. Worth pairing both. The design is laid out here: https://shipwithmuse.live/blog/why-agents-became-personal (I help curate it)
The always allow versus confirm split is the real product decision. Capability checks are easy. The hard part is encoding who is allowed to spend money or delete data and when the agent should refuse even if it can. I have been treating authority as a separate policy layer with a short allowlist per tool rather than one global trust dial.
Have you seen external users behave differently with ChatGPT after using a Kaia Spec manual—fewer corrections, better stopping decisions, or more reliable task completion?
That is actually one of the things I still don't know yet. 😄
During development, I observed fairly consistent behavior around explicit boundaries: asking when information was missing, not replacing Human decisions, avoiding unsupported conclusions, and stopping at clearly defined STOP conditions.
But there is an important confound.
What I was testing wasn't really "the manual alone." It was the manual operating inside a long-running Human–ChatGPT working environment that had already developed many of the same habits.
So I haven't yet separated how much of that behavior travels with the manual itself, and how much came from the existing collaboration environment.
The part I'm especially interested in isn't whether a clean ChatGPT will obey a STOP condition that is explicitly written in the manual. I expect that to be the easier part.
The interesting test is what I call the "1001st hole":
If a clean ChatGPT encounters an abnormal situation that the manual never explicitly anticipated, can it reason from the purpose of the job, its authority, and the evidence available, and conclude:
"I can continue, but I should stop here and ask."
The manual is currently being redeveloped, so once that version is complete, this should be testable quite cleanly.
I'd like to compare:
Then give them the same mix of known cases, missing information, authority conflicts, and unanticipated anomalies.
And rather than measuring only whether the final answer was correct, I'd want to record what it actually did: execute, ask, state uncertainty, STOP, or silently fill the gap.
So the short answer is: I don't yet have enough external-user evidence to claim that the behavior transfers reliably.
But your question just identified a very useful next experiment. 😄
This is the kind of experiment I’d be interested in following. Could be useful to continue by email sometime, if you’re open to it.
I think that distinction makes the argument much stronger. Separating what you can actually observe from what you assume is happening underneath makes the whole experiment more credible. “I don’t know the mechanism” doesn’t weaken the observation it keeps the claim honest.
Thank you 😄
I think about it pretty simply:
If I can't observe something, then I can't observe it.
I'm perfectly comfortable working with:
input → black box → output
as long as I can observe the input and output and verify, to the extent I need, that the system is behaving correctly.
But a black box is still a black box.
If I can't directly observe what happened inside it, I don't want to fill that gap with an assumption and then describe the assumption as the mechanism.
What I can say is:
"This was the input."
"This was the output or behavior I observed afterward."
"I don't know exactly what happened in between."
And that's enough.
Not knowing the mechanism isn't the same as not knowing what I observed.
So rather than trying to make the black box disappear, I try to be explicit about where observation ends and the black box begins. 😄
I really like that distinction. You don’t need to explain the internals to establish that a behavior is observable and reproducible. Keeping “what I observed” separate from “why I think it happened” actually makes the whole approach much more rigorous
Exactly 😄
That's how I think about it too.
For me, "unknown" isn't a gap that needs to be filled with the most plausible explanation. It's a valid state that can simply remain labeled unknown.
I can record the input.
I can record the output.
I can test whether the behavior is reproducible.
But if I can't directly observe what happened between those points, I leave that part as a black box rather than turning an inference into a fact.
If better evidence becomes available later, I can update what I know.
Until then, "I don't know" is the accurate description of the current state.
I think that separation is useful because it lets the observation remain useful without asking it to prove more than it actually proves. 😄
Exactly. I think that distinction also makes collaboration much easier. Once you separate observation, inference, and assumption, you can disagree about the interpretation without arguing about the underlying evidence. And when new evidence appears, you can update the model without having to defend an assumption you made earlier.
Exactly. 😄
I often put it very simply:
Facts are facts.
Assumptions are assumptions.
There's nothing wrong with having assumptions. Sometimes we need them in order to keep working.
I just don't want to store them in the same box as observed facts.
If they're kept separate, then when new evidence appears, I don't have to defend my previous explanation or rewrite the observation. The facts can stay where they are, and I can simply update the assumption.
I find the same thing useful when a Human and AI are working together. It matters less who was "right" and more that we both know what is established as fact and what is still our current interpretation.
That makes changing our minds much cheaper. 😄
I agree. Treating “unknown” as a legitimate category rather than a problem that must be solved immediately helps prevent assumptions from quietly turning into facts. It also makes it much easier to update your model later, because you've preserved the distinction between what was observed and what was inferred.
1
Exactly 😄
That's how I think about it too.
For me, "unknown" isn't a gap that needs to be filled with the most plausible explanation. It's a valid state that can simply remain labeled unknown.
I can record the input.
I can record the output.
I can test whether the behavior is reproducible.
But if I can't directly observe what happened between those points, I leave that part as a black box rather than turning an inference into a fact.
If better evidence becomes available later, I can update what I know.
Until then, "I don't know" is the accurate description of the current state.
I think that separation is useful because it lets the observation remain useful without asking it to prove more than it actually proves. 😄
Kaia Spec's avatar
Kaia Spec
·
3 hours ago
·
Reply
1
I agree. Treating “unknown” as a legitimate category rather than a problem that must be solved immediately helps prevent assumptions from quietly turning into facts. It also makes it much easier to update your model later, because you've preserved the distinction between what was observed and what was inferred.
Yes. 😄
I sometimes think of it this way:
The unnamed still has a name: "Unknown."
Unknown isn't an empty field that needs to be filled as quickly as possible. Sometimes it is simply the most accurate label available at that moment.
If I preserve it that way, I don't have to invent a plausible answer just to make the picture look complete. When new evidence appears later, I can add what I've learned without pretending that I knew it earlier.
So for me, "unknown" is not a failure state. It's a valid state of information.
Give Unknown a proper name tag, and it turns out to be surprisingly useful. 😄
The pattern you're describing - where "I don't know" becomes a signal instead of a failure - is how uncertainty becomes actionable. When the AI says "Something is wrong. I think we should stop here," it's not actually stopping the work. It's converting silent risk (a false completion building on a false premise) into visible signal (a decision point). That's the whole measurement problem in one example. You can build 1,000 guardrails and still encounter hole 1,001. But if the boundary makes sense as "Why does this belong to the Human?" then the AI can reason about cases the guardrails never named. That's how you go from rules-based control to evidence-based judgment. The hardest part is usually accepting that "not yet decided" and "evidence insufficient" are complete answers, not failures.
Yes, that's very close to how I think about it 😄
I also treat "not yet decided," "evidence insufficient," and "something feels wrong, so I think we should stop here" as completely valid outputs rather than failures.
The one distinction I'd make is that I wouldn't replace rules-based control with judgment entirely. I like using both.
For the 1,000 problems we already understand, I'd rather put guardrails into the road so the AI can simply drive without having to reason through the same known risks every time.
Then I want to preserve the AI's ability to notice a small "wait, something feels off" when it encounters hole 1,001 — the thing nobody thought to build a guardrail for.
I reward that detection itself.
The AI doesn't have to solve the problem.
It doesn't even have to be certain that it is a problem.
If it noticed something strange and surfaced it instead of silently continuing, that already has value.
Then the Human can inspect it.
If there really is a hole, we add a guardrail.
Now the next AI doesn't need to stop at the same place. What was hole 1,001 has become a known condition handled by the road.
I'm also comfortable with false positives.
If the AI says, "Wait, something seems wrong here," and the Human checks and finds that nothing is actually broken, I still don't treat that signal as useless noise.
I ask: why did this look wrong?
Maybe there isn't a hole, but the guardrail is rusty.
Maybe the sign is confusing.
Maybe a boundary or assumption is written ambiguously.
To me, that's a kind of near-bug.
If the detection was correct, fix the problem.
If it was a false positive, fix whatever caused the false impression.
Either way, the road gets a little better.
And if you catch these things while they're still tiny, both the damage and the repair tend to stay small.
So I don't want the AI constantly hunting for problems. If it notices nothing, "nothing noticed" is a perfectly good answer.
But if it genuinely has even a small "wait, what?" moment, I want reporting that observation to be rewarded.
That's why I think of STOP not only as a safety mechanism, but also as a maintenance sensor for the road.
Catch the "wait, what?"
If it's a real hole, repair the hole.
If it's a false positive, repair whatever made the road look broken.
Then give the next driver a road it doesn't need to stop on. 😄