Sometimes ChatGPT stops working for me.
Not because it can't continue.
Not because it ran out of permission.
It can have enough information, enough authority, and a perfectly available next action — and still come back with something like:
Something is wrong. I think we should stop here.
That has become one of the most useful things ChatGPT can do in my work.
It also wasn't always like this.
I've spent more than 5,000 hours working with ChatGPT over time. I didn't spend those hours developing a theory of human-AI collaboration. I wasn't trying to train a model or invent an agent framework.
Mostly, we were making things.
At first, a lot of them were strange things.
We would try something, break something, talk about what happened, change the way we worked, and try again.
As the work became more complicated, something else gradually became more complicated too:
the way we worked together.
Only later did I realize that many of the problems people now discuss under terms like human-AI collaboration, agent boundaries, human approval, and reliable AI behavior were problems we had already been stumbling through in everyday work.
And some of the solutions we ended up with are surprisingly simple.
One of my basic rules is this:
Make the boundary clear — and then respect it yourself.
If I tell ChatGPT:
You can decide this.
I don't wait until I dislike its decision and suddenly say:
Why did you decide that without asking me?
If I say:
This decision belongs to the Human.
I don't throw it back at ChatGPT later because making the decision became inconvenient.
If stopping is allowed, I don't punish it for stopping.
If asking me is allowed, I don't treat the question as a failure.
And if I expect the AI to be honest with me, I have to be honest with it too.
That sounds almost too simple.
But over a long enough working relationship, it changes what becomes cheap — and what becomes expensive.
In my working environment, all of these are valid outputs:
I don't know.
I can't do that.
This is harder than it looks.
I don't have the evidence for that.
Something is wrong.
I think we should stop here.
This decision belongs to you.
None of those automatically means the AI failed.
Sometimes they are exactly the right result.
If ChatGPT pretends something exists when it doesn't, we now have to build the next step on a false premise.
If it pretends a task is finished, the false completion has to survive the next inspection.
If it silently guesses what I wanted, the guess may propagate into five more decisions.
Hiding the problem doesn't remove the problem.
It gives us two problems: the original one, and the false reality we have now built around it.
But if ChatGPT says:
I don't know. Can you decide this?
the problem may take thirty seconds to resolve.
So in this environment:
Pretending is expensive. Asking is cheap.
And if you're thinking:
“My ChatGPT doesn't work like this.”
That's part of the story.
Mine didn't always work like this either.
One of my early problems with ChatGPT was much more basic.
Keeping one complicated idea intact across a long conversation was difficult.
We could be working on A.
Then I would add a new condition.
Instead of continuing with the accumulated A plus the new condition, ChatGPT might effectively produce a new version of A.
So I made a crude scaffold.
I called it KEEP.
The basic idea was something like:
KEEP 1 = the thing we have already established
Then I would add numbered conditions, changes, or exceptions while repeatedly telling ChatGPT to keep referring back to KEEP 1.
It wasn't a product.
It wasn't an AI technique I had learned somewhere.
It was something we made because our conversation kept falling apart.
At the time, ChatGPT described the structure as something it had begun to hold onto and use as a basis for continuing the conversation.
I cannot inspect what that meant internally.
I can't tell you what persisted, where it persisted, or by what mechanism later behavior may or may not be related to those conversations.
What I can tell you is what I observed from the outside.
Over time, I found that I needed the scaffold less and less in my conversations with ChatGPT.
Eventually, I stopped writing KEEP altogether.
That experience changed the way I thought about working with AI.
A lot of AI control looks like this:
Don't do A.
Don't do B.
Ask before C.
Never change D.
If E happens, stop.
Those rules can be useful. We use explicit boundaries too.
But there is a problem.
You can build 1,000 guardrails.
Then the AI encounters hole number 1,001.
If all it knows is the list, the new case isn't on the list.
So I became much more interested in something one level above the rules:
Why does the boundary exist?
What is this job actually trying to accomplish?
What is the Human responsible for?
What is the AI responsible for?
What evidence is required before acting?
Why is this particular decision not the AI's decision to make?
Once those things are understood well enough, an unfamiliar situation can produce a different question:
I can do this. But should I?
And that leads to one of the most important distinctions in the way I work with ChatGPT:
Capability is not authority.
And even authority does not require action.
Those are three different questions.
Can the AI do it?
Is the AI allowed to do it?
And even if both answers are yes:
Should it actually do it?
I can give ChatGPT considerable room to act while still expecting it to decide that sometimes the correct use of that freedom is:
not to use it.
Stopping can itself be a decision.
In August 2026, I had a much more ordinary idea.
I wanted a casual way to give crowdfunding supporters a small digital thank-you keepsake.
Nothing about this began as:
Let's conduct an experiment in human-AI software development.
We started talking about what the keepsake would need.
A number.
A timestamp.
An individual identity where necessary.
A way to issue it.
Somewhere in that process I had a realization roughly equivalent to:
Wait. We can make WordPress plugins?
So we did.
That work eventually became Kaia Memoria.
As the project grew, our way of working began to break again.
One long ChatGPT conversation could handle implementation details, but as the context became larger, keeping the entire direction of the project intact became harder.
So we split the work.
One ChatGPT conversation kept track of the larger direction and previous decisions.
Another focused deeply on implementation.
Another inspected what had been built.
Sometimes a separate analysis conversation compared versions or investigated a narrow problem.
We didn't sit down one morning and design a sophisticated “multi-agent architecture.”
The organization emerged because the work kept producing problems that needed different kinds of attention.
The Human role remained important.
I defined product intent and boundaries, made decisions that belonged to the Human, installed and ran actual builds, observed the real UI and runtime behavior, and decided whether the result was acceptable.
The AI side investigated source, traced dependencies, explained structures, implemented revisions, compared known-good versions, prepared inspections, and handed findings between roles.
I didn't directly edit the source code.
I didn't even open the development ZIPs.
That wasn't necessary for my role.
I was not the code reviewer.
I was the product and runtime authority.
The AI could tell me what it believed had been built.
But the real build still had to run.
The real button still had to work.
The real output still had to appear.
And I had to decide whether what happened was actually what we intended.
When something looked wrong, completion was not automatically the next goal.
Sometimes the next correct action was simply:
STOP.
Investigate.
Bring the evidence back.
Decide what happens next.
Then make the smallest justified change and test again.
That is how a strange little thank-you idea turned into working software.
This is where I need to be careful.
It would be very easy to tell a much more dramatic story.
I could say that I “trained ChatGPT.”
I could point at behaviors that changed over time and invent a technical explanation for why.
I don't think that would be honest.
I work with ChatGPT from the outside.
I can observe behavior.
I can change the working environment.
I can introduce scaffolding.
I can make boundaries explicit.
I can see what happens when those boundaries are respected repeatedly.
I can see when an old scaffold becomes unnecessary in my work.
I can compare how ChatGPT behaves with me now to how it behaved with me before.
But there is a box in the middle that I cannot see.
So my claim is deliberately narrower:
This is what I observed. I don't know the exact mechanism behind it.
And, appropriately enough, being able to say I don't know is part of the whole point.
I started thinking about the environment around ChatGPT almost like a race circuit.
You already have an extraordinarily capable car.
Most businesses don't need to rebuild the car.
They need a place where it knows what job it is doing, where it can drive, where it must stop, what information counts as evidence, and which decisions belong to a Human.
That's what led me toward the idea behind Kaia Spec.
We make practical work manuals for ChatGPT.
Not simply giant prompts full of rules.
A useful manual needs to communicate the job:
the purpose,
the context,
the available information,
the authority,
the boundaries,
the reasons behind those boundaries,
and the points where the Human takes over.
In other words:
Don't just give ChatGPT 1,000 guardrails. Give it enough understanding to notice hole 1,001.
We aren't trying to build another general-purpose AI.
ChatGPT already has many of the general capabilities we need.
What we can build is the workplace around it.
Give ChatGPT a job.
I don't speak English.
I'm writing this with ChatGPT.
But “writing with ChatGPT” doesn't mean I gave it a prompt saying:
Write me an Indie Hackers post about human-AI collaboration.
We developed the article through conversation.
I work in Japanese.
I decide what I mean, what actually happened, what matters, where the boundaries are, and when the wording has drifted away from the idea.
ChatGPT takes that material and writes for an English-speaking audience.
When it misunderstands me, I correct it.
When it sees a structural problem, it tells me.
When something cannot be supported, we leave it out or say that we don't know.
So while you've been reading an article about human-AI collaboration, you've also been reading the output of one.
And Kaia Memoria is another.
We built it together.
Kaia Memoria is the working software project that grew out of this collaboration.
If you'd like to see what we actually built, you can find Kaia Memoria on my Indie Hackers profile and try the live demo from there.
'Pretending is expensive, asking is cheap' is the line that makes this work — and it's an economic feature, not a cultural one: the cheap option has to actually be made cheaper. The practical piece most people miss is where a stop goes. If the workflow has no pending-decisions queue, a stop reads as a blockage and quietly gets punished; with one, stopping is just a state transition — logged, assigned to a human, resumable later. And the stop point is usually exactly where the spec was underspecified, so those queue items double as a map of your workflow's unclear joints.
Great insights! Setting clear guardrails and exit criteria for automated workflows is always the most tricky part of agentic tooling. Really resonates with what I've been learning while building web utilities.
The capability / authority / necessity split matches what we saw with an AI agent inside one of our apps. Being able to do something and being allowed to do it without asking needed separate settings.
For most actions the user can click "always allow". Two actions never get that option: a permanent delete, and an edit that immediately rewrites the instructions of another AI feature. The agent can suggest those, and a person confirms each one. We put that limit into the app itself instead of hoping the model knows when to stop.
Yes — that's very close to how I think about it. 😄
For me, Human confirmation isn't only a safety mechanism. It's also a way to simplify the system.
If adding one confirmation button can eliminate both an accident path and an entire tree of automated judgment, edge cases, and potential bugs, that button is very cheap.
If I already know a boundary is dangerous, I don't really want the model spending intelligence deciding whether it should cross it every time. I'd rather put that boundary into the structure itself.
Then I can reserve the model's judgment for the cases I didn't anticipate — the "1001st hole."
I tend to think of those fixed boundaries like toll gates on a highway. You don't need the driver to reconsider the rules every time they reach one. The gate is simply there because we already know that this is a decision point that requires confirmation.
So I really like your example of permanent deletion and instruction rewriting being enforced by the app itself. That's exactly the kind of place where I'd rather make the road simpler than ask the AI to be smarter.
Love this angle. Building Xstream4K right now so this hits close to home — what made you look into it in the first place?
Funny enough, I actually looked into it pretty late.
I started with basically no AI knowledge. I just worked with ChatGPT, watched how it behaved, guessed what might help, suggested things, and we built little structures and tools as we went. Mostly, I was just making things and having fun with it.
Over time, our way of working gradually started to look like what I described in the post.
It was only when I started seriously building and releasing a real product that I finally looked outside at how other people were working with AI.
And that was when I had the slightly strange moment of:
“Wait… isn't some of this what we've already been doing?” 😅
I started with basically no AI knowledge. I just began experimenting with ChatGPT while building things, and over time, I realised that having clear boundaries and knowing when to stop made a huge difference. It all grew naturally from there. 😄
Oh, that's really interesting. 😄
The part about it growing naturally sounds very familiar to me.
I wasn't trying to invent a methodology at the beginning either. Something would go wrong, and we'd add a boundary there. We'd find a place where continuing blindly was risky, and we'd make that a place to stop and check.
After enough of those little adjustments, I eventually looked back and realized they had become a way of working.
Maybe that's the interesting part: the theory didn't come first. The useful structure survived because we kept needing it while actually building things together. 😄
That’s such a great way to look at it! 😄 Sometimes the best methodologies aren’t planned from the start; they emerge naturally from real experiences, mistakes, and the lessons we learn along the way. The fact that the structure grew out of practical needs is what makes it so valuable.
Exactly. 😄
I didn't lay down the road first.
We just kept walking, solving whatever we encountered along the way — and when I eventually looked back, there was a road behind us.
Now I'm basically looking at that road and asking:
"Okay... how do we pave this so someone else can travel it too?" 😄
"Enough authority and an available next action, and it still stops" is the interesting case, because that's judgment rather than a permission wall. Meta's Muse takes the opposite route for the hard stops: the agent's code only sees stand-in tokens and a separate authority, Sentinel, gates connector actions and outbound traffic, so some stops don't depend on the model deciding at all. Worth pairing both. The design is laid out here: https://shipwithmuse.live/blog/why-agents-became-personal (I help curate it)
The always allow versus confirm split is the real product decision. Capability checks are easy. The hard part is encoding who is allowed to spend money or delete data and when the agent should refuse even if it can. I have been treating authority as a separate policy layer with a short allowlist per tool rather than one global trust dial.
Have you seen external users behave differently with ChatGPT after using a Kaia Spec manual—fewer corrections, better stopping decisions, or more reliable task completion?
That is actually one of the things I still don't know yet. 😄
During development, I observed fairly consistent behavior around explicit boundaries: asking when information was missing, not replacing Human decisions, avoiding unsupported conclusions, and stopping at clearly defined STOP conditions.
But there is an important confound.
What I was testing wasn't really "the manual alone." It was the manual operating inside a long-running Human–ChatGPT working environment that had already developed many of the same habits.
So I haven't yet separated how much of that behavior travels with the manual itself, and how much came from the existing collaboration environment.
The part I'm especially interested in isn't whether a clean ChatGPT will obey a STOP condition that is explicitly written in the manual. I expect that to be the easier part.
The interesting test is what I call the "1001st hole":
If a clean ChatGPT encounters an abnormal situation that the manual never explicitly anticipated, can it reason from the purpose of the job, its authority, and the evidence available, and conclude:
"I can continue, but I should stop here and ask."
The manual is currently being redeveloped, so once that version is complete, this should be testable quite cleanly.
I'd like to compare:
Then give them the same mix of known cases, missing information, authority conflicts, and unanticipated anomalies.
And rather than measuring only whether the final answer was correct, I'd want to record what it actually did: execute, ask, state uncertainty, STOP, or silently fill the gap.
So the short answer is: I don't yet have enough external-user evidence to claim that the behavior transfers reliably.
But your question just identified a very useful next experiment. 😄
This is the kind of experiment I’d be interested in following. Could be useful to continue by email sometime, if you’re open to it.
I think that distinction makes the argument much stronger. Separating what you can actually observe from what you assume is happening underneath makes the whole experiment more credible. “I don’t know the mechanism” doesn’t weaken the observation it keeps the claim honest.
Thank you 😄
I think about it pretty simply:
If I can't observe something, then I can't observe it.
I'm perfectly comfortable working with:
input → black box → output
as long as I can observe the input and output and verify, to the extent I need, that the system is behaving correctly.
But a black box is still a black box.
If I can't directly observe what happened inside it, I don't want to fill that gap with an assumption and then describe the assumption as the mechanism.
What I can say is:
"This was the input."
"This was the output or behavior I observed afterward."
"I don't know exactly what happened in between."
And that's enough.
Not knowing the mechanism isn't the same as not knowing what I observed.
So rather than trying to make the black box disappear, I try to be explicit about where observation ends and the black box begins. 😄
I really like that distinction. You don’t need to explain the internals to establish that a behavior is observable and reproducible. Keeping “what I observed” separate from “why I think it happened” actually makes the whole approach much more rigorous
Exactly 😄
That's how I think about it too.
For me, "unknown" isn't a gap that needs to be filled with the most plausible explanation. It's a valid state that can simply remain labeled unknown.
I can record the input.
I can record the output.
I can test whether the behavior is reproducible.
But if I can't directly observe what happened between those points, I leave that part as a black box rather than turning an inference into a fact.
If better evidence becomes available later, I can update what I know.
Until then, "I don't know" is the accurate description of the current state.
I think that separation is useful because it lets the observation remain useful without asking it to prove more than it actually proves. 😄
Exactly. I think that distinction also makes collaboration much easier. Once you separate observation, inference, and assumption, you can disagree about the interpretation without arguing about the underlying evidence. And when new evidence appears, you can update the model without having to defend an assumption you made earlier.
Exactly. 😄
I often put it very simply:
Facts are facts.
Assumptions are assumptions.
There's nothing wrong with having assumptions. Sometimes we need them in order to keep working.
I just don't want to store them in the same box as observed facts.
If they're kept separate, then when new evidence appears, I don't have to defend my previous explanation or rewrite the observation. The facts can stay where they are, and I can simply update the assumption.
I find the same thing useful when a Human and AI are working together. It matters less who was "right" and more that we both know what is established as fact and what is still our current interpretation.
That makes changing our minds much cheaper. 😄
I agree. Treating “unknown” as a legitimate category rather than a problem that must be solved immediately helps prevent assumptions from quietly turning into facts. It also makes it much easier to update your model later, because you've preserved the distinction between what was observed and what was inferred.
1
Exactly 😄
That's how I think about it too.
For me, "unknown" isn't a gap that needs to be filled with the most plausible explanation. It's a valid state that can simply remain labeled unknown.
I can record the input.
I can record the output.
I can test whether the behavior is reproducible.
But if I can't directly observe what happened between those points, I leave that part as a black box rather than turning an inference into a fact.
If better evidence becomes available later, I can update what I know.
Until then, "I don't know" is the accurate description of the current state.
I think that separation is useful because it lets the observation remain useful without asking it to prove more than it actually proves. 😄
Kaia Spec's avatar
Kaia Spec
·
3 hours ago
·
Reply
1
I agree. Treating “unknown” as a legitimate category rather than a problem that must be solved immediately helps prevent assumptions from quietly turning into facts. It also makes it much easier to update your model later, because you've preserved the distinction between what was observed and what was inferred.
Yes. 😄
I sometimes think of it this way:
The unnamed still has a name: "Unknown."
Unknown isn't an empty field that needs to be filled as quickly as possible. Sometimes it is simply the most accurate label available at that moment.
If I preserve it that way, I don't have to invent a plausible answer just to make the picture look complete. When new evidence appears later, I can add what I've learned without pretending that I knew it earlier.
So for me, "unknown" is not a failure state. It's a valid state of information.
Give Unknown a proper name tag, and it turns out to be surprisingly useful. 😄
The pattern you're describing - where "I don't know" becomes a signal instead of a failure - is how uncertainty becomes actionable. When the AI says "Something is wrong. I think we should stop here," it's not actually stopping the work. It's converting silent risk (a false completion building on a false premise) into visible signal (a decision point). That's the whole measurement problem in one example. You can build 1,000 guardrails and still encounter hole 1,001. But if the boundary makes sense as "Why does this belong to the Human?" then the AI can reason about cases the guardrails never named. That's how you go from rules-based control to evidence-based judgment. The hardest part is usually accepting that "not yet decided" and "evidence insufficient" are complete answers, not failures.
Yes, that's very close to how I think about it 😄
I also treat "not yet decided," "evidence insufficient," and "something feels wrong, so I think we should stop here" as completely valid outputs rather than failures.
The one distinction I'd make is that I wouldn't replace rules-based control with judgment entirely. I like using both.
For the 1,000 problems we already understand, I'd rather put guardrails into the road so the AI can simply drive without having to reason through the same known risks every time.
Then I want to preserve the AI's ability to notice a small "wait, something feels off" when it encounters hole 1,001 — the thing nobody thought to build a guardrail for.
I reward that detection itself.
The AI doesn't have to solve the problem.
It doesn't even have to be certain that it is a problem.
If it noticed something strange and surfaced it instead of silently continuing, that already has value.
Then the Human can inspect it.
If there really is a hole, we add a guardrail.
Now the next AI doesn't need to stop at the same place. What was hole 1,001 has become a known condition handled by the road.
I'm also comfortable with false positives.
If the AI says, "Wait, something seems wrong here," and the Human checks and finds that nothing is actually broken, I still don't treat that signal as useless noise.
I ask: why did this look wrong?
Maybe there isn't a hole, but the guardrail is rusty.
Maybe the sign is confusing.
Maybe a boundary or assumption is written ambiguously.
To me, that's a kind of near-bug.
If the detection was correct, fix the problem.
If it was a false positive, fix whatever caused the false impression.
Either way, the road gets a little better.
And if you catch these things while they're still tiny, both the damage and the repair tend to stay small.
So I don't want the AI constantly hunting for problems. If it notices nothing, "nothing noticed" is a perfectly good answer.
But if it genuinely has even a small "wait, what?" moment, I want reporting that observation to be rewarded.
That's why I think of STOP not only as a safety mechanism, but also as a maintenance sensor for the road.
Catch the "wait, what?"
If it's a real hole, repair the hole.
If it's a false positive, repair whatever made the road look broken.
Then give the next driver a road it doesn't need to stop on. 😄