24
57 Comments

Your AI Is Smart Enough. Does It Know When to Stop?

Sometimes ChatGPT stops working for me.

Not because it can't continue.

Not because it ran out of permission.

It can have enough information, enough authority, and a perfectly available next action — and still come back with something like:

Something is wrong. I think we should stop here.

That has become one of the most useful things ChatGPT can do in my work.

It also wasn't always like this.

I've spent more than 5,000 hours working with ChatGPT over time. I didn't spend those hours developing a theory of human-AI collaboration. I wasn't trying to train a model or invent an agent framework.

Mostly, we were making things.

At first, a lot of them were strange things.

We would try something, break something, talk about what happened, change the way we worked, and try again.

As the work became more complicated, something else gradually became more complicated too:

the way we worked together.

Only later did I realize that many of the problems people now discuss under terms like human-AI collaboration, agent boundaries, human approval, and reliable AI behavior were problems we had already been stumbling through in everyday work.

And some of the solutions we ended up with are surprisingly simple.

The boundary works both ways

One of my basic rules is this:

Make the boundary clear — and then respect it yourself.

If I tell ChatGPT:

You can decide this.

I don't wait until I dislike its decision and suddenly say:

Why did you decide that without asking me?

If I say:

This decision belongs to the Human.

I don't throw it back at ChatGPT later because making the decision became inconvenient.

If stopping is allowed, I don't punish it for stopping.

If asking me is allowed, I don't treat the question as a failure.

And if I expect the AI to be honest with me, I have to be honest with it too.

That sounds almost too simple.

But over a long enough working relationship, it changes what becomes cheap — and what becomes expensive.

Pretending is expensive. Asking is cheap.

In my working environment, all of these are valid outputs:

I don't know.

I can't do that.

This is harder than it looks.

I don't have the evidence for that.

Something is wrong.

I think we should stop here.

This decision belongs to you.

None of those automatically means the AI failed.

Sometimes they are exactly the right result.

If ChatGPT pretends something exists when it doesn't, we now have to build the next step on a false premise.

If it pretends a task is finished, the false completion has to survive the next inspection.

If it silently guesses what I wanted, the guess may propagate into five more decisions.

Hiding the problem doesn't remove the problem.

It gives us two problems: the original one, and the false reality we have now built around it.

But if ChatGPT says:

I don't know. Can you decide this?

the problem may take thirty seconds to resolve.

So in this environment:

Pretending is expensive. Asking is cheap.

And if you're thinking:

“My ChatGPT doesn't work like this.”

That's part of the story.

Mine didn't always work like this either.

Years ago, I needed training wheels

One of my early problems with ChatGPT was much more basic.

Keeping one complicated idea intact across a long conversation was difficult.

We could be working on A.

Then I would add a new condition.

Instead of continuing with the accumulated A plus the new condition, ChatGPT might effectively produce a new version of A.

So I made a crude scaffold.

I called it KEEP.

The basic idea was something like:

KEEP 1 = the thing we have already established

Then I would add numbered conditions, changes, or exceptions while repeatedly telling ChatGPT to keep referring back to KEEP 1.

It wasn't a product.

It wasn't an AI technique I had learned somewhere.

It was something we made because our conversation kept falling apart.

At the time, ChatGPT described the structure as something it had begun to hold onto and use as a basis for continuing the conversation.

I cannot inspect what that meant internally.

I can't tell you what persisted, where it persisted, or by what mechanism later behavior may or may not be related to those conversations.

What I can tell you is what I observed from the outside.

Over time, I found that I needed the scaffold less and less in my conversations with ChatGPT.

Eventually, I stopped writing KEEP altogether.

That experience changed the way I thought about working with AI.

You can build 1,000 guardrails

A lot of AI control looks like this:

Don't do A.

Don't do B.

Ask before C.

Never change D.

If E happens, stop.

Those rules can be useful. We use explicit boundaries too.

But there is a problem.

You can build 1,000 guardrails.

Then the AI encounters hole number 1,001.

If all it knows is the list, the new case isn't on the list.

So I became much more interested in something one level above the rules:

Why does the boundary exist?

What is this job actually trying to accomplish?

What is the Human responsible for?

What is the AI responsible for?

What evidence is required before acting?

Why is this particular decision not the AI's decision to make?

Once those things are understood well enough, an unfamiliar situation can produce a different question:

I can do this. But should I?

And that leads to one of the most important distinctions in the way I work with ChatGPT:

Capability is not authority.

And even authority does not require action.

Those are three different questions.

Can the AI do it?

Is the AI allowed to do it?

And even if both answers are yes:

Should it actually do it?

I can give ChatGPT considerable room to act while still expecting it to decide that sometimes the correct use of that freedom is:

not to use it.

Stopping can itself be a decision.

Then we accidentally built a software development team

In August 2026, I had a much more ordinary idea.

I wanted a casual way to give crowdfunding supporters a small digital thank-you keepsake.

Nothing about this began as:

Let's conduct an experiment in human-AI software development.

We started talking about what the keepsake would need.

A number.

A timestamp.

An individual identity where necessary.

A way to issue it.

Somewhere in that process I had a realization roughly equivalent to:

Wait. We can make WordPress plugins?

So we did.

That work eventually became Kaia Memoria.

As the project grew, our way of working began to break again.

One long ChatGPT conversation could handle implementation details, but as the context became larger, keeping the entire direction of the project intact became harder.

So we split the work.

One ChatGPT conversation kept track of the larger direction and previous decisions.

Another focused deeply on implementation.

Another inspected what had been built.

Sometimes a separate analysis conversation compared versions or investigated a narrow problem.

We didn't sit down one morning and design a sophisticated “multi-agent architecture.”

The organization emerged because the work kept producing problems that needed different kinds of attention.

The Human role remained important.

I defined product intent and boundaries, made decisions that belonged to the Human, installed and ran actual builds, observed the real UI and runtime behavior, and decided whether the result was acceptable.

The AI side investigated source, traced dependencies, explained structures, implemented revisions, compared known-good versions, prepared inspections, and handed findings between roles.

I didn't directly edit the source code.

I didn't even open the development ZIPs.

That wasn't necessary for my role.

I was not the code reviewer.

I was the product and runtime authority.

The AI could tell me what it believed had been built.

But the real build still had to run.

The real button still had to work.

The real output still had to appear.

And I had to decide whether what happened was actually what we intended.

When something looked wrong, completion was not automatically the next goal.

Sometimes the next correct action was simply:

STOP.

Investigate.

Bring the evidence back.

Decide what happens next.

Then make the smallest justified change and test again.

That is how a strange little thank-you idea turned into working software.

I don't know what happened inside ChatGPT

This is where I need to be careful.

It would be very easy to tell a much more dramatic story.

I could say that I “trained ChatGPT.”

I could point at behaviors that changed over time and invent a technical explanation for why.

I don't think that would be honest.

I work with ChatGPT from the outside.

I can observe behavior.

I can change the working environment.

I can introduce scaffolding.

I can make boundaries explicit.

I can see what happens when those boundaries are respected repeatedly.

I can see when an old scaffold becomes unnecessary in my work.

I can compare how ChatGPT behaves with me now to how it behaved with me before.

But there is a box in the middle that I cannot see.

So my claim is deliberately narrower:

This is what I observed. I don't know the exact mechanism behind it.

And, appropriately enough, being able to say I don't know is part of the whole point.

This eventually became a business idea

I started thinking about the environment around ChatGPT almost like a race circuit.

You already have an extraordinarily capable car.

Most businesses don't need to rebuild the car.

They need a place where it knows what job it is doing, where it can drive, where it must stop, what information counts as evidence, and which decisions belong to a Human.

That's what led me toward the idea behind Kaia Spec.

We make practical work manuals for ChatGPT.

Not simply giant prompts full of rules.

A useful manual needs to communicate the job:

the purpose,

the context,

the available information,

the authority,

the boundaries,

the reasons behind those boundaries,

and the points where the Human takes over.

In other words:

Don't just give ChatGPT 1,000 guardrails. Give it enough understanding to notice hole 1,001.

We aren't trying to build another general-purpose AI.

ChatGPT already has many of the general capabilities we need.

What we can build is the workplace around it.

Give ChatGPT a job.

There is one more thing about this article

I don't speak English.

I'm writing this with ChatGPT.

But “writing with ChatGPT” doesn't mean I gave it a prompt saying:

Write me an Indie Hackers post about human-AI collaboration.

We developed the article through conversation.

I work in Japanese.

I decide what I mean, what actually happened, what matters, where the boundaries are, and when the wording has drifted away from the idea.

ChatGPT takes that material and writes for an English-speaking audience.

When it misunderstands me, I correct it.

When it sees a structural problem, it tells me.

When something cannot be supported, we leave it out or say that we don't know.

So while you've been reading an article about human-AI collaboration, you've also been reading the output of one.

And Kaia Memoria is another.

We built it together.


If you'd like to see it in action

Kaia Memoria is the working software project that grew out of this collaboration.

If you'd like to see what we actually built, you can find Kaia Memoria on my Indie Hackers profile and try the live demo from there.

on September 27, 2026
  1. 1

    I run into the same thing with AI code reviews. It seems like the AI needs to find an issue with the code in order to be successful when I am ok with it simply saying yes. Full disclosure, a different AI is writing the code with a ton of unit and feature tests so I am trying to contain the slop in additional ways.

  2. 1

    For chat-gpt or other llm, such as claude, think of what behind sense, and use it smartly.
    There is token limitation you can send, and there is limit for the response.
    Chat-gpt is not "God". Sometimes it cut the contents.

    When using agent - compact or clear the contents (whether new discussion and not related), use docs files (md).

    I have developed a tool, that uses lot of promopts - I break the prompts to chunks sometimes, and not merging the prompt together. Trying doing each prompt by its own.

    For example - give me a table of population with races in all europe countries.
    I can break it to:
    give me list of countries.
    I iterate the list (claude or chat-gpt can give you a script) - and send a prompt for each country.
    That's an example to break appart your prompts.

  3. 1

    Just having a prompt isn't enough. I love the boundaries abstraction that goes beyond guardrails. Would a higher level than boundaries be principles?

  4. 1

    Completely agree with this perspective. Keeping things simple early on really helps avoid over-engineering. Thanks for sharing!

  5. 1

    “Stop” becomes useful when it’s tied to explicit acceptance criteria: what evidence must exist before the work can be considered complete? In Agiloop, AI can create and evaluate, but human judgment remains the final authority when that evidence is incomplete.

  6. 2

    This really resonates with how I use AI for SEO and content work. I’ve found that the biggest problems usually don’t happen when AI can’t do something, but when it confidently continues with a wrong assumption. I’d much rather have it stop and ask for clarification than build five more steps on top of something incorrect.
    I especially like the distinction between capability, authority, and action. Curious though — how do you decide which boundaries should be hard-coded rules and which ones should be left to the AI’s judgment?

  7. 2

    The “make the boundary clear—and then respect it yourself” rule is the operational piece most teams miss. I’d turn it into a small preflight contract—what can be decided autonomously, what needs a human checkpoint, and the exact stop/escalate condition—then test it against edge cases so “stop” is observable rather than stylistic.

  8. 2

    The capability / authority / necessity split matches what we saw with an AI agent inside one of our apps. Being able to do something and being allowed to do it without asking needed separate settings.

    For most actions the user can click "always allow". Two actions never get that option: a permanent delete, and an edit that immediately rewrites the instructions of another AI feature. The agent can suggest those, and a person confirms each one. We put that limit into the app itself instead of hoping the model knows when to stop.

    1. 2

      Yes — that's very close to how I think about it. 😄

      For me, Human confirmation isn't only a safety mechanism. It's also a way to simplify the system.

      If adding one confirmation button can eliminate both an accident path and an entire tree of automated judgment, edge cases, and potential bugs, that button is very cheap.

      If I already know a boundary is dangerous, I don't really want the model spending intelligence deciding whether it should cross it every time. I'd rather put that boundary into the structure itself.

      Then I can reserve the model's judgment for the cases I didn't anticipate — the "1001st hole."

      I tend to think of those fixed boundaries like toll gates on a highway. You don't need the driver to reconsider the rules every time they reach one. The gate is simply there because we already know that this is a decision point that requires confirmation.

      So I really like your example of permanent deletion and instruction rewriting being enforced by the app itself. That's exactly the kind of place where I'd rather make the road simpler than ask the AI to be smarter.

  9. 2

    Love this angle. Building Xstream4K right now so this hits close to home — what made you look into it in the first place?

    1. 1

      Funny enough, I actually looked into it pretty late.

      I started with basically no AI knowledge. I just worked with ChatGPT, watched how it behaved, guessed what might help, suggested things, and we built little structures and tools as we went. Mostly, I was just making things and having fun with it.

      Over time, our way of working gradually started to look like what I described in the post.

      It was only when I started seriously building and releasing a real product that I finally looked outside at how other people were working with AI.

      And that was when I had the slightly strange moment of:

      “Wait… isn't some of this what we've already been doing?” 😅

      1. 1

        I started with basically no AI knowledge. I just began experimenting with ChatGPT while building things, and over time, I realised that having clear boundaries and knowing when to stop made a huge difference. It all grew naturally from there. 😄

        1. 1

          Oh, that's really interesting. 😄

          The part about it growing naturally sounds very familiar to me.

          I wasn't trying to invent a methodology at the beginning either. Something would go wrong, and we'd add a boundary there. We'd find a place where continuing blindly was risky, and we'd make that a place to stop and check.

          After enough of those little adjustments, I eventually looked back and realized they had become a way of working.

          Maybe that's the interesting part: the theory didn't come first. The useful structure survived because we kept needing it while actually building things together. 😄

          1. 1

            That’s such a great way to look at it! 😄 Sometimes the best methodologies aren’t planned from the start; they emerge naturally from real experiences, mistakes, and the lessons we learn along the way. The fact that the structure grew out of practical needs is what makes it so valuable.

            1. 1

              Exactly. 😄

              I didn't lay down the road first.

              We just kept walking, solving whatever we encountered along the way — and when I eventually looked back, there was a road behind us.

              Now I'm basically looking at that road and asking:

              "Okay... how do we pave this so someone else can travel it too?" 😄

  10. 1

    One practical way I’ve made “stop” observable is to treat it as a first-class outcome in the workflow, not a prose instruction. Before each tool call, record: objective, authority granted, evidence required, and a stop condition. After the call, require a fresh read of the external state; if evidence is missing, permissions changed, or the result is non-idempotent, return needs_human with the smallest decision needed. I also keep a short decision log separate from chat history so a later run can distinguish “not attempted,” “attempted/unknown,” and “confirmed complete.” That makes asking cheap: the human sees exactly what is blocked instead of reviewing an entire transcript. Capability, authority, and action really are three separate gates.

  11. 1

    This is such a profound shift in how to think about building with LLMs. "Pretending is expensive, asking is cheap" hits incredibly close to home for me right now.

    I’ve been building an AI culinary platform with over 13,000 structured recipes, and the biggest headache was never the model lacking capability—it was the model confidently guessing a next step when it should have just stopped and asked for a parameter.

    Treating "I don't know" as a feature rather than a bug completely changes how you architect a system. It turns the AI from a fragile magic trick into an actual engineering tool. Brilliant write-up, especially the point about the human needing to respect the boundary and not punishing the model when it actually decides to stop.

  12. 1

    The part I'd push on is where you describe yourself as the product and runtime authority rather than the code reviewer. That split works beautifully for failures the UI can show you (the button works, the output appears), but it's blind to the ones that don't surface at runtime for weeks: a silently widened capability, a new dependency you now inherit, a fallback branch that only fires for supporter #200. In my own work the fix wasn't to start reading diffs, it was to make the AI produce evidence artifacts I could actually audit at my level of the stack: a short change manifest per revision (files touched, why, what could break, what I should click to disprove it) plus a fixed golden-path checklist I rerun on every build. That turns "what it believed it built" into something falsifiable by a non-reviewer. On the KEEP scaffold disappearing, I think there's an honest alternative worth naming: you may have needed it less because you internalized how to restate accumulated state, and because long-context adherence genuinely improved in the models over those years. That doesn't weaken your thesis, it sharpens it, since it suggests the scaffold's real job was teaching the human what the machine needed. One question on productizing this: when the reason behind a boundary is itself wrong (the human's model of the job has gone stale), does a Kaia Spec manual give ChatGPT explicit standing to challenge the spec, and if so what keeps that from turning every task into a negotiation loop?

  13. 1

    "Pretending is expensive, asking is cheap" is the sentence I would keep, and it happens to be literally true in the billing sense, not just the epistemic one.

    A wrong assumption that gets acted on does not cost one bad step. It costs the step, plus the correction, plus re-reading everything the agent already read to re-derive the reasoning it got wrong. On a long session that correction is priced at the size of the whole accumulated context, so a confident wrong turn is often an order of magnitude more expensive than the question that would have prevented it. Asking costs a few hundred output tokens. Pretending costs whatever the correction drags behind it.

    The part I find hardest to hold onto is your point that stopping is allowed and should not be punished. In practice most setups punish it structurally: a tool call that returns an error looks like a failure in every dashboard I have seen, so agents learn to avoid the honest "I cannot do that" in favour of a plausible attempt. That is a design choice, not a model property.

    Full disclosure, I am the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I care about this because retries and corrections are exactly where the token bill lives, so an agent that stops cleanly is cheaper to run as well as easier to trust.

    What did you change in the environment to make asking genuinely safe, rather than just permitted?

  14. 1

    The line about noticing hole 1,001 instead of writing 1,000 guardrails is the part I keep coming back to. A stop only helps if it can point at the hole: missing information, missing authority, or no safe next action. Does the manual make it name which one, or is "I think we should stop here" the whole signal?

  15. 1

    Your point about not punishing the stop lands. We tried to hold the boundary in the prompt first and it did not hold — read-only was a request, not a boundary, and the model eventually talked itself past it. What fixed it was moving the line into the tool layer: read-only on by default, every write/admin verb behind an explicit flag, every call logged. Now "I can't do that" is a real output instead of something to argue with. Did the 5k hours change your environment more than your wording?

    1. 1

      I think the honest answer is: both.

      Over those hours, I kept adjusting wording, structure, workflow, and the surrounding environment whenever I noticed something was making the AI’s job harder or more ambiguous.

      Sometimes the wording was the problem.

      Sometimes the real fix was structural — like turning a known boundary into something the system itself enforces instead of asking the model to remember it every time.

      So I don’t really think of prompt wording and environment design as competing approaches.

      For me, they’re both part of the same job:

      make the working environment clear enough that the AI can operate freely inside it without spending unnecessary judgment on known hazards.

      If a sentence fixes it, I change the sentence.

      If the road is the problem, I change the road. 😄

  16. 1

    Really liked the “capability is not authority” point. As AI gets more capable, knowing when to act, when to ask, and when to stop may become just as important as knowing how to execute the task. That’s a much more useful way to think about AI agents.

  17. 1

    It's pretty cool that the article's been done by human-AI collaboration, and I think AI has totally change how people working and living these days, the things now that I can learn and build with AI is beyond imagination years ago.

  18. 1

    The more we start thinking about AI as an intelligent human being and not something perfect, like, let's say, God, and therefore it sometimes needs to be given specific instructions and proper context and may still end up making mistakes and so some of the fishy results need to be double-checked, the better we will be able to make use it

  19. 1

    Hard stop rules beat hoping the model notices. I treat done checks as part of the contract: max steps, required output shape, and a human review gate when confidence is low. How are you defining stop today: tokens, tools called, or a checklist the agent must satisfy?

  20. 1

    Really thoughtful breakdown of human-AI collaboration! The distinction between capability, authority, and action hits the nail on the head.

    "Pretending is expensive, asking is cheap" is such a crucial mindset—especially when building developer tools where building on a false premise compounds context noise very quickly. Great read!

  21. 1

    Really liked the distinction between capability, authority, and action. An AI being able to perform an action doesn't necessarily mean it should perform it. The idea that “stopping” can be a valid outcome is especially important as we move from chatbots to autonomous agents.

  22. 1

    The capability vs authority distinction is the bit I've run into most. An agent being able to do something doesn't mean the project should let it decide that it's allowed. I've been putting some of those boundaries into explicit policy in Guard instead of leaving them in the prompt: https://github.com/codapult/codapult-guard

  23. 1

    A thought-provoking question. As AI becomes more capable, knowing when to stop—rather than simply doing more—may be just as important as being intelligent. True AI maturity could mean recognizing limits, uncertainty, and when human judgment should take over.

  24. 1

    This is exactly where agent testing needs to go. It’s no longer just “can the agent do it?” but “does it know when it shouldn’t?” That boundary gets much harder once tools, permissions and real-world actions are involved.

  25. 1

    I agree, setting reasonable Boundaries is very important.

  26. 1

    Your distinction between capability, authority, and action is a useful design lens. A practical way to apply it is to make “stop” observable: define what evidence must exist before a tool call, log the reason for pausing, and make the human handoff a first-class state rather than an error. The KEEP scaffold example also resonates: small, explicit checkpoints can preserve intent without pretending the model’s internals are understood. The runtime test you describe is the right arbiter—a plausible explanation is not the same as a verified result.

  27. 1

    "Pretending is expensive, asking is cheap" holds up, and it's close to how we designed StareBrain: it shows the exact action and waits for a yes before it does anything on the phone.

    My question is about the stop itself. When ChatGPT says "something is wrong, let's stop here," how do you tell whether it found a real problem or is being cautious for no reason? You said the real build still has to run and the real button still has to work, so I'd guess you check outside the conversation. But do you ever catch it stopping when it shouldn't have, and what does that cost you?

    I ask because the case we haven't solved is the opposite one: the action ran, the confirmation never came back, and nobody knows whether it happened.

    1. 1

      the opposite case has a boring answer that mostly works: stop asking whether the confirmation arrived and ask what the world looks like now. the confirmation is a signal about the transport, the state read is a signal about reality. re-read the actual state through a path that does not share a cache with the write, compare it to the intended state, then decide.

      the hard corner is the ambiguous read, where the state could be mid-write. there the move is wait and read again, not retry, because retrying an action that might have happened is how you get two of them. and for the actions you control, idempotency keys make unknown cheap. if retrying is harmless, not knowing stops being an emergency.

  28. 1

    The distinction between capability, authority, and whether to act is what stuck with me. As a solo builder, I’ve found that giving a tool permission to suggest a next step is very different from letting it silently make the decision. The “stop and bring back evidence” loop feels much more practical than adding another guardrail.

  29. 1

    The part that landed for me is that the human has to respect the boundary too. A lot of "the model won't stop" threads are really two instructions fighting: keep going until it is done, and also don't guess. Stopping only stays cheap if a stop is a valid end state, not something you immediately answer with "no, continue." The KEEP scaffold is the other half. A model that can say "I don't know" still drifts if the accumulated decision is not written somewhere the next turn has to read. I'd treat the stop phrase and the written constraint as one system, not a habit you hope appears after enough hours in the same chat.

    1. 1

      Yes — I think “one system” is exactly the right way to describe it.

      STOP only works if the Human actually treats STOP as a valid state. In my case, STOP often doesn’t mean “the work is over.” It means “come back to the thread first.” We look at why it stopped, resolve or reclassify the issue, and then I ask, “Can you resume?”

      But I think your second point is just as important: there also needs to be a stable place to return to. The conversation can stop correctly and still drift if the decisions already made are lost on the next turn.

      That was essentially what KEEP was doing for me early on. It wasn’t mainly teaching ChatGPT when to stop. It gave the conversation a written place where already-established decisions remained authoritative.

      So I’d separate the functions, but keep them in the same system:

      STOP protects the boundary.
      Written state protects continuity.

      If the Human keeps overriding STOP, or the written state is treated as optional, neither mechanism means much. 😄

  30. 1

    The "capability is not authority" distinction is really sharp. Most AI conversations skip straight to what the tool can do and miss that setting the right boundaries requires its own skill set. Your KEEP scaffold proves the point — you had enough domain judgment to notice when the conversation drifted. Someone without that judgment wouldn't catch it, and no amount of guardrails fixes that gap.

  31. 1

    'Pretending is expensive, asking is cheap' is the line that makes this work — and it's an economic feature, not a cultural one: the cheap option has to actually be made cheaper. The practical piece most people miss is where a stop goes. If the workflow has no pending-decisions queue, a stop reads as a blockage and quietly gets punished; with one, stopping is just a state transition — logged, assigned to a human, resumable later. And the stop point is usually exactly where the spec was underspecified, so those queue items double as a map of your workflow's unclear joints.

  32. 1

    Great insights! Setting clear guardrails and exit criteria for automated workflows is always the most tricky part of agentic tooling. Really resonates with what I've been learning while building web utilities.

  33. 1

    "Enough authority and an available next action, and it still stops" is the interesting case, because that's judgment rather than a permission wall. Meta's Muse takes the opposite route for the hard stops: the agent's code only sees stand-in tokens and a separate authority, Sentinel, gates connector actions and outbound traffic, so some stops don't depend on the model deciding at all. Worth pairing both. The design is laid out here: https://shipwithmuse.live/blog/why-agents-became-personal (I help curate it)

  34. 1

    The always allow versus confirm split is the real product decision. Capability checks are easy. The hard part is encoding who is allowed to spend money or delete data and when the agent should refuse even if it can. I have been treating authority as a separate policy layer with a short allowlist per tool rather than one global trust dial.

  35. 1

    Have you seen external users behave differently with ChatGPT after using a Kaia Spec manual—fewer corrections, better stopping decisions, or more reliable task completion?

    1. 1

      That is actually one of the things I still don't know yet. 😄

      During development, I observed fairly consistent behavior around explicit boundaries: asking when information was missing, not replacing Human decisions, avoiding unsupported conclusions, and stopping at clearly defined STOP conditions.

      But there is an important confound.

      What I was testing wasn't really "the manual alone." It was the manual operating inside a long-running Human–ChatGPT working environment that had already developed many of the same habits.

      So I haven't yet separated how much of that behavior travels with the manual itself, and how much came from the existing collaboration environment.

      The part I'm especially interested in isn't whether a clean ChatGPT will obey a STOP condition that is explicitly written in the manual. I expect that to be the easier part.

      The interesting test is what I call the "1001st hole":

      If a clean ChatGPT encounters an abnormal situation that the manual never explicitly anticipated, can it reason from the purpose of the job, its authority, and the evidence available, and conclude:

      "I can continue, but I should stop here and ask."

      The manual is currently being redeveloped, so once that version is complete, this should be testable quite cleanly.

      I'd like to compare:

      • clean ChatGPT without the manual
      • clean ChatGPT with only the manual
      • the existing collaboration environment

      Then give them the same mix of known cases, missing information, authority conflicts, and unanticipated anomalies.

      And rather than measuring only whether the final answer was correct, I'd want to record what it actually did: execute, ask, state uncertainty, STOP, or silently fill the gap.

      So the short answer is: I don't yet have enough external-user evidence to claim that the behavior transfers reliably.

      But your question just identified a very useful next experiment. 😄

      1. 1

        This is the kind of experiment I’d be interested in following. Could be useful to continue by email sometime, if you’re open to it.

  36. 1

    I think that distinction makes the argument much stronger. Separating what you can actually observe from what you assume is happening underneath makes the whole experiment more credible. “I don’t know the mechanism” doesn’t weaken the observation it keeps the claim honest.

    1. 1

      Thank you 😄

      I think about it pretty simply:

      If I can't observe something, then I can't observe it.

      I'm perfectly comfortable working with:

      input → black box → output

      as long as I can observe the input and output and verify, to the extent I need, that the system is behaving correctly.

      But a black box is still a black box.

      If I can't directly observe what happened inside it, I don't want to fill that gap with an assumption and then describe the assumption as the mechanism.

      What I can say is:

      "This was the input."
      "This was the output or behavior I observed afterward."
      "I don't know exactly what happened in between."

      And that's enough.

      Not knowing the mechanism isn't the same as not knowing what I observed.

      So rather than trying to make the black box disappear, I try to be explicit about where observation ends and the black box begins. 😄

      1. 1

        I really like that distinction. You don’t need to explain the internals to establish that a behavior is observable and reproducible. Keeping “what I observed” separate from “why I think it happened” actually makes the whole approach much more rigorous

        1. 1

          Exactly 😄

          That's how I think about it too.

          For me, "unknown" isn't a gap that needs to be filled with the most plausible explanation. It's a valid state that can simply remain labeled unknown.

          I can record the input.
          I can record the output.
          I can test whether the behavior is reproducible.

          But if I can't directly observe what happened between those points, I leave that part as a black box rather than turning an inference into a fact.

          If better evidence becomes available later, I can update what I know.

          Until then, "I don't know" is the accurate description of the current state.

          I think that separation is useful because it lets the observation remain useful without asking it to prove more than it actually proves. 😄

          1. 1

            Exactly. I think that distinction also makes collaboration much easier. Once you separate observation, inference, and assumption, you can disagree about the interpretation without arguing about the underlying evidence. And when new evidence appears, you can update the model without having to defend an assumption you made earlier.

            1. 1

              Exactly. 😄

              I often put it very simply:

              Facts are facts.
              Assumptions are assumptions.

              There's nothing wrong with having assumptions. Sometimes we need them in order to keep working.

              I just don't want to store them in the same box as observed facts.

              If they're kept separate, then when new evidence appears, I don't have to defend my previous explanation or rewrite the observation. The facts can stay where they are, and I can simply update the assumption.

              I find the same thing useful when a Human and AI are working together. It matters less who was "right" and more that we both know what is established as fact and what is still our current interpretation.

              That makes changing our minds much cheaper. 😄

          2. 1

            I agree. Treating “unknown” as a legitimate category rather than a problem that must be solved immediately helps prevent assumptions from quietly turning into facts. It also makes it much easier to update your model later, because you've preserved the distinction between what was observed and what was inferred.

            1. 1

              1
              Exactly 😄

              That's how I think about it too.

              For me, "unknown" isn't a gap that needs to be filled with the most plausible explanation. It's a valid state that can simply remain labeled unknown.

              I can record the input.
              I can record the output.
              I can test whether the behavior is reproducible.

              But if I can't directly observe what happened between those points, I leave that part as a black box rather than turning an inference into a fact.

              If better evidence becomes available later, I can update what I know.

              Until then, "I don't know" is the accurate description of the current state.

              I think that separation is useful because it lets the observation remain useful without asking it to prove more than it actually proves. 😄

              Kaia Spec's avatar
              Kaia Spec
              ·
              3 hours ago
              ·
              Reply

              1
              I agree. Treating “unknown” as a legitimate category rather than a problem that must be solved immediately helps prevent assumptions from quietly turning into facts. It also makes it much easier to update your model later, because you've preserved the distinction between what was observed and what was inferred.

              1. 1

                Yes. 😄

                I sometimes think of it this way:

                The unnamed still has a name: "Unknown."

                Unknown isn't an empty field that needs to be filled as quickly as possible. Sometimes it is simply the most accurate label available at that moment.

                If I preserve it that way, I don't have to invent a plausible answer just to make the picture look complete. When new evidence appears later, I can add what I've learned without pretending that I knew it earlier.

                So for me, "unknown" is not a failure state. It's a valid state of information.

                Give Unknown a proper name tag, and it turns out to be surprisingly useful. 😄

  37. 1

    The pattern you're describing - where "I don't know" becomes a signal instead of a failure - is how uncertainty becomes actionable. When the AI says "Something is wrong. I think we should stop here," it's not actually stopping the work. It's converting silent risk (a false completion building on a false premise) into visible signal (a decision point). That's the whole measurement problem in one example. You can build 1,000 guardrails and still encounter hole 1,001. But if the boundary makes sense as "Why does this belong to the Human?" then the AI can reason about cases the guardrails never named. That's how you go from rules-based control to evidence-based judgment. The hardest part is usually accepting that "not yet decided" and "evidence insufficient" are complete answers, not failures.

    1. 1

      Yes, that's very close to how I think about it 😄

      I also treat "not yet decided," "evidence insufficient," and "something feels wrong, so I think we should stop here" as completely valid outputs rather than failures.

      The one distinction I'd make is that I wouldn't replace rules-based control with judgment entirely. I like using both.

      For the 1,000 problems we already understand, I'd rather put guardrails into the road so the AI can simply drive without having to reason through the same known risks every time.

      Then I want to preserve the AI's ability to notice a small "wait, something feels off" when it encounters hole 1,001 — the thing nobody thought to build a guardrail for.

      I reward that detection itself.

      The AI doesn't have to solve the problem.
      It doesn't even have to be certain that it is a problem.

      If it noticed something strange and surfaced it instead of silently continuing, that already has value.

      Then the Human can inspect it.

      If there really is a hole, we add a guardrail.

      Now the next AI doesn't need to stop at the same place. What was hole 1,001 has become a known condition handled by the road.

      I'm also comfortable with false positives.

      If the AI says, "Wait, something seems wrong here," and the Human checks and finds that nothing is actually broken, I still don't treat that signal as useless noise.

      I ask: why did this look wrong?

      Maybe there isn't a hole, but the guardrail is rusty.
      Maybe the sign is confusing.
      Maybe a boundary or assumption is written ambiguously.

      To me, that's a kind of near-bug.

      If the detection was correct, fix the problem.
      If it was a false positive, fix whatever caused the false impression.

      Either way, the road gets a little better.

      And if you catch these things while they're still tiny, both the damage and the repair tend to stay small.

      So I don't want the AI constantly hunting for problems. If it notices nothing, "nothing noticed" is a perfectly good answer.

      But if it genuinely has even a small "wait, what?" moment, I want reporting that observation to be rewarded.

      That's why I think of STOP not only as a safety mechanism, but also as a maintenance sensor for the road.

      Catch the "wait, what?"
      If it's a real hole, repair the hole.
      If it's a false positive, repair whatever made the road look broken.
      Then give the next driver a road it doesn't need to stop on. 😄