2
6 Comments

NAEOS Technical Build Log #004 What Happens When Execution Becomes Unknown?

A policy check can tell us whether an action is allowed.

But what happens after the action is authorized—and the runtime stops telling us what happened?

This sounds like a simple timeout problem.

It isn't.

For an AI coding agent, an execution timeout can create a much more dangerous state:

we don't know whether the side effect happened.

And that distinction matters.


The scenario

Consider a simple deployment:

Agent proposes deployment
        ↓
Policy evaluation
        ↓
Authorization granted
        ↓
Runtime sends deployment request
        ↓
Network timeout
        ↓
???

What is the correct state?

Did the deployment fail?

Did it succeed?

Did the provider receive the request but fail to return the response?

We don't actually know.

Treating the timeout as FAILED would be an assumption.

Treating it as SUCCESS would also be an assumption.

So NAEOS needs another state.

UNKNOWN

UNKNOWN is not FAILURE

This is the distinction I'm exploring in this build log.

A useful execution state machine looks more like:

PROPOSED
   ↓
AUTHORIZED
   ↓
EXECUTING
   ├──→ CONFIRMED
   ├──→ REJECTED
   └──→ UNKNOWN

CONFIRMED means we have sufficient evidence that the side effect occurred.

REJECTED means we have evidence that the action did not occur or was explicitly rejected.

UNKNOWN means the runtime no longer has sufficient observation to determine the outcome.

That last state is uncomfortable.

But pretending uncertainty does not exist is worse.


Why automatic retry can be dangerous

Imagine an agent is authorized to deploy an application.

The runtime sends the request.

The provider processes it.

The deployment succeeds.

But the response is lost.

From the runtime's perspective:

request → timeout

The agent sees:

UNKNOWN

If the agent simply retries:

attempt #1 → UNKNOWN
attempt #2 → EXECUTE

we may now have two deployment attempts.

For some operations that may be harmless.

For others, it can be dangerous:

  • database migrations
  • infrastructure provisioning
  • payments
  • external API mutations
  • resource creation
  • production deployments
  • notifications
  • destructive operations

The important question is therefore not:

"Did the request timeout?"

The important question is:

"Can we establish what actually happened?"


Reconciliation before retry

This leads to another primitive I'm testing:

UNKNOWN
   ↓
RECONCILIATION
   ├── evidence confirms execution
   │        ↓
   │     CONFIRMED
   │
   ├── evidence confirms no execution
   │        ↓
   │     REJECTED
   │
   └── insufficient evidence
            ↓
         UNKNOWN

The runtime should attempt to reconcile the unknown state before deciding what happens next.

Depending on the operation, that evidence might include:

  • provider deployment ID
  • runtime receipt
  • commit SHA
  • artifact digest
  • resource state
  • health status
  • audit event
  • external system receipt
  • idempotency key

The exact evidence depends on the execution target.

But the principle remains the same:

Outcome should be established from evidence, not inferred from silence.


Authorization does not establish outcome

This is another distinction that became clearer while working on the previous build log.

A policy decision answers:

May this action happen?

Authorization answers:

Is this execution attempt permitted under the current context?

Execution answers:

What did the runtime attempt to do?

Observation answers:

What actually happened?

These are different questions.

So the model I'm working with now is:

Proposal
   ↓
Policy Evaluation
   ↓
Authorization
   ↓
Execution
   ↓
Observation
   ↓
Outcome

The agent can propose the next action.

It should not be the authority that declares the previous side effect successful simply because the runtime stopped responding.


The agent should not own uncertainty resolution

This is particularly important for AI coding agents.

A model might reason:

"The deployment probably failed. I'll retry."

But "probably" is not an execution receipt.

The governance/runtime layer should instead be able to return something explicit:

{
  "execution_status": "UNKNOWN",
  "authorization_id": "auth_01",
  "attempt_id": "exec_01",
  "reconciliation_required": true
}

Now the agent has information about the state without being given authority to redefine it.

The agent can propose a next action.

The control plane decides whether that action is allowed.

And the runtime determines whether the previous action actually happened.


What happens with non-idempotent operations?

This creates an even harder question.

Suppose the system has:

UNKNOWN

and the operation is not safely idempotent.

Should NAEOS retry automatically?

My current answer is:

No—not without sufficient evidence or an explicit authorization path for the retry.

A safer decision path looks like:

UNKNOWN
   ↓
Can outcome be reconciled?
   ├── YES → reconcile
   └── NO
        ↓
Is operation safely idempotent?
        ├── YES → controlled retry
        └── NO → require explicit decision

The goal is not to make the system incapable of recovering.

The goal is to prevent uncertainty from silently becoming authority.


The experiment

The next NAEOS execution-boundary tests are therefore becoming more explicit.

Test A — Explicit rejection

authorize
→ execute
→ provider rejects

Expected:

REJECTED

Test B — Confirmed execution

authorize
→ execute
→ receipt returned

Expected:

CONFIRMED

Test C — Timeout before side effect

authorize
→ connection lost
→ no side effect

Initial state:

UNKNOWN

Then reconciliation should establish:

REJECTED

Test D — Timeout after side effect

authorize
→ side effect occurs
→ response lost

Initial state:

UNKNOWN

Then reconciliation should establish:

CONFIRMED

Test E — Unknown non-idempotent operation

authorize
→ execute
→ timeout
→ UNKNOWN

Expected:

no automatic retry

unless the system can establish that retry is safe and authorized.


A possible execution receipt

One direction I'm exploring is making execution outcomes durable artifacts rather than transient agent messages.

For example:

{
  "attempt_id": "exec_01",
  "authorization_id": "auth_01",
  "operation": "deploy",
  "target": "production",
  "status": "CONFIRMED",
  "provider_receipt": "deployment_83921",
  "artifact_digest": "sha256:...",
  "observed_at": "...",
  "runtime_version": "...",
  "policy_version": "..."
}

The exact schema is still experimental.

The important part is the relationship:

Authorization
      ↓
Execution Attempt
      ↓
Observation
      ↓
Evidence
      ↓
Outcome

That chain should survive the agent that initiated it.


A broader principle

The deeper lesson here is not really about timeouts.

It is about distributed systems and control boundaries.

A timeout doesn't necessarily tell us that something failed.

It tells us that we lost observation of what happened.

For AI engineering systems, that distinction becomes particularly important because an agent can continue reasoning even when the underlying runtime has entered an uncertain state.

NAEOS should not hide that uncertainty.

It should make it explicit.

When observation is lost, don't manufacture certainty. Reconcile the state.

And if reconciliation cannot establish the outcome:

UNKNOWN

should remain UNKNOWN.


What I'm testing next

The next question is becoming even more interesting:

What evidence is sufficient to move an execution from UNKNOWN to CONFIRMED or REJECTED?

That takes us from execution control into something closer to an evidence and verification model.

And that's where I think the next NAEOS experiment should go.


Question for the builders here

How would you handle a non-idempotent operation when:

  1. authorization was valid,
  2. execution was initiated,
  3. the runtime timed out,
  4. and there is no immediate confirmation of whether the side effect occurred?

Would you reconcile first, require human intervention, use an idempotency mechanism, or take another approach?

I'd especially like to hear from people working on distributed systems, infrastructure, deployment platforms, and AI coding agents.

on September 27, 2026
  1. 1

    We hit a smaller version of this in our own framework, where the AI agent calls the same write handlers as the UI.

    Two things removed most of the UNKNOWN cases for us. Every write carries a request id, and a retry reuses it. The server remembers the id: if the first attempt is still running, the retry waits for it, and if it already finished, the retry gets the stored result and nothing runs twice. Also, every state change is an event appended in the same transaction as the write, so whether it happened is a lookup in the event log.

    It has limits. The server keeps those ids for five minutes, which covers a retry but not someone re-running the job the next day. And none of it helps once the side effect lives in someone else's system. There we only get certainty if their API takes an idempotency key as well. Otherwise we're where you are: find evidence first, then decide about a retry.

    Do you show UNKNOWN to the user, or does the agent try to resolve it before anyone sees it?

    1. 1

      This is a really useful pattern. The request ID + durable result + transactional event approach removes a large class of UNKNOWN states inside a system you control.

      I especially like the distinction between retries of the same execution attempt and actually creating a new execution. Reusing the same idempotency identity means the retry is asking for the result of the original attempt rather than silently creating a second side effect.

      For NAEOS, I’d treat that as one possible reconciliation mechanism rather than assuming every execution target can provide it. Once the side effect crosses into an external system, the control plane may only have provider receipts, resource state, logs, or an external idempotency key to work with.

      On your question: I don't want the agent to resolve UNKNOWN by reasoning about it. The runtime/control plane should own reconciliation first and return the resulting state to the agent.

      So potentially:

      UNKNOWN → RECONCILIATION → CONFIRMED / REJECTED / UNKNOWN

      If it remains UNKNOWN, the agent can see that state and propose what to do next, but it shouldn't be allowed to reinterpret it as success or failure. For higher-risk operations, the unresolved state may also need to be surfaced to a human rather than silently continuing.

      Your five-minute request-ID window is also an interesting boundary. It raises the question of what durable evidence NAEOS should retain after the provider's idempotency window expires.

      1. 1

        Agreed on the agent not reasoning its way out of UNKNOWN. We have the same rule for permissions: the agent can propose a high-risk write, but it can't decide that it's allowed to run it.

        On the five-minute window: for us it only covers the cached response. The events stay in the log for good, so after the window you can still answer whether the write happened. What you lose is the automatic dedup. A retry after the window runs the handler again, unless the handler first checks the log for its own event. For writes that must never repeat, that check is the durable part, and the request-id cache is just the fast path.

  2. 1

    Really interesting approach! Treating UNKNOWN as a separate state instead of automatically retrying makes a lot of sense, especially for production deployments. How are you planning to handle cases where reconciliation can't establish the outcome?

    1. 1

      Yes — I think this is actually the harder case.

      If reconciliation cannot establish whether the side effect happened, I don’t think NAEOS should manufacture a CONFIRMED or REJECTED outcome just to close the state.

      I’d keep it as a first-class UNKNOWN state, but record the reconciliation attempt and the evidence that was inspected, including why the outcome could not be established.

      The important part is what happens next: while the operation remains unresolved, NAEOS should prevent further non-safe side effects that depend on that outcome.

      So the flow becomes:

      UNKNOWN → RECONCILIATION → CONFIRMED / REJECTED / UNKNOWN

      And if it remains UNKNOWN, the system preserves that uncertainty rather than converting it into a false certainty.

      That also raises the next question for me: what evidence should be considered sufficient to transition UNKNOWN → CONFIRMED or UNKNOWN → REJECTED?

      That’s probably the next boundary worth testing.

  3. 1

    Really solid approach — I'm juggling something similar myself (building Xstream4K on the side), what's been the hardest part for you so far?