A policy check can tell us whether an action is allowed.
But what happens after the action is authorized—and the runtime stops telling us what happened?
This sounds like a simple timeout problem.
It isn't.
For an AI coding agent, an execution timeout can create a much more dangerous state:
we don't know whether the side effect happened.
And that distinction matters.
Consider a simple deployment:
Agent proposes deployment
↓
Policy evaluation
↓
Authorization granted
↓
Runtime sends deployment request
↓
Network timeout
↓
???
What is the correct state?
Did the deployment fail?
Did it succeed?
Did the provider receive the request but fail to return the response?
We don't actually know.
Treating the timeout as FAILED would be an assumption.
Treating it as SUCCESS would also be an assumption.
So NAEOS needs another state.
UNKNOWN
This is the distinction I'm exploring in this build log.
A useful execution state machine looks more like:
PROPOSED
↓
AUTHORIZED
↓
EXECUTING
├──→ CONFIRMED
├──→ REJECTED
└──→ UNKNOWN
CONFIRMED means we have sufficient evidence that the side effect occurred.
REJECTED means we have evidence that the action did not occur or was explicitly rejected.
UNKNOWN means the runtime no longer has sufficient observation to determine the outcome.
That last state is uncomfortable.
But pretending uncertainty does not exist is worse.
Imagine an agent is authorized to deploy an application.
The runtime sends the request.
The provider processes it.
The deployment succeeds.
But the response is lost.
From the runtime's perspective:
request → timeout
The agent sees:
UNKNOWN
If the agent simply retries:
attempt #1 → UNKNOWN
attempt #2 → EXECUTE
we may now have two deployment attempts.
For some operations that may be harmless.
For others, it can be dangerous:
The important question is therefore not:
"Did the request timeout?"
The important question is:
"Can we establish what actually happened?"
This leads to another primitive I'm testing:
UNKNOWN
↓
RECONCILIATION
├── evidence confirms execution
│ ↓
│ CONFIRMED
│
├── evidence confirms no execution
│ ↓
│ REJECTED
│
└── insufficient evidence
↓
UNKNOWN
The runtime should attempt to reconcile the unknown state before deciding what happens next.
Depending on the operation, that evidence might include:
The exact evidence depends on the execution target.
But the principle remains the same:
Outcome should be established from evidence, not inferred from silence.
This is another distinction that became clearer while working on the previous build log.
A policy decision answers:
May this action happen?
Authorization answers:
Is this execution attempt permitted under the current context?
Execution answers:
What did the runtime attempt to do?
Observation answers:
What actually happened?
These are different questions.
So the model I'm working with now is:
Proposal
↓
Policy Evaluation
↓
Authorization
↓
Execution
↓
Observation
↓
Outcome
The agent can propose the next action.
It should not be the authority that declares the previous side effect successful simply because the runtime stopped responding.
This is particularly important for AI coding agents.
A model might reason:
"The deployment probably failed. I'll retry."
But "probably" is not an execution receipt.
The governance/runtime layer should instead be able to return something explicit:
{
"execution_status": "UNKNOWN",
"authorization_id": "auth_01",
"attempt_id": "exec_01",
"reconciliation_required": true
}
Now the agent has information about the state without being given authority to redefine it.
The agent can propose a next action.
The control plane decides whether that action is allowed.
And the runtime determines whether the previous action actually happened.
This creates an even harder question.
Suppose the system has:
UNKNOWN
and the operation is not safely idempotent.
Should NAEOS retry automatically?
My current answer is:
No—not without sufficient evidence or an explicit authorization path for the retry.
A safer decision path looks like:
UNKNOWN
↓
Can outcome be reconciled?
├── YES → reconcile
└── NO
↓
Is operation safely idempotent?
├── YES → controlled retry
└── NO → require explicit decision
The goal is not to make the system incapable of recovering.
The goal is to prevent uncertainty from silently becoming authority.
The next NAEOS execution-boundary tests are therefore becoming more explicit.
authorize
→ execute
→ provider rejects
Expected:
REJECTED
authorize
→ execute
→ receipt returned
Expected:
CONFIRMED
authorize
→ connection lost
→ no side effect
Initial state:
UNKNOWN
Then reconciliation should establish:
REJECTED
authorize
→ side effect occurs
→ response lost
Initial state:
UNKNOWN
Then reconciliation should establish:
CONFIRMED
authorize
→ execute
→ timeout
→ UNKNOWN
Expected:
no automatic retry
unless the system can establish that retry is safe and authorized.
One direction I'm exploring is making execution outcomes durable artifacts rather than transient agent messages.
For example:
{
"attempt_id": "exec_01",
"authorization_id": "auth_01",
"operation": "deploy",
"target": "production",
"status": "CONFIRMED",
"provider_receipt": "deployment_83921",
"artifact_digest": "sha256:...",
"observed_at": "...",
"runtime_version": "...",
"policy_version": "..."
}
The exact schema is still experimental.
The important part is the relationship:
Authorization
↓
Execution Attempt
↓
Observation
↓
Evidence
↓
Outcome
That chain should survive the agent that initiated it.
The deeper lesson here is not really about timeouts.
It is about distributed systems and control boundaries.
A timeout doesn't necessarily tell us that something failed.
It tells us that we lost observation of what happened.
For AI engineering systems, that distinction becomes particularly important because an agent can continue reasoning even when the underlying runtime has entered an uncertain state.
NAEOS should not hide that uncertainty.
It should make it explicit.
When observation is lost, don't manufacture certainty. Reconcile the state.
And if reconciliation cannot establish the outcome:
UNKNOWN
should remain UNKNOWN.
The next question is becoming even more interesting:
What evidence is sufficient to move an execution from UNKNOWN to CONFIRMED or REJECTED?
That takes us from execution control into something closer to an evidence and verification model.
And that's where I think the next NAEOS experiment should go.
How would you handle a non-idempotent operation when:
Would you reconcile first, require human intervention, use an idempotency mechanism, or take another approach?
I'd especially like to hear from people working on distributed systems, infrastructure, deployment platforms, and AI coding agents.
We hit a smaller version of this in our own framework, where the AI agent calls the same write handlers as the UI.
Two things removed most of the UNKNOWN cases for us. Every write carries a request id, and a retry reuses it. The server remembers the id: if the first attempt is still running, the retry waits for it, and if it already finished, the retry gets the stored result and nothing runs twice. Also, every state change is an event appended in the same transaction as the write, so whether it happened is a lookup in the event log.
It has limits. The server keeps those ids for five minutes, which covers a retry but not someone re-running the job the next day. And none of it helps once the side effect lives in someone else's system. There we only get certainty if their API takes an idempotency key as well. Otherwise we're where you are: find evidence first, then decide about a retry.
Do you show UNKNOWN to the user, or does the agent try to resolve it before anyone sees it?
This is a really useful pattern. The request ID + durable result + transactional event approach removes a large class of UNKNOWN states inside a system you control.
I especially like the distinction between retries of the same execution attempt and actually creating a new execution. Reusing the same idempotency identity means the retry is asking for the result of the original attempt rather than silently creating a second side effect.
For NAEOS, I’d treat that as one possible reconciliation mechanism rather than assuming every execution target can provide it. Once the side effect crosses into an external system, the control plane may only have provider receipts, resource state, logs, or an external idempotency key to work with.
On your question: I don't want the agent to resolve UNKNOWN by reasoning about it. The runtime/control plane should own reconciliation first and return the resulting state to the agent.
So potentially:
UNKNOWN → RECONCILIATION → CONFIRMED / REJECTED / UNKNOWN
If it remains UNKNOWN, the agent can see that state and propose what to do next, but it shouldn't be allowed to reinterpret it as success or failure. For higher-risk operations, the unresolved state may also need to be surfaced to a human rather than silently continuing.
Your five-minute request-ID window is also an interesting boundary. It raises the question of what durable evidence NAEOS should retain after the provider's idempotency window expires.
Agreed on the agent not reasoning its way out of UNKNOWN. We have the same rule for permissions: the agent can propose a high-risk write, but it can't decide that it's allowed to run it.
On the five-minute window: for us it only covers the cached response. The events stay in the log for good, so after the window you can still answer whether the write happened. What you lose is the automatic dedup. A retry after the window runs the handler again, unless the handler first checks the log for its own event. For writes that must never repeat, that check is the durable part, and the request-id cache is just the fast path.
Really interesting approach! Treating UNKNOWN as a separate state instead of automatically retrying makes a lot of sense, especially for production deployments. How are you planning to handle cases where reconciliation can't establish the outcome?
Yes — I think this is actually the harder case.
If reconciliation cannot establish whether the side effect happened, I don’t think NAEOS should manufacture a CONFIRMED or REJECTED outcome just to close the state.
I’d keep it as a first-class
UNKNOWNstate, but record the reconciliation attempt and the evidence that was inspected, including why the outcome could not be established.The important part is what happens next: while the operation remains unresolved, NAEOS should prevent further non-safe side effects that depend on that outcome.
So the flow becomes:
UNKNOWN → RECONCILIATION → CONFIRMED / REJECTED / UNKNOWN
And if it remains UNKNOWN, the system preserves that uncertainty rather than converting it into a false certainty.
That also raises the next question for me: what evidence should be considered sufficient to transition UNKNOWN → CONFIRMED or UNKNOWN → REJECTED?
That’s probably the next boundary worth testing.
Really solid approach — I'm juggling something similar myself (building Xstream4K on the side), what's been the hardest part for you so far?