2
5 Comments

The rise and fall of the human in the loop

A couple of weeks ago I posted 4 questions I run before an agent touches production. The replies and the emails were all about question 4, which input moved the number. Fair, it's the one that fails a review. But none of the 4 is the actual test, and neither is the first incident.

Everyone survives the first incident. Someone notices, someone fixes the record by hand, someone writes "we added a check" in the retro. I've done all 3. The second incident is where you find out where that check went: into the write path, or into a person's head.

Some of what it looked like on my side, all boring:

Someone bulk edits 4 tickets to done in 2 minutes and each one reads as a real change. A decision arrives from the meeting notes, then again from the chat thread an hour later, and the second copy opens a second review. A CRM stage moves and the person who moved it isn't the one who owns the account. And the fix from incident 1 turns out to be a rule that one person remembers.

What changed after incident 2, and none of it is the model:

The same signal from a second message never opens a second review. A dropped duplicate is a logged decision, not a silent skip.

A review stays open until a person decides it. A lost card or a missed message never closes one, and 2 people never work the same row.

Undo is compensation, never deletion. Reversing a confirmed write is a new proposed action that goes through the same gate as the first one, under the name of the person who approves it, from the ticket or the chat they already have open.

The company sees which questions the layer couldn't answer, per source, so the number that gets measured is per system, never per person.

The approval gate is the part people ask about and it's honestly not the point here. The point is that incident 2 was boring. The setup is written up here if it's useful: https://renezander.com/case-studies/operational-context-layer-governed-actions/

The part I haven't settled, and the reason I'm posting: "a review stays open until a person decides it" is the right rule, and it's also how you end up with a queue nobody owns. Week 1 everyone works it. Week 3 there are cards open for 5 days and the people who could decide them have switched the notifications off. Right now a card never expires, by design, so the queue only ever grows unless someone works it. If you run anything with a human decision in the loop, what happens to a card nobody picks up? Does it escalate, expire, or block the next action, and who decided that?

on September 17, 2026
  1. 1

    The queue-growth problem seems like the real next test. Have you seen what actually happens to unresolved reviews in practice—silent accumulation, escalation, or something else?

    1. 1

      yes - they don't stick around very long. newer fact proposals push them out.

      1. 1

        That displacement behavior is the interesting signal, especially if unresolved items naturally lose priority. If you’re open to it, what’s the best email to reach you on?

        1. 1

          you are with Beryxa, right? we already spoke, and you didn't follow through on my offer.

          Where I do see a possible overlap is downstream: Beryxa helps founders reach a decision, while I help companies turn AI and workflow decisions into production systems. If that need comes up with one of your clients, there may be a useful referral fit.

          1. 1

            You’re right — I didn’t follow through after our earlier conversation. The downstream referral fit makes sense, so let’s close the loop properly this time.