2
3 Comments

Our AI kept overruling the corrections users made

We let people correct us. Mark a sender as important. Fix a classification we got wrong. Tell us something we flagged as revenue isn't actually revenue.

Then a background job would rescan that same email later and quietly put its own answer back.

So someone tells the system "this isn't revenue," and the next day there it is again. From their side that isn't a bug, it's a personality. Being wrong once is forgivable. Being wrong again after they fixed it means the thing doesn't listen.

We'd already caught one process doing this and patched it. Went looking anyway and found a second writer sitting in the main classification flow. That's the part that still bothers me — nobody reported it. We just had a whole second thing writing to those fields that nobody had inventoried.

The fix itself is boring. Human answer wins, and we store where each value came from, model or person.

What I keep chewing on is that a system learning from corrections and a system overwriting them look identical from the inside. Same tables, same code path. The only difference is whether anything tracks provenance, and we weren't tracking it because we'd spent all our attention on accuracy and none on who gets the last word.

So for anyone else building on a model — how are you handling this? Do you store where a value came from, or does newest write just win? I don't think we've solved it. We've stopped the bleeding.

on September 4, 2026
  1. 1

    I use a slightly different approach when working with AI.

    I don't give human corrections priority over AI-generated values. Instead, I have a rule that an existing completed version should never be edited or overwritten directly.

    When something needs to change, I first make a copy, give that copy a new identity before making any changes, and then edit only the new version. The previous version remains untouched.

    One reason I adopted this rule is that I have seen AI make mistakes not during the correction itself, but during the cleanup afterward.

    For example, it can effectively go:

    "I fixed it."
    → "Now I should get rid of the old one."
    → "Let me find the old one."
    → "Found it. Delete it."

    And then it accidentally deletes the newly corrected version because it misidentifies which one is the old version.

    So rather than relying on whether a value came from a human or AI, I separate the old and new objects before any editing begins:

    OLD VERSION
    → COPY
    → NEW IDENTITY
    → EDIT
    → VERIFY

    If the new version turns out to be broken, I also don't necessarily repair that failed version directly. I can go back to the last known-good version, copy it again, give it a new identity, and selectively transfer only the intended changes from the failed version.

    Because of this, I don't need to treat human and AI decisions differently at the storage level. Whether a change comes from me or from AI, the previous state is preserved rather than overwritten.

    I can still record who or what made each change as metadata, but that is separate from the basic revision rule.

  2. 1

    The second writer is the part that really stands out.

    Once the user corrected the value, that effectively changed the authority for what should be written next — yet another process still had the technical ability to overwrite it.

    Have you looked at enforcing that boundary at execution time, rather than only tracking provenance afterward?

    I’m curious whether you’ve seen the same pattern anywhere else in the workflow.