3
2 Comments

The diff looks fine. The production impact doesn't.

A 5 line change to shared auth middleware sails through review. It's also imported by every service that logs a user in.

That's the kind of PR that worries me with AI written code.

The agent writes it in a couple of minutes, CI goes green, the diff looks reasonable, and it gets merged. The review process hasn't really changed, even though the amount of code going through it has.

The auth helper isn't the only case where the diff hides the real impact:

A migration renames or drops a column. Tests pass because the fixtures were updated in the same PR, but the old version of the app is still running during deployment.
A small change to a DB client wrapper or retry utility that half the codebase imports.
Config or environment changes that none of your tests load.
Tests written by the same agent that wrote the code. They pass because they were shaped around the implementation, so green tells you less than usual.

None of that is really a code review failure.

The reviewer is reading the diff. The diff just doesn't contain the answer.

That's the problem Production Reliability Index (PRI) at Tomosu is aimed at.

PRI gives a PR a 0–100 reliability score and a ready / not ready signal, based on seven areas: fragility, drift, governance compliance, runtime signals, code volatility, deployment velocity, and escalation.

There are limits too. A plain repo scan only gives you the code side. Runtime signals and escalation stay empty until you connect observability and ticketing data. It's also advisory by default, so it doesn't block a merge unless you configure a policy to do that.

I wrote the longer version here:

https://tomosu.ai/blogs/what-is-production-reliability-software-engineering.html

on September 30, 2026
  1. 1

    The “diff doesn’t contain the answer” framing resonates. One practical guardrail I’d add is an impact manifest before review: importers/consumers, schema compatibility, config or env keys, rollout and rollback plan, and the runtime dashboards that should move. Require it when a change crosses shared boundaries, then make canary → expand → rollback sequencing part of the ready signal. That keeps local changes fast while demanding deeper evidence where the blast radius is hidden.

    1. 1

      yeah, this is a good point. i'd also be careful about making the manifest another thing people have to fill out manually.

      the hidden consumers are exactly the problem. if the repo can tell you about importers, schema changes and config dependencies, that part should be automatic.

      the human part is more about rollout and rollback, and whether anyone is actually watching the right things after the change.

      that's also why i like having deployment velocity as one of the PRI signals. a small change shipped on its own is a very different situation from the same change bundled into a big release.