Hey IH,
Starting a build-in-public log here. Day 1.
Quick context on what I'm doing: I'm building Relay, a harness for AI agents.
Not a wrapper around an API, not a chat UI. The actual infrastructure layer
that lets an agent run for a long time hours, not seconds recover when it
breaks something, verify its own output before it ships, and keep what it
learned instead of starting from zero next run.
Why I'm doing this: I think the actual bottleneck in AI right now isn't model
intelligence, it's endurance. Every agent I've used is impressive for 30
seconds and then falls apart the moment something goes wrong, because there's no recovery loop and no verification step. You end up babysitting it, which efeats the point.
Where it's at today: it runs coding tasks end to end environment setup,
attempt, run checks in a sandbox, recover on failure, ship the PR through git.
Free to run locally, no account needed. Pro tier exists for running it
unattended across repos if you want that.
I picked coding agents as the first wedge because it's the fastest place to
prove the recovery/verification loop actually works under real conditions.
Long term I think this pattern (environment + memory + recovery +
verification) is what needs to exist under a lot more than just coding.
Not here to pitch genuinely want feedback. If you've built or used AI agents
for real work (not demos), I'd love to know: where do they actually break for
you? What's the failure mode that makes you stop trusting them?
relayevals.com if you want to poke at it.
The endurance framing matches what I see, but the failure that makes me stop trusting a long run is context loss, not just a crash. After a recovery, the next session (sometimes a different coding agent) can't see why the previous one chose that approach, so it "fixes" the symptom again. I'd want the memory you mentioned to be queryable decision history, not a chat dump — otherwise recovery replays the same wrong lesson. Silent-wrongness from the comments above is the harder case: checks the agent can see become checks it can satisfy.
The silent-wrongness case gets worse when the same loop both attempts the fix and runs the checks, since anything the agent can see becomes something it can satisfy. Pinning the checks before the attempt, or deriving them from the issue rather than the diff, is the part I'd want hardened, otherwise 'recover on failure' quietly optimizes for green instead of correct. Same for the memory: a wrong lesson that persists across runs is worse than starting from zero.
For me the trust-breaker is silent wrongness, not loud failure. Recovery loops handle crashes well, but the thing that makes me stop trusting an agent is when it says it's done and everything looks right until you diff carefully — output that passes the checks but is subtly wrong. I'm curious how you're handling that case: when the agent's PR passes tests but fixes the symptom rather than the cause. That's where verification that understands intent starts to matter more than just re-running checks.
Have users of real coding agents shown one failure mode that repeatedly makes them stop trusting the agent, rather than just problems that are technically interesting to solve?
The recovery and verification loop feels like the interesting part here. I’ve also found that agents can look impressive until something unexpected happens, then you end up babysitting them anyway.
This resonates a lot — how long did it take before you saw any real signal on it?
Pretty early, but the signal got stronger once you use you start using regularly (founder)