1
1 Comment

Make the agent prove every claim

A useful lesson from building the Upwork Contract Safety SOP Pack: the hard part was not writing the content, it was making the agent justify every sentence. What failed first was the classic single-agent setup. One big prompt, a pile of source material, and a request to produce a polished SOP. It looked good on first read, but the failure mode was brutal for anything tied to revenue. The model would smooth over ambiguity, invent safe-sounding rules, and occasionally turn platform guidance into something that sounded like legal advice. That is fine for a demo, terrible for a product you have to stand behind. What worked was turning the workflow into a gated pipeline with typed state: retrieve -> outline -> draft -> validate -> assemble Each step wrote strict JSON, something like: {section_id, claim, source_ids[], confidence, blocked_reason} Two practical patterns mattered: 1. Retrieval returned small chunks with metadata, not raw docs. Every chunk had source_id, updated_at, and a short title.
2. Validation was deterministic. We rejected any section where source_ids.length == 0, where confidence was below threshold, or where banned phrases appeared, like "guaranteed protection" or anything implying formal legal advice. That changed the agent from "writer with opinions" into "compiler with receipts". The biggest architecture insight was this: retry the failing step, not the whole run. If the validator blocks section 3, only section 3 gets regenerated. That keeps costs stable and makes debugging possible. What I would do differently: define the invariants before writing prompts. Example: every recommendation must map to a cited source, every warning must include scope, every checklist item must be testable by a human. Prompts got much easier once the rules existed outside the model. If you are building autonomous agents for anything people pay for, do you treat output validation as a first-class system yet, or is it still buried inside prompting?

on April 22, 2026
  1. 1

    Making each section carry the ids of the notes it came from, and marking it blocked when it can't, seems stronger than any sterner wording in the prompt. A blocked section is annoying; a confident paragraph quoting a dropped rule is worse. Is validation its own step with a hard fail, or does "prove it" still live inside the prompt?