
Two weeks ago Anthropic launched Managed Agents on the Claude Platform. Hosted sessions, harnesses, sandboxes. The pitch is simple: stop gluing together your own agent infrastructure, let them run it for you.
I run a small managed agent shop with my co-founder Brandon. We have five customer agents in production right now. Total. Not five thousand. Five. So when the company that ships the model also ships the runtime, the question lands hard: does our category still exist next quarter?
Here is what we did the day the announcement dropped, and what changed (and didn't) two weeks in.
Brandon read the docs first. I read his Slack messages. By 4pm he had the same take three other founders DM'd me: "this is a horizontal default, not a vertical replacement."
Anthropic is great at the model. Their hosted runtime is also good, and it will become the default for teams that want to run a Claude agent and nothing else. Real audience.
Our customers are not running "a Claude agent" though. They are running a long-lived assistant that has to talk to a payments platform they already pay for, survive a node restart at 3am without paging anyone, show the operator what the agent did in plain English the next morning, and run for 30+ days for a non-technical owner who would rather lose a thumb than touch a terminal.
A horizontal hosted runtime ships fast. A vertical, boring-reliable runtime for non-technical operators is a different shape of work. We sell the second one.
I owe IH honest numbers, not a victory lap.
The MRR is small because the product is. We are not chasing thousands of agents. We are trying to be the place a non-technical operator parks one agent and forgets about it for a year. Anthropic's launch does not change that customer.
Three things.
First, we rewrote our homepage. Anthropic's launch made it concrete: "self-host vs managed" is no longer a binary. There are at least four lanes (DIY, Anthropic-managed, vendor-managed-on-Anthropic, fully-managed-with-our-own-runtime). I broke down the four lanes in this piece on managed vs self-hosted agents. Linking it because three IH founders DM'd asking the same question this week.
Second, we doubled down on observability. Anthropic gives you logs and traces, great for engineers. Our customers do not read logs. They read a one-paragraph "yesterday your agent did X, Y, Z" email. We had a half-built version. We finished it. The diff between raw events and a narrative summary your bookkeeper reads on Monday is the actual product. Wrote up our approach to AI agent observability if anyone is fighting the same translation problem.
Third, we stopped pitching to developers entirely. Every demo since the announcement has been a non-technical operator. Plumber. Yoga studio. Two-location bakery. None of them care about the runtime. They care that the agent answers texts the same way on Tuesday that it did on Monday.
If you build on top of Anthropic and they just shipped your feature: fine, you are now competing with an OK default instead of a bad default. That is good for the category. The bad default kept inertia high.
The trap is benchmarking against them on horizontal axes. Faster API. Cheaper tokens. More models. You will lose those fights. The vertical axis (boring, reliable, runs for a year for someone who does not want to think about it) is wide open. Anthropic is not going to call your customer when their agent stalls at 2am.
For us, the boutique angle reads sharper than it did three weeks ago. Smaller is the moat for a while. If you want to see how we frame the offer for non-technical operators, it lives at managed AI agents.
The part where I said our 5-customer audience does not overlap with Anthropic's launch. In 12 months there will be a customer who started on hosted Claude, outgrew it, and bounced to a managed shop. Maybe that becomes customer six. I do not know yet.
Will report back next month with whatever broke. Boutique pace.
The moat is not “managed agents.” It’s managed outcomes.
Anthropic can ship the runtime.
They’re not shipping “the agent handled 47 customer texts, escalated 2 edge cases, and your staff didn’t touch it.”
That’s the actual product.
Once the buyer is non-technical, infra stops being the differentiator.
Reliability translation becomes the differentiator.
That’s where most agent teams still overbuild for developers and underbuild for operators.
The winners here probably won’t look like better agent infra.
They’ll look like boring vertical software with an agent quietly buried inside it.
This is a really useful framing.
The “horizontal default vs vertical operator workflow” distinction feels like the key point.
Competing on runtime, logs, token cost, or model access seems like a losing game for a small team.
But translating all of that into something a non-technical customer can trust every day is a very different product.
The one-paragraph daily summary example is especially good. It makes the value feel less like infrastructure and more like peace of mind.
Curious how you decide which verticals are worth serving, since the deeper you go into operator workflows, the more specific the product probably has to become.