Hermes agent is pretty bonkers when it comes to building harnesses, self-improving agents. We were pretty skeptical about this process as well, and we just tested out building a simple security agent for ourselves.
We just created the whole self-evolving loop. And then we deployed this guy towards a bunch of other targets, and holy shit !!!!!!, he literally hacked multiple listed companies, apps and sshit. I asked it to go reach out to the CISOs to sell this agent, and this agent started doing that too.
I mean, obviously not as easy as I am telling here. I did a lot of improvements myself to make it good. There is a human in the loop. It is not like it just asked the agent and it did everything automatically like a superhuman. It's not like that, but practically speaking I did not expect Hermes agent to be this good, because the biggest issue is that if you have to get a model on its own, it is not as good as this.
This is multi-model, and also even if it's a single model, it is significantly better at getting or doing something than a model doing without the harness.
This is interesting, especially the part about the harness making the model feel “more capable” than the base model alone.
That said, I’d be careful with framing it as “hacking companies” — even in a test context it can come off pretty risky/ambiguous. I think the real signal here is the orchestration layer + human-in-the-loop making the agent actually usable in practice.
im going to look into this now and get it onto my tower asap. great post