
Hey Indie Hackers, my team and I build Mavibot, an automation platform with an AI agent that walks people through setup by talking to them instead of dropping them into a settings dashboard. We wrote about it here before and got some great questions in the comments about reliability and how it handles vague requests.
We just shipped a pretty big update and wanted to share what changed.
The agent got noticeably better at knowing when it can't handle something on its own and needs to check with a human first, instead of guessing and moving forward anyway. We also reworked the dashboard metrics, and speed is way up, setting up a bot or a site now takes about 2 minutes.
The bigger shift is what it's turning into: basically a personal assistant for the business owner. Every morning it puts together a report of what it did overnight, what needs attention, how many deals are stalled, and what could be done about them. All of that is now in the mobile app too, so you're not tied to a desk to check in on it.
Next up we're working on voice control, not just typing a task, but actual real-time conversation with the agent while it helps think through a problem.
Curious how other builders here are thinking about the line between "agent handles it" and "agent asks first." Are we headed the right direction by leaning more into the second one, or does that slow things down too much for your use case?
I think the “ask first” approach makes more sense as the stakes increase. An agent making a small reversible change can probably act autonomously, but anything involving money, customers, or irreversible decisions should have a clear human checkpoint.
The tricky part is defining that boundary without making the agent feel like a glorified approval system. I’d be interested in whether you’re using different confidence thresholds based on the type of action, rather than one global rule for when the agent asks for confirmation.
right now it's more of a user-set mode (ask before everything, auto-apply but confirm deletions, or fully hands-off advice-only) rather than the system dynamically adjusting confidence thresholds per action type on its own. so the boundary is more "the user decided where they want it" than "the agent decided based on risk scoring." honestly your point about defining that boundary without it feeling like an approval bottleneck is the harder problem, still figuring out the right balance there
Yeah, I think the interesting part is that “user control” and “risk-aware autonomy” aren't necessarily the same thing.
A fixed mode gives users predictability, but it also puts the burden on them to correctly decide how much autonomy an agent should have. A dynamic system could theoretically reduce that burden, but then you introduce a whole new trust problem: who decides what counts as risky, and how do I know the agent's threshold is reasonable?
The sweet spot probably isn't asking for approval less often, but making the approval moments feel meaningful — interrupt when the consequence is genuinely hard to reverse, otherwise just act and explain afterward.
The UX challenge is basically making autonomy feel like a safety feature rather than an approval queue.
Leaning into "agent asks first" is the right call for onboarding specifically. The cost of a bad autonomous decision during setup is high — you've corrupted the user's first impression and possibly their data. The right tradeoff flips later in the lifecycle when the user trusts the system and the tasks are routine.
The morning report idea is where this gets interesting. Once users are reviewing a daily summary, they develop a mental model of what the agent does and doesn't do well. That context is what makes them comfortable approving more autonomy over time. You're essentially building trust in increments, which is exactly how it should work.
really good framing, "building trust in increments" is a cleaner way to put it than how I'd been thinking about it. the morning report as the thing that actually builds that mental model over time makes sense too, hadn't connected those two dots explicitly
Leaning towards "asks first" makes sense where nothing can check the answer. Where code can check it, like whether a phone number parses for the country or an email domain exists, I'd let code decide and have the agent ask the person only when the check fails, so you keep the speed without the guessing.
yeah, that makes sense, gives me something to think about, thanks
Love this angle. Building Xstream4K right now so this hits close to home — what made you look into it in the first place?
honestly it came from seeing how often "ai builds it for you" tools fall apart the moment the input data isn't clean, they just guess and produce something broken instead of flagging it. wanted ours to say what's missing and stop instead of pretending everything's fine