I've been researching a simple question:
What happens between detecting an operational problem and actually knowing what to do next?
The more conversations I have with incident managers, SREs, platform engineers, product leaders, and technical project managers, the less I think the problem is simply "not enough alerts."
Organizations already have plenty of signals.
The harder part is connecting those signals into something people can understand and act on.
So I'm experimenting with a workflow around:
Signal → Context → Evidence → Ownership → Action
The idea is simple.
Something happens.
For example:
«Payment service latency increases.»
What else is connected to the signal?
Instead of immediately making a conclusion, bring the relevant evidence together.
For example:
Now the system has a problem:
The evidence doesn't completely agree.
Rather than automatically declaring:
«"Team B owns it."»
the experiment exposes the uncertainty.
Ownership confidence: AMBIGUOUS
The goal is to make the evidence visible before someone acts on an assumption.
Once the relevant people and context are clear, the system can help identify a reasonable next step.
Not:
«"Automation has decided what you should do."»
But more like:
«"Here is what we know, here is where the evidence conflicts, here is who appears relevant, and here is a possible next step."»
What I'm learning
My earlier thinking focused heavily on ownership.
The research has made the problem more nuanced.
Sometimes the owner is documented.
Sometimes the owner is known.
But teams can still spend time reconstructing what happened, reconciling different signals, confirming responsibility, and deciding what evidence is trustworthy enough to act on.
That has pushed me toward evidence before automation.
I'm not trying to build another monitoring system.
I'm not trying to replace ServiceNow, observability, incident-management or existing systems of record.
And I'm deliberately not trying to build an AI that blindly decides who is responsible.
I'm testing whether connecting the right evidence and making uncertainty explicit can reduce some of the reconstruction work that happens before coordinated action.
The experiment is still being built, so none of this is validated yet.
The question I'm most interested in now is:
Does this remove work from an existing incident workflow — or does it simply create another workflow people have to maintain?
That's the part I want to find out.
If you're building infrastructure, developer tools, SRE tooling, or internal platforms, I'd be interested in how you'd approach this.
Great advice on shipping fast and talking to users. The feedback loop is everything in the early stage.
This resonates. I ran into a similar thing while building my own app, the signals were never the problem, figuring out which one actually mattered right now was. What ended up working better for you, a scoring system or just tighter context around each alert?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
The real risk is exactly the one you named: a new workflow always loses to removing work unless it lives inside the tool people already have open. Running Henson Group's MSP business across six continents, the incident tooling that actually got used was embedded in the ticketing system already open, never a separate dashboard nobody remembers to check during a live incident. I'd test this as a panel inside ServiceNow or PagerDuty rather than standalone, adoption dies the moment someone needs a second tab mid-incident.
The hard part of evidence-to-action is rarely the evidence — it's
making the action step opinionated enough that users actually take
it. Most tools die in the gap between "here's your dashboard" and
"here's what you do on Monday morning."
Curious where your users drop off: at collecting the evidence, or at
acting on it? In my experience the second one kills way more
workflows than the first.
The “evidence before automation” idea really resonates. I think there’s a similar problem in a lot of AI products: getting an answer is easy, but knowing whether there’s enough evidence behind that answer to actually act on it is much harder.
I’m exploring this from a different angle with an AI visibility platform I’m building. It lets users enter a brand and its competitors, run generated prompts through AI tools, then paste the actual responses back in for analysis.
One thing I’m deliberately interested in is separating the observation from the conclusion — e.g. “the brand appeared in 3/20 responses” is evidence, while “the brand has poor AI visibility” is an interpretation.
Your Signal → Context → Evidence → Ownership → Action framework makes me think there’s a similar pattern for AI analysis: Observation → Evidence → Interpretation → Action.
The question you ended with is probably the hardest one too: does the system actually remove work from the existing workflow, or just create another dashboard people have to maintain?
The "evidence doesn't completely agree" stage is the one that maps cleanly outside SRE too — I've seen the same pattern when product teams try to act on competitor signals. A pricing page change, a changelog bullet, and a sales call anecdote can all point different directions, and forcing a single "what they did" conclusion too early is how briefs get ignored.
One tactic that's helped: stamp every evidence item with (a) when it was true, (b) source type (primary page vs secondhand), and (c) the decision it could unblock. Then the weekly memo only includes items that clear a confidence bar, with conflicts left visible instead of flattened.
Curious whether your interviews have found a lightweight owner for that evidence pack — PM, PMM, or on-call — or whether the pack dies when it's "everyone's job."
This matches what I'm seeing building Mantis AI: the hardest part isn't collecting evidence, it's deciding how much confidence to assign it before it becomes a claim someone acts on. Curious how you're weighting recency vs. source reliability when they conflict.
That confidence threshold is a really interesting problem. I’m running into something similar with the AI visibility platform I’m building.
I’m trying to distinguish between what the AI response actually says and the conclusion we draw from it. For example, a brand appearing in 2 out of 20 responses is an observation; deciding that it has “poor AI visibility” is already an interpretation.
I think recency vs. source reliability gets even more interesting when the sources disagree. A recent source might reflect the current state better, but an older authoritative source could still be more trustworthy.
I’d probably want to preserve both signals rather than collapse them into one score, so you can actually see why confidence changed rather than just getting a number.
Great breakdown.
The explicit “Ownership confidence: AMBIGUOUS” step feels especially valuable—surfacing conflicting evidence seems safer than forcing a confident assignment. I also like that you’re testing whether this removes reconstruction work instead of assuming another layer of automation is automatically helpful. Have your interviews surfaced a lightweight way to measure that time saved?
That's one of the things I'm trying to pin down now. The interviews have consistently surfaced ownership clarification, context reconstruction, evidence reconciliation and repeated coordination as sources of delay, but I don't want to assume that an evidence layer actually saves time. I'm currently thinking about measuring things like time to first correct action, number of ownership handoffs/reassignments, and how often teams have to retrieve or request the same evidence more than once. The experiment is still being built, so I'm treating those as hypotheses to test rather than validated results.
The “evidence before automation” approach makes a lot of sense. Especially the part about exposing conflicting signals instead of forcing the system to choose an answer. That seems important for building trust with engineering teams.
Yeah, exactly. I’ve been thinking along similar lines with something I’m building around AI visibility for brands.
Instead of trying to automate the whole process or give a single “visibility score,” it lets you test a brand against its competitors by running prompts through AI models, pasting the responses back in, and analysing what actually shows up — including cases where the signals are conflicting.
The idea is to make the evidence visible first, rather than hiding the uncertainty behind an automated answer. Would be curious to hear how you think about this from the engineering side.
Really like the framing of evidence → action. We’re seeing a similar challenge while building for AI agents — generating more signals or eval results is relatively easy, but making them reproducible and actionable for engineering teams is where the real value starts to show up. Would be interesting to follow how your workflow evolves as you get more real-world usage.
Yeah, this resonates a lot. I’m seeing something similar while building my own product around AI visibility for brands.
The basic idea is to have brands test themselves against competitors across AI prompts, capture the actual responses, and then analyse the evidence to see how often and in what context each brand shows up. I’m deliberately keeping the workflow evidence-first rather than turning everything into a single automated score.
I’ve found that the interesting part isn’t generating the signals — it’s making them useful enough that you can actually understand why a brand is or isn’t showing up and what to do about it.
Great breakdown. What feedback have you had from early users?
I like the “evidence before automation” idea. How are you thinking about measuring whether it actually saves teams time?
One stage that bit me an hour ago: the evidence can lag the action.
I made a change, checked it from outside, got the old state back, and nearly reverted something that had worked fine. The read came from a cache with a five minute life. Nothing disagreed and nothing errored, the check just wasn't answering about now.
Worth stamping every piece of evidence with when it was true rather than when you fetched it. Freshness is the field people skip and then argue over.
This is a really useful distinction. “Retrieved at” and “true at” can tell very different stories, especially when the system itself hasn't produced an error. I hadn't separated those clearly enough in the experiment. Evidence freshness may need to be treated as part of the evidence itself, alongside the source and what that source is authoritative for. That's something I'll add to the things I'm testing.
That “evidence can lag the action” point really resonates. The timestamp distinction is easy to overlook, especially when the evidence itself looks perfectly valid.
I’m building a product around AI visibility for brands, where users run prompts about their brand and competitors and analyse the AI responses. Your point makes me think about how important it is to know exactly when a response was generated when comparing visibility over time.
Really interesting observation — especially the idea that stale evidence can look completely trustworthy while still answering the wrong question.
The "Evidence doesn't completely agree" stage is where most incident response tools fall over in my experience. Humans are actually good at reconciling conflicting signals, but they need the disagreement surfaced honestly instead of getting a flattened answer. One thing I'd watch: once you connect ownership into the same flow, teams can start treating suggested owners as blame assignments, which changes how honestly evidence gets entered. Keeping ownership as a tentative suggestion, rather than a conclusion, probably matters a lot for trust. Following this — the Signal to Context jump alone is a genuinely hard problem.
This is a really important point, especially the distinction between a suggested owner and a conclusion. I don't want Gnobu to turn an operational signal into “Team X is responsible” when the evidence doesn't support that certainty. The intent is closer to showing which evidence points where, where it conflicts, and why the system remains uncertain. The trust implication around ownership becoming a blame signal is something I hadn't considered deeply enough. I'll keep that in the experiment as a specific risk to test.
Really interesting point about suggested ownership turning into blame assignment. I hadn’t thought about that angle, but it makes sense — once a system attaches a person/team to a signal, people can start treating the suggestion as a conclusion.
I’m working on a related problem with my product, but specifically around AI visibility for brands. We compare how a brand and its competitors appear across AI-generated responses, and I’m deliberately keeping the raw responses and conflicting signals visible rather than collapsing everything into a definitive score.
The more I build it, the more I’m convinced that the “Signal → Context” step is where the real product value is. The signal itself is relatively easy to collect; understanding what it actually means without overclaiming is the hard part.
The evidence can disagree part is really important. I've seen systems jump from "we found this signal" straight to "this must be the cause" and then everything after that is built on the assumption. Keeping the conflict visible is probably more useful than pretending the system knows more than it does.
100% agree. That’s actually something I’ve been thinking about while building my product around AI visibility for brands.
When you compare how a brand shows up across different AI prompts/models, the signals can be pretty inconsistent — and I think that inconsistency is useful information in itself. Instead of forcing it into “your brand is visible” or “your brand isn’t visible,” I’m trying to surface the actual responses and where the evidence conflicts.
It feels much more trustworthy to let the user see the uncertainty and make sense of it rather than having the system confidently explain something that may not actually be true.
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
This is useful. How are you finding your first users so far?
Interesting. How are you measuring whether it is working?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Makes sense. Are you planning to charge for it, or keep it free for now?
Great breakdown. What feedback have you had from early users?
Nice work shipping it. What has been the biggest challenge since launch?
What made you pick this stack over the alternatives?
Interesting. How are you measuring whether it is working?
Appreciate the honesty here, most people only share the wins.
Interesting. How are you measuring whether it is working?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Clear and practical, thanks. Did anything surprise you along the way?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
This is useful. How are you finding your first users so far?
Makes sense. Are you planning to charge for it, or keep it free for now?
What made you pick this stack over the alternatives?
Appreciate the honesty here, most people only share the wins.
Interesting take. Would you still recommend this approach to someone starting today?
Good write-up. What would you do differently if you started again?
Thanks for writing this up. Bookmarking it for later.
Love this angle, honestly. What made you look into it in the first place?
In your conversations so far, have teams identified a concrete incident step that this evidence layer would replace, or is the value still mostly conceptual?
That's exactly what I'm trying to validate.
So far, the research has identified recurring friction around ownership, context reconstruction, evidence reconciliation, and coordination. But I haven't yet validated that an evidence layer actually replaces a concrete incident step.
That's the next experiment I'm building toward: whether it can reduce the work between the initial signal and the first correct coordinated action, without creating another workflow teams have to maintain.
The Signal → Context → Evidence → Ownership → Action chain makes the “does this remove work?” question testable. I’d start with one incident archetype and baseline time-to-first-correct-action, number of handoffs/reassignments, and how often someone has to ask for the same evidence twice. For ambiguous ownership, I’d avoid forcing a winner: show the top two candidates, the conflicting fields, and the one missing piece that would resolve it, then measure how often a human confirms it without opening another coordination loop. Evidence also needs a freshness/authority rule — a current service catalog and a recent deployment should not silently count the same as a stale spreadsheet. A shadow-mode pilot on existing incidents could show whether the workflow reduces reconstruction time before it becomes another system teams must maintain.
This is exactly the kind of thinking I’ve been applying to my own tool around AI brand visibility.
The workflow is essentially: give it a brand and its competitors, run a set of prompts through AI, paste the actual responses back in, and then analyse the evidence — which brands are mentioned, in what context, how consistently they appear, and where the responses disagree.
What I like about your Signal → Context → Evidence framing is that it maps really well to this problem too. It’s tempting to collapse everything into a simple “your brand is visible” score, but the useful part is actually being able to inspect the underlying evidence and understand why the result looks the way it does.
The “one missing piece that would resolve it” idea is particularly interesting. That feels like a much more actionable output than pretending the system has a definitive answer when the evidence is ambiguous.
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Clear and practical, thanks. Did anything surprise you along the way?
Great breakdown. What feedback have you had from early users?
What made you pick this stack over the alternatives?
Great breakdown. What feedback have you had from early users?
Clear and practical, thanks. Did anything surprise you along the way?