SonarOps

Monitoring that shows what broke, not just that it did.

Visit Website
June 17, 2026 I built the website monitor I always wanted — one that shows what broke, not just that it did

I'm a solo developer, and SonarOps is the monitoring tool I couldn't buy.

Every uptime tool I tried gave me a red dot and one probe's opinion: down, but never why. Half the false alarms were the probe's own network, not my site. So I built one that shows what actually broke.

It checks sites every 60 seconds from probes in the EU and USA. When one probe sees an outage, a second probe in another region confirms it before the alert fires, so a local blip doesn't page me at 3am. Beyond up/down it shows the signal journey: DNS, TLS handshake, first byte, full load. It tracks SSL expiry with reminders at 60, 30, 14 and 7 days and full chain validation, and crawls for broken links. Alerts go to email, Telegram or webhook.

Where I'm at: the platform is feature-complete and running in production on my own infrastructure. I'm hardening it and wiring up billing before I open public signups.

If monitoring that explains itself sounds useful, I'd love feedback on what you'd want to see. You can follow along at sonarops.it.

3 Comments

  1. 1

    The dual-probe confirmation is interesting.

    Are you weighting regions equally when they disagree, or treating one as primary and the other as validation?

    Feels like that choice could change how false positives behave under real traffic spikes.

    1. 1
      Neither, exactly — and that's the part worth explaining, because the two probes aren't answering the same question. There's a primary and a validator, and the asymmetry is one-directional: the second probe can cancel an alert, never create one. If the primary sees a failure and the second probe sees the site up, I stop and page nobody. If the second probe agrees it's down, that doesn't page anyone either — it just lets the check continue. It isn't a vote, and a second opinion is never sufficient on its own. The full path for a single failure: retry from the same probe after 2s — kills the one-off cross-probe check from a second probe, same region preferred second probe says up → stop it agrees (or times out) → re-check from the primary ~30s later Step 4 is the one that answers your question. You're right that the choice changes false-positive behaviour, but I'd frame it differently: the spatial step and the temporal step protect against different failure modes, and a traffic spike is the second kind. Spatial rules out something local to one probe's path — a bad route, a blip on one network. Two probes disagreeing is genuinely informative. But both probes sample roughly the same 12-second window. If the target itself stalls under load, both see it, they agree, and that agreement says nothing about whether it's still true thirty seconds later. So under a real traffic spike the cross-probe step gives you no protection at all — by design, it isn't the step meant to. The protection is the temporal recheck. Weighting the regions equally wouldn't help there; it would just make two probes agree on the same transient with more confidence. Two things I'd add because they're the parts I can't paper over: When the second probe says "up", there are two readings I can't distinguish at that instant — a blip local to the primary's path, or the target genuinely refusing that one address (rate limiting, IP reputation, a WAF rule). They look identical from here. So I record the observation ("secondary saw up"), not the conclusion. It matters if you're ever debugging why an outage didn't page you. And the suppression isn't unconditional: if a monitor recovers on the recheck but does it 3 times within 15 minutes, it escalates to an incident anyway. A site that flaps every few minutes has a problem, even though each individual check "recovered". On regions: the validator is picked in the same region as the primary when there's a healthy one, and only falls back across regions otherwise. Where the region is part of the question being asked, I'd rather lose the arbitration entirely than take the wrong arbiter — a cross-region validator confirming an outage nobody in the original region experienced is worse than no validator. The cost is real: a confirmed incident takes about half a minute longer to page than a naive checker would. I think that trade is obvious, but it is a trade, and it's the reason I don't claim faster detection than tools that alert on the first failed check.
      1. 2
        I see. That’s a much deeper answer than I expected. Happy to continue the conversation privately — what’s the best email to reach you on

About

Every uptime tool I tried gave me a red dot and a single probe's opinion: down, but never why. Half the false alarms were the probe's own network, not my site. I wanted monitoring that shows what actually broke.