At one point, I thought I just needed better coverage.
So I added more tools. One for logs, one for errors, one for uptime, and a couple more because they “might be useful later.”
On paper, it looked solid. In reality, 𝘪𝘵 𝘮𝘢𝘥𝘦 𝘵𝘩𝘪𝘯𝘨𝘴 𝘸𝘰𝘳𝘴𝘦.
I had 𝗺𝗼𝗿𝗲 𝗮𝗹𝗲𝗿𝘁𝘀, 𝗺𝗼𝗿𝗲 𝗱𝗮𝘀𝗵𝗯𝗼𝗮𝗿𝗱𝘀, 𝗺𝗼𝗿𝗲 𝗱𝗮𝘁𝗮… but 𝗹𝗲𝘀𝘀 𝗰𝗹𝗮𝗿𝗶𝘁𝘆 about what actually needed my attention.
The problem with stacking monitoring tools is simple.
You start getting duplicate alerts for the same issue. Some tools flag things too early, others too late. A few send warnings that sound serious but don’t require action.
After a while, you stop reacting the way you should. You skim. You delay. Sometimes you ignore. And that’s where things break.
Because the issues that really matter are rarely loud.
They show up as small inconsistencies. A request that fails more often than usual. A flow that behaves slightly differently. A pattern that looks off if you’re paying attention.
When everything is competing for your attention, those signals get buried.
𝗪𝗵𝗮𝘁 𝘄𝗼𝗿𝗸𝗲𝗱 𝗳𝗼𝗿 𝗺𝗲 𝘄𝗮𝘀 𝗰𝗵𝗮𝗻𝗴𝗶𝗻𝗴 𝘁𝗵𝗲 𝗮𝗽𝗽𝗿𝗼𝗮𝗰𝗵.
Instead of trying to watch everything, I focused on reducing what I had to pay attention to.
Now I rely on something that filters the noise and flags only the issues that matter, especially the ones that could turn into security or reliability problems if ignored. I use Gordon (https://trygordon.ai/) for this, mostly so I don’t have to constantly check multiple tools and guess what’s important.
If you're facing similar issues...be honest, how many alerts did you ignore today that could come back later?
More tools, more dashboards, more alerts, and yet less clarity about what actually needed fixing. At some point, you realise you're not managing your infrastructure anymore, but managing your monitoring setup. The alert fatigue point is very much underrated!!
Absolutely