Skip to main content

Monitoring you can explain during an incident

A dashboard full of charts does not help if nobody knows which chart to look at first.

An engineer points to one clear alert panel for an operations manager, a dimmed wall of charts beside them

Monitoring is often installed like a shopping list: enable everything, turn on every alert. The result is fifty notifications a day that nobody reads after two weeks — and a real incident passes through them unnoticed.

What we install is smaller and sharper. Four numbers that actually mean something to users: does the page open, how fast, what share of requests fail, and is the queue moving. Alerts exist only for those four, and every alert names the first step to take.

Everything else is still recorded but rings no bells. It is useful when tracing things afterwards, and that is exactly its job.