I made a Grafana Master-dashboard to monitor my Kubernetes cluster.

Features:

  1. Health bar at the top, that will turn amber or red when things go wrong, and send push notifications to my phone
  2. CPU, Memory and Network utilization
  3. Database and Storage backup statuses

I also added CPU/Memory usage request monitoring so that I can tune how much CPU/Memory a pod requests in the cluster.
image

And active alerts
image
I’ve muted these two, I need to add a 3rd node with more CPU and ram, currently if one node goes down there isn’t enough redundancy for the cluster to just keep working.

Lots more detail here: https://erasmus.works/ It’s all Open-Source, leave a Star on my repo if you like what you see.