I made a Grafana Master-dashboard to monitor my Kubernetes cluster.

Features:

  1. Health bar at the top, that will turn amber or red when things go wrong, and send push notifications to my phone
  2. CPU, Memory and Network utilization
  3. Database and Storage backup statuses

I also added CPU/Memory usage request monitoring so that I can tune how much CPU/Memory a pod requests in the cluster.
image

And active alerts
image
I’ve muted these two, I need to add a 3rd node with more CPU and ram, currently if one node goes down there isn’t enough redundancy for the cluster to just keep working.

Lots more detail here: https://erasmus.works/ It’s all Open-Source, leave a Star on my repo if you like what you see.

  • INeedMana@piefed.zip
    link
    fedilink
    English
    arrow-up
    4
    ·
    2 hours ago

    running pods: 98

    Oof, my lab is lightweight compared to yours but I’ve been thinking of putting up a grafana for mine too. Would you mind sharing how much resources it takes to run and roughly how much storage does, let’s say, a month of data take? I know that it all depends on what one puts inside but just as a reference point

    • Legion739@piefed.zipOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      2 hours ago

      See the “monitoring” namespace

      image

      kube-prometheus-stack plus VictoriaLogs for logs, about 100 pods on 2 nodes.

      The whole stack uses about 2.2 GB RAM. Prometheus takes about 1.1 GB of that and Grafana about 700 MB. CPU is basically idle, around 0.2 cores.

      Prometheus has about 150k series and uses about 7 GB for 15 days of metrics, so roughly 14 GB a month. Logs are about 3 GB a month. Grafana itself is under 1 GB.

      It’s mostly the default exporters, so a smaller lab would need less. A 20 GB volume would be plenty.

      Prometheus is set to retention: 15d and retentionSize: 18GB. Anything older than 15 days is dropped, and so is the oldest data if the total goes over 18 GB, whichever happens first. It’s sitting at a steady ~7 GB now that it has a full 15 days stored.

  • Legion739@piefed.zipOP
    link
    fedilink
    English
    arrow-up
    12
    ·
    3 hours ago

    Preemptively responding to this, because I just know there’s going to be a comment about this.

    Yes Yes I know, GitHub, screw GitHub, Screw Microslop. I am trying to move away from Big Tech as much as possible, but it’s a marathon not a sprint, and every win should be celebrated.

    Perfection is the enemy of good.
    .

    • hoshikarakitaridia@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 hours ago

      every win should be celebrated

      Yeah we could have some more of that here on Lemmy because people tend to be quite negative.

      Honestly that’s awesome, I didn’t manage to get grafana working when I tried exactly this, I’ll definitely check out any guides you used ^^

      • Legion739@piefed.zipOP
        link
        fedilink
        English
        arrow-up
        1
        ·
        2 hours ago

        Thanks!

        I took a lot of inspiration from here https://github.com/onedr0p/home-ops, there is also a discord server on there where you can ask people for help, they are quite helpful.

        I mostly shopped around and found parts and implementations I liked and coppied that and made it my own