The Observability Stack

Observability in my homelab has two pillars: metrics and logs. Metrics run on Prometheus with Thanos for long-term storage. Logs run on Fluent Bit and Fluentd shipping to Loki. Both surface in Grafana, which also owns alerting: Grafana-managed rules fire anything that breaks to Slack.

This post is the high-level map. Each pillar has its own write-up with the real configuration, linked at the bottom.


The two pillars

flowchart LR
    P[Prometheus] --> T[Thanos] --> G[Grafana]
    FB[Fluent Bit] --> FD[Fluentd] --> L[Loki] --> G
    T -.blocks.-> M[(MinIO)]
    L -.chunks.-> M

Thanos blocks and Loki chunks both live in MinIO.

Grafana is the only thing I actually open. Everything behind it is plumbing: Prometheus and Loki are the storage engines, MinIO holds the durable data, and the rest is glue. The two pillars are independent, so I can break metrics without touching logs, and the other way around.


What's deployed

Component Pillar Role
Prometheus metrics Scrapes and stores metrics, 4-day local retention
Thanos metrics Long-term metric storage and unified querying
Fluent Bit logs DaemonSet that tails container logs, forwards them
Fluentd logs Aggregator that buffers to disk and ships to Loki
Loki logs Stores and indexes logs
Grafana both Dashboards, ad-hoc queries, and alerting
MinIO both Object storage for Thanos blocks and Loki chunks

Everything is a Helm wrapper chart deployed via ArgoCD, the same GitOps flow as the rest of the homelab. The metrics side comes from kube-prometheus-stack, which bundles Prometheus, Grafana, Node Exporter, and kube-state-metrics in one chart, plus a separate Bitnami Thanos chart on top. The logs side runs fluent-operator, which manages Fluent Bit and Fluentd through CRDs. Alerting is Grafana-managed rules provisioned by grafana-operator; there's no Alertmanager in the cluster.


Two clusters

The homelab runs two clusters, develop and production. Each runs its own full stack: its own Prometheus, Thanos, Loki, and MinIO. Production's Thanos Query federates develop's over the LAN, so a single Grafana can see metrics from both. The details live in the metrics write-up.


Going deeper

Each pillar gets its own write-up:

  • The Metrics Stack covers Prometheus, Thanos, multi-cluster federation, and where everything runs.
  • The Log Stack covers the Fluent Bit and Fluentd forwarder/aggregator pair, HA Loki on the community chart, and the runtime bugs the rebuild surfaced.
  • Alerting to Slack covers the Grafana-managed rules, the notification policy, and why Alertmanager left.

Keeping metrics and logs independent means I can rebuild one without dragging the other down with it. In a homelab where I'm constantly breaking things on purpose, that separation has paid off more than once.