The Observability Stack
Observability in my homelab has two pillars: metrics and logs. Metrics run on Prometheus with Thanos for long-term storage. Logs run on Fluent Bit and Fluentd shipping to Loki. Both surface in Grafana, which also owns alerting: Grafana-managed rules fire anything that breaks to Slack.
This post is the high-level map. Each pillar has its own write-up with the real configuration, linked at the bottom.
The two pillars
flowchart LR
P[Prometheus] --> T[Thanos] --> G[Grafana]
FB[Fluent Bit] --> FD[Fluentd] --> L[Loki] --> G
T -.blocks.-> M[(MinIO)]
L -.chunks.-> M
Thanos blocks and Loki chunks both live in MinIO.
Grafana is the only thing I actually open. Everything behind it is plumbing: Prometheus and Loki are the storage engines, MinIO holds the durable data, and the rest is glue. The two pillars are independent, so I can break metrics without touching logs, and the other way around.
What's deployed
| Component | Pillar | Role |
|---|---|---|
| Prometheus | metrics | Scrapes and stores metrics, 4-day local retention |
| Thanos | metrics | Long-term metric storage and unified querying |
| Fluent Bit | logs | DaemonSet that tails container logs, forwards them |
| Fluentd | logs | Aggregator that buffers to disk and ships to Loki |
| Loki | logs | Stores and indexes logs |
| Grafana | both | Dashboards, ad-hoc queries, and alerting |
| MinIO | both | Object storage for Thanos blocks and Loki chunks |
Everything is a Helm wrapper chart deployed via ArgoCD, the same GitOps flow as the rest of the homelab. The metrics side comes from kube-prometheus-stack, which bundles Prometheus, Grafana, Node Exporter, and kube-state-metrics in one chart, plus a separate Bitnami Thanos chart on top. The logs side runs fluent-operator, which manages Fluent Bit and Fluentd through CRDs. Alerting is Grafana-managed rules provisioned by grafana-operator; there's no Alertmanager in the cluster.
Two clusters
The homelab runs two clusters, develop and production. Each runs its own full stack: its own Prometheus, Thanos, Loki, and MinIO. Production's Thanos Query federates develop's over the LAN, so a single Grafana can see metrics from both. The details live in the metrics write-up.
Going deeper
Each pillar gets its own write-up:
- The Metrics Stack covers Prometheus, Thanos, multi-cluster federation, and where everything runs.
- The Log Stack covers the Fluent Bit and Fluentd forwarder/aggregator pair, HA Loki on the community chart, and the runtime bugs the rebuild surfaced.
- Alerting to Slack covers the Grafana-managed rules, the notification policy, and why Alertmanager left.
Keeping metrics and logs independent means I can rebuild one without dragging the other down with it. In a homelab where I'm constantly breaking things on purpose, that separation has paid off more than once.