← All notes

A small Prometheus and Grafana setup

Until this spring, my monitoring strategy was noticing that something felt slow. That worked until the root filesystem filled up with logs while I was away for a week. It was time for something slightly more grown up.

The setup is modest: Prometheus and Grafana run in containers on the server, node_exporter runs on every machine in the house, and smartctl_exporter keeps an eye on the drives. There is also a small exporter I wrote that reports the age of the latest backup snapshot, which is probably the single most useful metric I have.

In Grafana I kept things simple. One dashboard shows the overview: CPU, memory, disk usage, temperatures and network traffic. A second one is about storage: pool health, scrub results and drive temperatures over time. I resisted the urge to graph everything.

Alerting goes through Alertmanager to email. There are only five rules: disk above 85 percent, a drive reporting SMART errors, a machine down for more than ten minutes, backups older than two days and the CPU running hot. Fewer alerts means I actually read them.