From 8506f12bce5dfece12c45e4dfa3d4203e5cfa0b9 Mon Sep 17 00:00:00 2001 From: Sho Ishida Date: Sat, 18 Jul 2026 08:57:02 +0000 Subject: [PATCH] Document implentation of graphical dashboard for monitoring VMs --- README.md | 1 + docs/07-observability.md | 96 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 97 insertions(+) create mode 100644 docs/07-observability.md diff --git a/README.md b/README.md index 4ed6679..fe0b7b2 100644 --- a/README.md +++ b/README.md @@ -105,3 +105,4 @@ graph TD * [**Centralized Identity & DNS Management**](./docs/04-identity.md) * [**Application Container Deployments**](./docs/05-applications.md) * [**Ingress & Secure Tunnels**](./docs/06-tunnels.md) +* [**Observability & Telemetry**](./docs/07-observability.md) diff --git a/docs/07-observability.md b/docs/07-observability.md new file mode 100644 index 0000000..11a2c92 --- /dev/null +++ b/docs/07-observability.md @@ -0,0 +1,96 @@ +# 📊 Observability & Telemetry + +This document details the configuation and design of the cluster-wide telemetry scraping infrastructure utilizing **Prometheus**, **Grafana**, and native **Node Exporter**. + +--- + +## 📈 Observability Architecture + +Prometheus scrapes node metrics periodically via Node Exporter agents listening on port `9100` across the bridge network: + +```mermaid +graph TD + %% My Color Palette + classDef hostNode fill:#161d1c,stroke:#415854,color:#f8f8f2,stroke-width:2px; + classDef ipdNode fill:#2b3b38,stroke:#ff9580,color:#ff9580,stroke-width:1.5px; + classDef vmNode fill:#2b3b38,stroke:#70a99f,color:#f8f8f2,stroke-width:1.5px; + classDef obsNode fill:#2b3b38,stroke:#8aff80,color:#8aff80,stroke-width:1.5px; + + VM1["🔑 freeipa.lab.local (172.30.1.85)
LDAP / Kerberos / BIND DNS"]:::ipdNode + + subgraph Nodes ["Monitored Nodes (Port: 9100)"] + HostOS["🖥️ Hypervisor Host (172.30.1.200)"]:::hostNode + VM2["📄 portfolio VM (172.30.1.93)"]:::vmNode + VM3["⚔️ minecraft VM (172.30.1.91)"]:::vmNode + VM4["🎵 navidrome VM (172.30.1.92)"]:::vmNode + end + + subgraph PortfolioServices ["portfolio VM Services"] + Prom["📈 Prometheus TSDB"]:::obsNode + Grafana["📊 Grafana Dashboard"]:::obsNode + end + + HostOS & VM1 & VM2 & VM3 & VM4 -.->|Node Exporter Scrape| Prom + Prom -->|Data Source Query| Grafana + + Nodes --->|DNS Lookups| VM1 + + %% Subgraph Colors + style Nodes fill:#212c2a,stroke:#70a99f,stroke-width:1px; + style PortfolioServices fill:#161d1c,stroke:#415854,stroke-width:2px; +``` + +--- + +## 📄 Service Configurations + +Telemetry collectors are automated via `ansible/playbooks/07_observability.yml`: + +### 1. Prometheus Node Exporter (Daemon Node) + +* **Engine**: Downloads the native `node_exporter-1.8.2.linux-amd64` release binary and places it under `/usr/local/bin/node_exporter`. +* **Systemd Integration (`node_exporter.service`)**: Deploys a background daemon unit configured to run the exporter immediately after network interfaces load. +* **SELinux Contexts (AlmaLinux)**: Automatically runs `restorecon` to preserve SELinux contexts for the binary and service files. +* **Security**: Opens port `9100/tcp` on firewalld on RedHat-family systems to allow scrapers to read metric inputs. + +### 2. Prometheus Engine (`prometheus.yml.j2`) + +* **Engine**: Containerized using `docker.io/prom/prometheus:latest` running with `--net=host` on the `portfolio` VM. +* **Storage Mounts**: Maps the host folder `/home/sho/containers/prometheus/data` to `/prometheus` with tag properties `:z,U`(ensuring rootless Podman SELinux permissions map correctly to local storage directories). +* **Scrape Loop Specifications**: Sets a scrape and evaluation interval of 15 seconds: + +```yaml +global: + scrape_interval: 15s + evaluation_interval: 15s + +scrape_configs: + - job_name: 'homelab-nodes' + static_configs: + - targets: + - '172.30.1.200:9100' # Host + - '172.30.1.85:9100' # freeipa + - '172.30.1.93:9100' # portfolio + - '172.30.1.91:9100' # minecraft + - '172.30.1.92:9100' # navidrome +``` + +### 3. Grafana Dashboard + +* **Engine**: Runs `docker.io/grafana/grafana-oss:latest` in a container mapping port `3000:3000`. +* **Data Persistence**: Mounts `/home/sho/containers/grafana/data` to preserve custom dashboards, datasources, and user configurations between restarts. + +--- + +## 🚀 Execution & Monitoring + +To deploy the observability stack across the cluster nodes, run the following: + +```bash +ansible-playbook site.yml --tags "observability" --ask-vault-pass +``` + +### Checking Scraping Sinks + +1. **Prometheus Targets Console**: Access the Prometheus TUI interface by opening `http://172.30.1.93:9000/targets` and verify that all 5 target hosts report `UP`. +2 **Grafana Portal**: Navigate to `http://172.30.1.93:3000` to create custom query dashboards. *(Default port is mapped externally via Cloudflared Tunnel).*