Cluster monitoring
Track live CPU, memory, storage and daemon health across every cluster host and storage device over time.
Cluster monitoring
Purpose
- Watch live host CPU and memory and storage utilisation across all clusters.
- Spot degraded or down hosts, storage devices and daemons at a glance.
- Review host state, storage state and daemon health over a chosen time window.
- See pending patch counts alongside health in one KPI strip.
Prerequisites
- Root-account-only feature in the Platform section.
- Minimum permission to view: platform.clusters.managed.read
- Metrics are Graphite-backed and refresh automatically about every 60 seconds while the page is open.
- Time ranges available: 1h, 24h, 7d and 30d.
View all-cluster monitoring
The all-cluster dashboard aggregates every cluster's hosts and storage.
- Open Platform → Cluster Monitoring (root account only).
- Read the KPI strip: Hosts (healthy / online), Storage healthy, Daemon issues, Storage used, Peak host and Patches pending.
- Use the time-range tabs (1h, 24h, 7d, 30d) to set the window for the trend charts and state timelines.
- Review the trend charts: Host CPU %, Host Memory %, Storage Used % and Storage Used (GB), one line per host or device.
- Filter by Cluster, Host or Storage device to narrow the view; selections persist in the page URL.
- Expected result: charts and timelines redraw for the selected range and filters, while KPI tiles stay on current state.
Read state timelines and daemon health
Interpret the host/storage state bands and drill into daemons for a single cluster.
- Open a cluster's Monitoring tab (Platform → Clusters, open a cluster, select Monitoring), or scroll the all-cluster page.
- Read the Host state and Storage state timelines: Good (green), Degraded (amber), Down (red) and Unknown (grey); the most recent band is the current state.
- In the Daemons matrix, expand a host to see each daemon's state, then expand a daemon to load its breakdown metrics.
- Expected result: current-state views (KPI tiles and the daemon matrix) stay fixed regardless of the selected time range.
The KPI tiles and daemon matrix always reflect current state and are not affected by the time-range tabs; only the charts and timelines follow the range.
Best Practices
- Start from the all-cluster KPI strip, then drill into a cluster's Monitoring tab for host- and daemon-level detail.
- Investigate anything Down (red) first, then Degraded (amber), before healthy items.
- Use the range tabs to tell a transient spike apart from a sustained trend.
- Leave the page open to rely on the automatic refresh rather than reloading.
Ready to rethink private cloud?
Lower costs. Simplify operations. Deliver more.