Cluster Diagnostics
Use Diagnostics inside Clusters → Observability to investigate live Kubernetes-level issues inside a cluster. It is the first Observability tab and opens by default. It reads the Kubernetes API directly, so it works without VictoriaMetrics or VictoriaLogs.
This page is separate from Cluster Settings → Events:
- Diagnostics focuses on cluster health, Kubernetes warning events, K3s control plane signals, and problematic pods.
- Events focuses on Edka cluster operation telemetry such as provisioning, updates, and deletions.
What Diagnostics shows
Section titled “What Diagnostics shows”The page lists, from top to bottom:
- Health: the cluster’s status and one row per current warning.
- Kubernetes Events: every event Kubernetes still holds for the cluster.
- Pods Requiring Attention: pods that are failing, with the reason from each container.
- Workloads: Deployments, StatefulSets, and DaemonSets that are not fully ready.
- Persistent Volumes: every volume claim with its usage, fullest first.
- K3s Signals: control plane indicators such as certificate expiry, etcd snapshot health, and the etcd leader.
- Connectivity Incidents: the times Edka could not reach the Kubernetes API. An ongoing incident moves to the top of the page.
Health
Section titled “Health”The Health line shows Healthy, a count of warnings or issues, Unreachable or Disconnected, or Unknown while Edka has no node data yet. It also shows when the snapshot was taken, and Refresh takes a new one. The banner at the top of the cluster page explains an outage.
Each warning row opens the section with its detail:
| Warning | Opens |
|---|---|
| Pods | Pods Requiring Attention |
| Nodes | Explorer → Nodes, filtered to nodes with issues |
| etcd, Certificates, Snapshots | K3s Signals |
| Events | Kubernetes Events |
| Workloads | Workloads |
| Storage | Persistent Volumes |
| NAT gateway | Cluster Settings → NAT gateway |
A pod needs attention when a container is crash-looping or cannot start, when the pod failed, when its state is unknown, when it has been starting or not ready for more than 3 minutes, or when it has been terminating for more than 5 minutes. Past restarts do not count once the pod runs again.
A NAT gateway row appears when the cluster’s NAT gateway reports degraded, stale, or drifted, with the first reason it gives.
A Monitoring note means Edka’s own connectivity check is running late, so outage detection can be delayed. It is listed but does not change the cluster’s status.
The same warnings appear in the Health panel on the cluster Overview, and its rows open the same sections.
Kubernetes Events
Section titled “Kubernetes Events”The Kubernetes Events panel lists the events Kubernetes reports for the whole cluster.
You can:
- filter to warnings only
- switch scope by namespace
- refresh the feed on demand
Kubernetes keeps events for about one hour, so older events no longer appear.
Typical issues you can catch here include:
- image pull failures
- certificate or challenge failures
- scheduling problems
- repeated backoff events
A pod can show the Kubernetes phase Pending while its events name the real
issue, such as ImagePullBackOff.
Pods Requiring Attention
Section titled “Pods Requiring Attention”This section lists pods by their failure reason instead of the raw pod phase.
Examples include:
CrashLoopBackOffImagePullBackOffErrImagePull- terminated containers with errors
- pods that remain pending because they cannot schedule
For each pod, Diagnostics shows the pod message, restart counts, and container-level reasons when available. The list opens by itself when pods need attention.
Workloads
Section titled “Workloads”Workloads lists the Deployments, StatefulSets, and DaemonSets that are not fully ready, with their ready replica count and the reason Explorer found, such as a stuck rollout. Each row opens the workload in Explorer. When every workload is ready, the section shows a single line with the total.
Persistent Volumes
Section titled “Persistent Volumes”Persistent Volumes lists every PersistentVolumeClaim in the cluster, fullest first:
- namespace and claim name, which opens the claim in Explorer
- bind phase
- used space and capacity
- storage class
- pods that mount the claim
Usage comes from the node that mounts each claim, so a claim no pod mounts shows no usage. A claim at 85% of its capacity raises a Storage warning in Health, and at 95% the warning turns into an error.
K3s Signals
Section titled “K3s Signals”K3s Signals shows control plane health from the K3s API server’s own metrics:
- certificate expiry windows
- etcd snapshot success vs failure
- etcd leader visibility
- inflight API requests
- etcd request latency
If metrics are unavailable, Diagnostics shows that state explicitly.
Connectivity Incidents
Section titled “Connectivity Incidents”Connectivity Incidents lists each time Edka could not reach the Kubernetes API, with its start, duration, cause, and what Edka found when it investigated. Past incidents appear at the end of the page. The section is hidden when there are none.
Availability and refresh behavior
Section titled “Availability and refresh behavior”- Diagnostics is available inside the Observability workspace.
- The page is intended for users with cluster visibility who need fast debugging context.
- Kubernetes-backed diagnostics become available once cluster Kubernetes access is ready.
- Snapshot-based sections can be refreshed manually.
When to use Diagnostics vs Events
Section titled “When to use Diagnostics vs Events”Use Diagnostics when you want to answer questions like:
- Why is this pod failing?
- Why are there warnings in the cluster?
- Are K3s control plane signals healthy?
- Which namespace is producing warning events?
Use Cluster Settings → Events when you want to answer questions like:
- What happened during provisioning?
- Did a cluster update fail?
- Which lifecycle operation produced this error?
- What progress events were emitted during deletion or upgrade?