Built-in metrics endpoint
KubeSolo has its own Prometheus endpoint that reports the health of the control plane it runs: whether each component is up, when it became ready, certificate expiry and the size of the database. It is off by default and listens on localhost when enabled:
The endpoint serves /metrics and /healthz over plain HTTP with no authentication, so bind it to a non-local address only on a network you trust. Probes run every 15 seconds.
Metric reference
| Metric | Type | Labels | Meaning |
|---|---|---|---|
| kubesolo_component_up | gauge | component | 1 if the component is reachable and healthy, 0 otherwise. |
| kubesolo_component_ready_timestamp_seconds | gauge | component | When the component first signalled readiness. 0 until ready. |
| kubesolo_component_last_probe_timestamp_seconds | gauge | component | When the component was last probed. |
| kubesolo_certificate_expiry_timestamp_seconds | gauge | name | When the certificate expires. Absent if it cannot be read. |
| kubesolo_certificate_valid | gauge | name | 1 if the certificate is readable and inside its validity window. |
| kubesolo_kine_db_size_bytes | gauge | Size of the SQLite database file. | |
| kubesolo_build_info | gauge | version, commit, build_date, go_version, arch | Always 1; build metadata in the labels. |
| kubesolo_start_time_seconds | gauge | When the metrics endpoint started. | |
| kubesolo_uptime_seconds | gauge | Seconds since the metrics endpoint started. |
The standard Go runtime (go_*) and process (process_*) collectors are included too.
| Label | Values |
|---|---|
| component | runtime, kine, apiserver, controller, kubelet, kubeproxy, coredns, webhook |
| name (certificates) | ca, apiserver, controller-manager, kubelet, admin, webhook, request-header-ca, request-header-client, plus d2k-server and d2k-client when d2k is enabled |
Scraping and alerting
Grafana dashboards
The repository ships a KubeSolo Cluster Overview dashboard in examples/grafana. It covers the cluster and node overview, API server request rate, errors and latency, admission webhook latency, kubelet pods and containers, per-pod CPU, memory, network and throttling from cAdvisor, controller workqueues, and Go runtime health.
The dashboard reads the Kubernetes components' own metrics, not the KubeSolo endpoint above. Its data comes from two scrape targets: the API server's /metrics, and kubelet plus cAdvisor metrics through the API server's node proxy.
| File | Use |
|---|---|
| kubesolo-grafana-dashboard-classic.json | Classic dashboard JSON. Recommended for Grafana OSS, Enterprise and Grafana Cloud. |
| kubesolo-grafana-dashboard-apiv2.json | The same dashboard in the dashboard.grafana.app/v2 format, for Grafana instances that support Dashboard API v2. |
| alloy-config.yaml | Grafana Alloy deployment that scrapes both targets and remote-writes to a Prometheus-compatible endpoint, with the RBAC it needs. |
| otel-collector-config.yaml | The same pipeline with the OpenTelemetry Collector. Needs the contrib distribution. |
To import, in Grafana go to Dashboards → New → Import, upload the JSON and select your Prometheus data source for ${datasource}. If Grafana reports unsupported apiVersion/kind, use the classic file.
Then set GRAFANA_CLOUD_URL, GRAFANA_CLOUD_USERNAME and GRAFANA_CLOUD_API_KEY on the Alloy deployment, ideally from a Kubernetes Secret, or change its prometheus.remote_write block to point at your own endpoint.
Logs
KubeSolo logs every component to one stream. With systemd:
On OpenRC the service logs to /var/log/messages, on SysV init to /var/log/syslog, and in daemon mode to /var/log/kubesolo.log. Set logging.debug: true for debug output.
Profiling with pprof
logging.pprof: true (or --pprof-server=true at install) starts Go's pprof server on port 6060 on all interfaces, without authentication. Enable it only while profiling, and firewall the port.