Source: https://git.zabbix.com/projects/ZBX/repos/zabbix/browse/templates/app/kubernetes_http?at=release/7.4
A template set for monitoring your Kubernetes cluster via Zabbix 7.0 and higher. Zabbix provides a powerful automated solution for monitoring the Kubernetes cluster components. You need to deploy Zabbix Helm Chart with Zabbix Proxy and Zabbix agents to monitor the cluster.
Zabbix is an open-source product that can be installed on a majority of Unix-like distributions at no cost - see full list of supported distributions. Alternatively, Zabbix is available on certain cloud services.
Templates for Kubernetes monitoring
| Name | Readme | Template |
|---|---|---|
| Kubernetes nodes by HTTP | Readme | Template |
| Kubernetes cluster state by HTTP | Readme | Template |
| Kubernetes API server by HTTP | Readme | Template |
| Kubernetes Controller manager by HTTP | Readme | Template |
| Kubernetes Scheduler by HTTP | Readme | Template |
| Kubernetes kubelet by HTTP | Readme | Template |
Templates are of two types:
Cluster node monitoring
Kubernetes nodes by HTTP template discovers cluster nodes, creates hosts in Zabbix based on prototypes and assigns the "Linux by Zabbix agent" template to them. The template collects basic node metrics via the Kubernetes API.
Main cluster components monitoring
Kubernetes cluster state by HTTP discovers cluster components and control plane nodes, creates Zabbix hosts and assigns the required templates to them.
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
It is strongly recommended to use filtering macros when configuring templates, as on a large cluster the number of discoverable objects can reduce the performance of the monitoring system.
Source: https://git.zabbix.com/projects/ZBX/repos/zabbix/browse/templates/app/kubernetes_http?at=release/7.2
A template set for monitoring your Kubernetes cluster via Zabbix 7.0 and higher. Zabbix provides a powerful automated solution for monitoring the Kubernetes cluster components. You need to deploy Zabbix Helm Chart with Zabbix Proxy and Zabbix agents to monitor the cluster.
Zabbix is an open-source product that can be installed on a majority of Unix-like distributions at no cost - see full list of supported distributions. Alternatively, Zabbix is available on certain cloud services.
Templates for Kubernetes monitoring
| Name | Readme | Template |
|---|---|---|
| Kubernetes nodes by HTTP | Readme | Template |
| Kubernetes cluster state by HTTP | Readme | Template |
| Kubernetes API server by HTTP | Readme | Template |
| Kubernetes Controller manager by HTTP | Readme | Template |
| Kubernetes Scheduler by HTTP | Readme | Template |
| Kubernetes kubelet by HTTP | Readme | Template |
Templates are of two types:
Cluster node monitoring
Kubernetes nodes by HTTP template discovers cluster nodes, creates hosts in Zabbix based on prototypes and assigns the "Linux by Zabbix agent" template to them. The template collects basic node metrics via the Kubernetes API.
Main cluster components monitoring
Kubernetes cluster state by HTTP discovers cluster components and control plane nodes, creates Zabbix hosts and assigns the required templates to them.
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
It is strongly recommended to use filtering macros when configuring templates, as on a large cluster the number of discoverable objects can reduce the performance of the monitoring system.
Source: https://git.zabbix.com/projects/ZBX/repos/zabbix/browse/templates/app/kubernetes_http?at=release/7.0
A template set for monitoring your Kubernetes cluster via Zabbix 7.0 and higher. Zabbix provides a powerful automated solution for monitoring the Kubernetes cluster components. You need to deploy Zabbix Helm Chart with Zabbix Proxy and Zabbix agents to monitor the cluster.
Zabbix is an open-source product that can be installed on a majority of Unix-like distributions at no cost - see full list of supported distributions. Alternatively, Zabbix is available on certain cloud services.
Templates for Kubernetes monitoring
| Name | Readme | Template |
|---|---|---|
| Kubernetes nodes by HTTP | Readme | Template |
| Kubernetes cluster state by HTTP | Readme | Template |
| Kubernetes API server by HTTP | Readme | Template |
| Kubernetes Controller manager by HTTP | Readme | Template |
| Kubernetes Scheduler by HTTP | Readme | Template |
| Kubernetes kubelet by HTTP | Readme | Template |
Templates are of two types:
Cluster node monitoring
Kubernetes nodes by HTTP template discovers cluster nodes, creates hosts in Zabbix based on prototypes and assigns the "Linux by Zabbix agent" template to them. The template collects basic node metrics via the Kubernetes API.
Main cluster components monitoring
Kubernetes cluster state by HTTP discovers cluster components and control plane nodes, creates Zabbix hosts and assigns the required templates to them.
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
It is strongly recommended to use filtering macros when configuring templates, as on a large cluster the number of discoverable objects can reduce the performance of the monitoring system.
Source: https://git.zabbix.com/projects/ZBX/repos/zabbix/browse/templates/app/kubernetes_http?at=release/6.4
A template set for monitoring your Kubernetes cluster via Zabbix 6.4 and higher. Zabbix provides a powerful automated solution for monitoring the Kubernetes cluster components. You need to deploy Zabbix Helm Chart with Zabbix Proxy and Zabbix agents to monitor the cluster.
Zabbix is an open-source product that can be installed on a majority of Unix-like distributions at no cost - see full list of supported distributions. Alternatively, Zabbix is available on certain cloud services.
Templates for Kubernetes monitoring
| Name | Readme | Template |
|---|---|---|
| Kubernetes nodes by HTTP | Readme | Template |
| Kubernetes cluster state by HTTP | Readme | Template |
| Kubernetes API server by HTTP | Readme | Template |
| Kubernetes Controller manager by HTTP | Readme | Template |
| Kubernetes Scheduler by HTTP | Readme | Template |
| Kubernetes kubelet by HTTP | Readme | Template |
Templates are of two types:
Cluster node monitoring
Kubernetes nodes by HTTP template discovers cluster nodes, creates hosts in Zabbix based on prototypes and assigns the "Linux by Zabbix agent" template to them. The template collects basic node metrics via the Kubernetes API.
Main cluster components monitoring
Kubernetes cluster state by HTTP discovers cluster components and control plane nodes, creates Zabbix hosts and assigns the required templates to them.
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
It is strongly recommended to use filtering macros when configuring templates, as on a large cluster the number of discoverable objects can reduce the performance of the monitoring system.
Source: https://git.zabbix.com/projects/ZBX/repos/zabbix/browse/templates/app/kubernetes_http?at=release/6.2
A template set for monitoring your Kubernetes cluster via Zabbix 6.2 and higher. Zabbix provides a powerful automated solution for monitoring the Kubernetes cluster components. You need to deploy Zabbix Helm Chart with Zabbix Proxy and Zabbix agents to monitor the cluster.
Zabbix is an open-source product that can be installed on a majority of Unix-like distributions at no cost - see full list of supported distributions. Alternatively, Zabbix is available on certain cloud services.
Templates for Kubernetes monitoring
| Name | Readme | Template |
|---|---|---|
| Kubernetes nodes by HTTP | Readme | Template |
| Kubernetes cluster state by HTTP | Readme | Template |
| Kubernetes API server by HTTP | Readme | Template |
| Kubernetes Controller manager by HTTP | Readme | Template |
| Kubernetes Scheduler by HTTP | Readme | Template |
| Kubernetes kubelet by HTTP | Readme | Template |
Templates are of two types:
Cluster node monitoring
Kubernetes nodes by HTTP template discovers cluster nodes, creates hosts in Zabbix based on prototypes and assigns the "Linux by Zabbix agent" template to them. The template collects basic node metrics via the Kubernetes API.
Main cluster components monitoring
Kubernetes cluster state by HTTP discovers cluster components and control plane nodes, creates Zabbix hosts and assigns the required templates to them.
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
It is strongly recommended to use filtering macros when configuring templates, as on a large cluster the number of discoverable objects can reduce the performance of the monitoring system.
Source: https://git.zabbix.com/projects/ZBX/repos/zabbix/browse/templates/app/kubernetes_http?at=release/6.0
A template set for monitoring your Kubernetes cluster via Zabbix 6.0 and higher. Zabbix provides a powerful automated solution for monitoring the Kubernetes cluster components. You need to deploy Zabbix Helm Chart with Zabbix Proxy and Zabbix agents to monitor the cluster.
Zabbix is an open-source product that can be installed on a majority of Unix-like distributions at no cost - see full list of supported distributions. Alternatively, Zabbix is available on certain cloud services.
Templates for Kubernetes monitoring
| Name | Readme | Template |
|---|---|---|
| Kubernetes nodes by HTTP | Readme | Template |
| Kubernetes cluster state by HTTP | Readme | Template |
| Kubernetes API server by HTTP | Readme | Template |
| Kubernetes Controller manager by HTTP | Readme | Template |
| Kubernetes Scheduler by HTTP | Readme | Template |
| Kubernetes kubelet by HTTP | Readme | Template |
Templates are of two types:
Cluster node monitoring
Kubernetes nodes by HTTP template discovers cluster nodes, creates hosts in Zabbix based on prototypes and assigns the "Linux by Zabbix agent" template to them. The template collects basic node metrics via the Kubernetes API.
Main cluster components monitoring
Kubernetes cluster state by HTTP discovers cluster components and control plane nodes, creates Zabbix hosts and assigns the required templates to them.
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
It is strongly recommended to use filtering macros when configuring templates, as on a large cluster the number of discoverable objects can reduce the performance of the monitoring system.
The template to monitor Kubernetes API server that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes API server by HTTP - collects metrics by HTTP agent from API server /metrics endpoint.
Zabbix version: 7.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.API.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: Some metrics may not be collected depending on your Kubernetes API server instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.API.SERVER.URL} | Kubernetes API server metrics endpoint URL. |
https://localhost:6443/metrics |
| {$KUBE.API.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.API.HTTP.SERVER.ERROR} | Maximum number of HTTP server requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get API instance metrics | Get raw metrics from API instance /metrics endpoint. |
HTTP agent | kubernetes.api.get_metrics Preprocessing
|
| Audit events, total | Accumulated number audit events generated and sent to the audit backend. |
Dependent item | kubernetes.api.audit_event_total Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.api.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.api.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.api.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.api.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.api.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.api.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.api.max_fds Preprocessing
|
| gRPCs client started, rate | Total number of RPCs started per second. |
Dependent item | kubernetes.api.grpc_client_started.rate Preprocessing
|
| gRPCs messages received, rate | Total number of gRPC stream messages received per second. |
Dependent item | kubernetes.api.grpc_client_msg_received.rate Preprocessing
|
| gRPCs messages sent, rate | Total number of gRPC stream messages sent per second. |
Dependent item | kubernetes.api.grpc_client_msg_sent.rate Preprocessing
|
| Request terminations, rate | Number of requests which apiserver terminated in self-defense per second. |
Dependent item | kubernetes.api.apiserver_request_terminations Preprocessing
|
| TLS handshake errors, rate | Number of requests dropped with 'TLS handshake error from' error per second. |
Dependent item | kubernetes.api.apiserver_tls_handshake_errors_total.rate Preprocessing
|
| API server requests: 5xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_500.rate Preprocessing
|
| API server requests: 4xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_400.rate Preprocessing
|
| API server requests: 3xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_300.rate Preprocessing
|
| API server requests: 0 | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_0.rate Preprocessing
|
| API server requests: 2xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_200.rate Preprocessing
|
| HTTP requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_500.rate Preprocessing
|
| HTTP requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_400.rate Preprocessing
|
| HTTP requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_300.rate Preprocessing
|
| HTTP requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_200.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API server: Too many server errors | "Kubernetes API server is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.apiserver_request_total_500.rate,5m)>{$KUBE.API.HTTP.SERVER.ERROR} |
Warning | |
| Kubernetes API server: Too many client errors | "Kubernetes API client is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.rest_client_requests_total_500.rate,5m)>{$KUBE.API.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running requests | Discovery of long-running requests by verb, resource and scope. |
Dependent item | kubernetes.api.longrunning_gauge.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running ["{#VERB}"] requests ["{#RESOURCE}"]: {#SCOPE} | Gauge of all active long-running apiserver requests broken out by verb, resource and scope. Not all requests are tracked this way. |
Dependent item | kubernetes.api.longrunning_gauge["{#RESOURCE}","{#SCOPE}","{#VERB}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Request duration histogram | Discovery raw data and percentile items of request duration. |
Dependent item | kubernetes.api.requests_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#VERB}"] Requests bucket: {#LE} | Response latency distribution in seconds for each verb. |
Dependent item | kubernetes.api.request_duration_seconds_bucket[{#LE},"{#VERB}"] Preprocessing
|
| ["{#VERB}"] Requests, p90 | 90 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p90["{#VERB}"] |
| ["{#VERB}"] Requests, p95 | 95 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p95["{#VERB}"] |
| ["{#VERB}"] Requests, p99 | 99 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p99["{#VERB}"] |
| ["{#VERB}"] Requests, p50 | 50 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p50["{#VERB}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests inflight discovery | Discovery requests inflight by kind. |
Dependent item | kubernetes.api.inflight_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests current: {#KIND} | Maximal number of currently used inflight request limit of this apiserver per request kind in last second. |
Dependent item | kubernetes.api.current_inflight_requests["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPC completed requests discovery | Discovery grpc completed requests by grpc code. |
Dependent item | kubernetes.api.grpc_client_handled.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPCs completed: {#GRPC_CODE}, rate | Total number of RPCs completed by the client regardless of success or failure per second. |
Dependent item | kubernetes.api.grpc_client_handled_total.rate["{#GRPC_CODE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts discovery | Discovery authentication attempts by result. |
Dependent item | kubernetes.api.authentication_attempts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts: {#RESULT}, rate | Authentication attempts by result per second. |
Dependent item | kubernetes.api.authentication_attempts.rate["{#RESULT}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication requests discovery | Discovery authentication attempts by name. |
Dependent item | kubernetes.api.authenticated_user_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authenticated requests: {#NAME}, rate | Counter of authenticated requests broken out by username per second. |
Dependent item | kubernetes.api.authenticated_user_requests.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers metrics discovery | Discovery watchers by kind. |
Dependent item | kubernetes.api.apiserver_registered_watchers.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers: {#KIND} | Number of currently registered watchers for a given resource. |
Dependent item | kubernetes.api.apiserver_registered_watchers["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Etcd objects metrics discovery | Discovery etcd objects by resource. |
Dependent item | kubernetes.api.etcd_object_counts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| etcd objects: {#RESOURCE} | Number of stored objects at the time of last check split by kind. |
Dependent item | kubernetes.api.etcd_object_counts["{#RESOURCE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Discovery workqueue metrics by name. |
Dependent item | kubernetes.api.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#NAME}"] Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.api.workqueue_depth["{#NAME}"] Preprocessing
|
| ["{#NAME}"] Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.api.workqueue_adds_total.rate["{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes API server that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes API server by HTTP - collects metrics by HTTP agent from API server /metrics endpoint.
Zabbix version: 7.2 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.API.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. Some metrics may not be collected depending on your Kubernetes API server instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.SERVER.URL} | Kubernetes API server metrics endpoint URL. |
https://localhost:6443/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.API.CERT.EXPIRATION} | Number of days for alert of client certificate used for trigger. |
7 |
| {$KUBE.API.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.API.HTTP.SERVER.ERROR} | Maximum number of HTTP server requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get API instance metrics | Get raw metrics from API instance /metrics endpoint. |
HTTP agent | kubernetes.api.get_metrics Preprocessing
|
| Audit events, total | Accumulated number audit events generated and sent to the audit backend. |
Dependent item | kubernetes.api.audit_event_total Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.api.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.api.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.api.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.api.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.api.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.api.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.api.max_fds Preprocessing
|
| gRPCs client started, rate | Total number of RPCs started per second. |
Dependent item | kubernetes.api.grpc_client_started.rate Preprocessing
|
| gRPCs messages received, rate | Total number of gRPC stream messages received per second. |
Dependent item | kubernetes.api.grpc_client_msg_received.rate Preprocessing
|
| gRPCs messages sent, rate | Total number of gRPC stream messages sent per second. |
Dependent item | kubernetes.api.grpc_client_msg_sent.rate Preprocessing
|
| Request terminations, rate | Number of requests which apiserver terminated in self-defense per second. |
Dependent item | kubernetes.api.apiserver_request_terminations Preprocessing
|
| TLS handshake errors, rate | Number of requests dropped with 'TLS handshake error from' error per second. |
Dependent item | kubernetes.api.apiserver_tls_handshake_errors_total.rate Preprocessing
|
| API server requests: 5xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_500.rate Preprocessing
|
| API server requests: 4xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_400.rate Preprocessing
|
| API server requests: 3xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_300.rate Preprocessing
|
| API server requests: 0 | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_0.rate Preprocessing
|
| API server requests: 2xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_200.rate Preprocessing
|
| HTTP requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_500.rate Preprocessing
|
| HTTP requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_400.rate Preprocessing
|
| HTTP requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_300.rate Preprocessing
|
| HTTP requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_200.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API server: Too many server errors | "Kubernetes API server is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.apiserver_request_total_500.rate,5m)>{$KUBE.API.HTTP.SERVER.ERROR} |
Warning | |
| Kubernetes API server: Too many client errors | "Kubernetes API client is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.rest_client_requests_total_500.rate,5m)>{$KUBE.API.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running requests | Discovery of long-running requests by verb, resource and scope. |
Dependent item | kubernetes.api.longrunning_gauge.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running ["{#VERB}"] requests ["{#RESOURCE}"]: {#SCOPE} | Gauge of all active long-running apiserver requests broken out by verb, resource and scope. Not all requests are tracked this way. |
Dependent item | kubernetes.api.longrunning_gauge["{#RESOURCE}","{#SCOPE}","{#VERB}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Request duration histogram | Discovery raw data and percentile items of request duration. |
Dependent item | kubernetes.api.requests_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#VERB}"] Requests bucket: {#LE} | Response latency distribution in seconds for each verb. |
Dependent item | kubernetes.api.request_duration_seconds_bucket[{#LE},"{#VERB}"] Preprocessing
|
| ["{#VERB}"] Requests, p90 | 90 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p90["{#VERB}"] |
| ["{#VERB}"] Requests, p95 | 95 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p95["{#VERB}"] |
| ["{#VERB}"] Requests, p99 | 99 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p99["{#VERB}"] |
| ["{#VERB}"] Requests, p50 | 50 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p50["{#VERB}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests inflight discovery | Discovery requests inflight by kind. |
Dependent item | kubernetes.api.inflight_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests current: {#KIND} | Maximal number of currently used inflight request limit of this apiserver per request kind in last second. |
Dependent item | kubernetes.api.current_inflight_requests["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPC completed requests discovery | Discovery grpc completed requests by grpc code. |
Dependent item | kubernetes.api.grpc_client_handled.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPCs completed: {#GRPC_CODE}, rate | Total number of RPCs completed by the client regardless of success or failure per second. |
Dependent item | kubernetes.api.grpc_client_handled_total.rate["{#GRPC_CODE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts discovery | Discovery authentication attempts by result. |
Dependent item | kubernetes.api.authentication_attempts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts: {#RESULT}, rate | Authentication attempts by result per second. |
Dependent item | kubernetes.api.authentication_attempts.rate["{#RESULT}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication requests discovery | Discovery authentication attempts by name. |
Dependent item | kubernetes.api.authenticated_user_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authenticated requests: {#NAME}, rate | Counter of authenticated requests broken out by username per second. |
Dependent item | kubernetes.api.authenticated_user_requests.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers metrics discovery | Discovery watchers by kind. |
Dependent item | kubernetes.api.apiserver_registered_watchers.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers: {#KIND} | Number of currently registered watchers for a given resource. |
Dependent item | kubernetes.api.apiserver_registered_watchers["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Etcd objects metrics discovery | Discovery etcd objects by resource. |
Dependent item | kubernetes.api.etcd_object_counts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| etcd objects: {#RESOURCE} | Number of stored objects at the time of last check split by kind. |
Dependent item | kubernetes.api.etcd_object_counts["{#RESOURCE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Discovery workqueue metrics by name. |
Dependent item | kubernetes.api.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#NAME}"] Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.api.workqueue_depth["{#NAME}"] Preprocessing
|
| ["{#NAME}"] Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.api.workqueue_adds_total.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Client certificate expiration histogram | Discovery raw data of client certificate expiration |
Dependent item | kubernetes.api.certificate_expiration.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Certificate expiration seconds bucket, {#LE} | Distribution of the remaining lifetime on the certificate used to authenticate a request. |
Dependent item | kubernetes.api.client_certificate_expiration_seconds_bucket[{#LE}] Preprocessing
|
| Client certificate expiration, p1 | 1 percentile of the remaining lifetime on the certificate used to authenticate a request. |
Calculated | kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}] |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API server: Kubernetes client certificate is expiring | A client certificate used to authenticate to the apiserver is expiring in {$KUBE.API.CERT.EXPIRATION} days. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < {$KUBE.API.CERT.EXPIRATION}*24*60*60 |
Warning | Depends on:
|
| Kubernetes API server: Kubernetes client certificate expires soon | A client certificate used to authenticate to the apiserver is expiring in less than 24.0 hours. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < 24*60*60 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes API server that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes API server by HTTP - collects metrics by HTTP agent from API server /metrics endpoint.
Zabbix version: 7.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.API.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: Some metrics may not be collected depending on your Kubernetes API server instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.SERVER.URL} | Kubernetes API server metrics endpoint URL. |
https://localhost:6443/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.API.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.API.HTTP.SERVER.ERROR} | Maximum number of HTTP server requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get API instance metrics | Get raw metrics from API instance /metrics endpoint. |
HTTP agent | kubernetes.api.get_metrics Preprocessing
|
| Audit events, total | Accumulated number audit events generated and sent to the audit backend. |
Dependent item | kubernetes.api.audit_event_total Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.api.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.api.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.api.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.api.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.api.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.api.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.api.max_fds Preprocessing
|
| gRPCs client started, rate | Total number of RPCs started per second. |
Dependent item | kubernetes.api.grpc_client_started.rate Preprocessing
|
| gRPCs messages received, rate | Total number of gRPC stream messages received per second. |
Dependent item | kubernetes.api.grpc_client_msg_received.rate Preprocessing
|
| gRPCs messages sent, rate | Total number of gRPC stream messages sent per second. |
Dependent item | kubernetes.api.grpc_client_msg_sent.rate Preprocessing
|
| Request terminations, rate | Number of requests which apiserver terminated in self-defense per second. |
Dependent item | kubernetes.api.apiserver_request_terminations Preprocessing
|
| TLS handshake errors, rate | Number of requests dropped with 'TLS handshake error from' error per second. |
Dependent item | kubernetes.api.apiserver_tls_handshake_errors_total.rate Preprocessing
|
| API server requests: 5xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_500.rate Preprocessing
|
| API server requests: 4xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_400.rate Preprocessing
|
| API server requests: 3xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_300.rate Preprocessing
|
| API server requests: 0 | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_0.rate Preprocessing
|
| API server requests: 2xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_200.rate Preprocessing
|
| HTTP requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_500.rate Preprocessing
|
| HTTP requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_400.rate Preprocessing
|
| HTTP requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_300.rate Preprocessing
|
| HTTP requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_200.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API server: Too many server errors | "Kubernetes API server is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.apiserver_request_total_500.rate,5m)>{$KUBE.API.HTTP.SERVER.ERROR} |
Warning | |
| Kubernetes API server: Too many client errors | "Kubernetes API client is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.rest_client_requests_total_500.rate,5m)>{$KUBE.API.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running requests | Discovery of long-running requests by verb, resource and scope. |
Dependent item | kubernetes.api.longrunning_gauge.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running ["{#VERB}"] requests ["{#RESOURCE}"]: {#SCOPE} | Gauge of all active long-running apiserver requests broken out by verb, resource and scope. Not all requests are tracked this way. |
Dependent item | kubernetes.api.longrunning_gauge["{#RESOURCE}","{#SCOPE}","{#VERB}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Request duration histogram | Discovery raw data and percentile items of request duration. |
Dependent item | kubernetes.api.requests_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#VERB}"] Requests bucket: {#LE} | Response latency distribution in seconds for each verb. |
Dependent item | kubernetes.api.request_duration_seconds_bucket[{#LE},"{#VERB}"] Preprocessing
|
| ["{#VERB}"] Requests, p90 | 90 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p90["{#VERB}"] |
| ["{#VERB}"] Requests, p95 | 95 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p95["{#VERB}"] |
| ["{#VERB}"] Requests, p99 | 99 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p99["{#VERB}"] |
| ["{#VERB}"] Requests, p50 | 50 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p50["{#VERB}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests inflight discovery | Discovery requests inflight by kind. |
Dependent item | kubernetes.api.inflight_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests current: {#KIND} | Maximal number of currently used inflight request limit of this apiserver per request kind in last second. |
Dependent item | kubernetes.api.current_inflight_requests["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPC completed requests discovery | Discovery grpc completed requests by grpc code. |
Dependent item | kubernetes.api.grpc_client_handled.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPCs completed: {#GRPC_CODE}, rate | Total number of RPCs completed by the client regardless of success or failure per second. |
Dependent item | kubernetes.api.grpc_client_handled_total.rate["{#GRPC_CODE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts discovery | Discovery authentication attempts by result. |
Dependent item | kubernetes.api.authentication_attempts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts: {#RESULT}, rate | Authentication attempts by result per second. |
Dependent item | kubernetes.api.authentication_attempts.rate["{#RESULT}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication requests discovery | Discovery authentication attempts by name. |
Dependent item | kubernetes.api.authenticated_user_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authenticated requests: {#NAME}, rate | Counter of authenticated requests broken out by username per second. |
Dependent item | kubernetes.api.authenticated_user_requests.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers metrics discovery | Discovery watchers by kind. |
Dependent item | kubernetes.api.apiserver_registered_watchers.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers: {#KIND} | Number of currently registered watchers for a given resource. |
Dependent item | kubernetes.api.apiserver_registered_watchers["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Etcd objects metrics discovery | Discovery etcd objects by resource. |
Dependent item | kubernetes.api.etcd_object_counts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| etcd objects: {#RESOURCE} | Number of stored objects at the time of last check split by kind. |
Dependent item | kubernetes.api.etcd_object_counts["{#RESOURCE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Discovery workqueue metrics by name. |
Dependent item | kubernetes.api.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#NAME}"] Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.api.workqueue_depth["{#NAME}"] Preprocessing
|
| ["{#NAME}"] Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.api.workqueue_adds_total.rate["{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes API server that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes API server by HTTP - collects metrics by HTTP agent from API server /metrics endpoint.
Zabbix version: 6.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.API.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. Some metrics may not be collected depending on your Kubernetes API server instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.SERVER.URL} | Kubernetes API server metrics endpoint URL. |
https://localhost:6443/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.API.CERT.EXPIRATION} | Number of days for alert of client certificate used for trigger. |
7 |
| {$KUBE.API.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.API.HTTP.SERVER.ERROR} | Maximum number of HTTP server requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Get API instance metrics | Get raw metrics from API instance /metrics endpoint. |
HTTP agent | kubernetes.api.get_metrics Preprocessing
|
| Kubernetes API: Audit events, total | Accumulated number audit events generated and sent to the audit backend. |
Dependent item | kubernetes.api.audit_event_total Preprocessing
|
| Kubernetes API: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.api.process_virtual_memory_bytes Preprocessing
|
| Kubernetes API: Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.api.process_resident_memory_bytes Preprocessing
|
| Kubernetes API: CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.api.cpu.util Preprocessing
|
| Kubernetes API: Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.api.go_goroutines Preprocessing
|
| Kubernetes API: Go threads | Number of OS threads created. |
Dependent item | kubernetes.api.go_threads Preprocessing
|
| Kubernetes API: Fds open | Number of open file descriptors. |
Dependent item | kubernetes.api.open_fds Preprocessing
|
| Kubernetes API: Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.api.max_fds Preprocessing
|
| Kubernetes API: gRPCs client started, rate | Total number of RPCs started per second. |
Dependent item | kubernetes.api.grpc_client_started.rate Preprocessing
|
| Kubernetes API: gRPCs messages received, rate | Total number of gRPC stream messages received per second. |
Dependent item | kubernetes.api.grpc_client_msg_received.rate Preprocessing
|
| Kubernetes API: gRPCs messages sent, rate | Total number of gRPC stream messages sent per second. |
Dependent item | kubernetes.api.grpc_client_msg_sent.rate Preprocessing
|
| Kubernetes API: Request terminations, rate | Number of requests which apiserver terminated in self-defense per second. |
Dependent item | kubernetes.api.apiserver_request_terminations Preprocessing
|
| Kubernetes API: TLS handshake errors, rate | Number of requests dropped with 'TLS handshake error from' error per second. |
Dependent item | kubernetes.api.apiserver_tls_handshake_errors_total.rate Preprocessing
|
| Kubernetes API: API server requests: 5xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_500.rate Preprocessing
|
| Kubernetes API: API server requests: 4xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_400.rate Preprocessing
|
| Kubernetes API: API server requests: 3xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_300.rate Preprocessing
|
| Kubernetes API: API server requests: 0 | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_0.rate Preprocessing
|
| Kubernetes API: API server requests: 2xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_200.rate Preprocessing
|
| Kubernetes API: HTTP requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_500.rate Preprocessing
|
| Kubernetes API: HTTP requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_400.rate Preprocessing
|
| Kubernetes API: HTTP requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_300.rate Preprocessing
|
| Kubernetes API: HTTP requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_200.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API: Too many server errors | "Kubernetes API server is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.apiserver_request_total_500.rate,5m)>{$KUBE.API.HTTP.SERVER.ERROR} |
Warning | |
| Kubernetes API: Too many client errors | "Kubernetes API client is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.rest_client_requests_total_500.rate,5m)>{$KUBE.API.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running requests | Discovery of long-running requests by verb, resource and scope. |
Dependent item | kubernetes.api.longrunning_gauge.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Long-running ["{#VERB}"] requests ["{#RESOURCE}"]: {#SCOPE} | Gauge of all active long-running apiserver requests broken out by verb, resource and scope. Not all requests are tracked this way. |
Dependent item | kubernetes.api.longrunning_gauge["{#RESOURCE}","{#SCOPE}","{#VERB}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Request duration histogram | Discovery raw data and percentile items of request duration. |
Dependent item | kubernetes.api.requests_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: ["{#VERB}"] Requests bucket: {#LE} | Response latency distribution in seconds for each verb. |
Dependent item | kubernetes.api.request_duration_seconds_bucket[{#LE},"{#VERB}"] Preprocessing
|
| Kubernetes API: ["{#VERB}"] Requests, p90 | 90 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p90["{#VERB}"] |
| Kubernetes API: ["{#VERB}"] Requests, p95 | 95 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p95["{#VERB}"] |
| Kubernetes API: ["{#VERB}"] Requests, p99 | 99 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p99["{#VERB}"] |
| Kubernetes API: ["{#VERB}"] Requests, p50 | 50 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p50["{#VERB}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests inflight discovery | Discovery requests inflight by kind. |
Dependent item | kubernetes.api.inflight_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Requests current: {#KIND} | Maximal number of currently used inflight request limit of this apiserver per request kind in last second. |
Dependent item | kubernetes.api.current_inflight_requests["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPC completed requests discovery | Discovery grpc completed requests by grpc code. |
Dependent item | kubernetes.api.grpc_client_handled.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: gRPCs completed: {#GRPC_CODE}, rate | Total number of RPCs completed by the client regardless of success or failure per second. |
Dependent item | kubernetes.api.grpc_client_handled_total.rate["{#GRPC_CODE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts discovery | Discovery authentication attempts by result. |
Dependent item | kubernetes.api.authentication_attempts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Authentication attempts: {#RESULT}, rate | Authentication attempts by result per second. |
Dependent item | kubernetes.api.authentication_attempts.rate["{#RESULT}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication requests discovery | Discovery authentication attempts by name. |
Dependent item | kubernetes.api.authenticated_user_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Authenticated requests: {#NAME}, rate | Counter of authenticated requests broken out by username per second. |
Dependent item | kubernetes.api.authenticated_user_requests.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers metrics discovery | Discovery watchers by kind. |
Dependent item | kubernetes.api.apiserver_registered_watchers.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Watchers: {#KIND} | Number of currently registered watchers for a given resource. |
Dependent item | kubernetes.api.apiserver_registered_watchers["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Etcd objects metrics discovery | Discovery etcd objects by resource. |
Dependent item | kubernetes.api.etcd_object_counts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: etcd objects: {#RESOURCE} | Number of stored objects at the time of last check split by kind. |
Dependent item | kubernetes.api.etcd_object_counts["{#RESOURCE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Discovery workqueue metrics by name. |
Dependent item | kubernetes.api.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: ["{#NAME}"] Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.api.workqueue_depth["{#NAME}"] Preprocessing
|
| Kubernetes API: ["{#NAME}"] Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.api.workqueue_adds_total.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Client certificate expiration histogram | Discovery raw data of client certificate expiration |
Dependent item | kubernetes.api.certificate_expiration.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Certificate expiration seconds bucket, {#LE} | Distribution of the remaining lifetime on the certificate used to authenticate a request. |
Dependent item | kubernetes.api.client_certificate_expiration_seconds_bucket[{#LE}] Preprocessing
|
| Kubernetes API: Client certificate expiration, p1 | 1 percentile of the remaining lifetime on the certificate used to authenticate a request. |
Calculated | kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}] |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API: Kubernetes client certificate is expiring | A client certificate used to authenticate to the apiserver is expiring in {$KUBE.API.CERT.EXPIRATION} days. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < {$KUBE.API.CERT.EXPIRATION}*24*60*60 |
Warning | Depends on:
|
| Kubernetes API: Kubernetes client certificate expires soon | A client certificate used to authenticate to the apiserver is expiring in less than 24.0 hours. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < 24*60*60 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
For Zabbix version: 6.2 and higher
The template to monitor InfluxDB by Zabbix that works without any external scripts.
Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes API server by HTTP — collects metrics by HTTP agent from API server /metrics endpoint.
This template was tested on:
See Zabbix template operation for basic instructions.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.API.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values. NOTE. Some metrics may not be collected depending on your Kubernetes API server instance version and configuration.
No specific Zabbix configuration is required.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.CERT.EXPIRATION} | Number of days for alert of client certificate used for trigger |
7 |
| {$KUBE.API.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger |
2 |
| {$KUBE.API.HTTP.SERVER.ERROR} | Maximum number of HTTP client requests failures used for trigger |
2 |
| {$KUBE.API.SERVER.URL} | instance URL |
http://localhost:8086/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token |
`` |
There are no template links in this template.
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts discovery | Discovery authentication attempts by result. |
DEPENDENT | kubernetes.api.authentication_attempts.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Authentication requests discovery | Discovery authentication attempts by name. |
DEPENDENT | kubernetes.api.authenticated_user_requests.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Client certificate expiration histogram | Discovery raw data of client certificate expiration |
DEPENDENT | kubernetes.api.certificate_expiration.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| Etcd objects metrics discovery | Discovery etcd objects by resource. |
DEPENDENT | kubernetes.api.etcd_object_counts.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| gRPC completed requests discovery | Discovery grpc completed requests by grpc code. |
DEPENDENT | kubernetes.api.grpc_client_handled.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Long-running requests | Discovery of long-running requests by verb, resource and scope. |
DEPENDENT | kubernetes.api.longrunning_gauge.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Request duration histogram | Discovery raw data and percentile items of request duration. |
DEPENDENT | kubernetes.api.requests_bucket.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| Requests inflight discovery | Discovery requests inflight by kind. |
DEPENDENT | kubernetes.api.inflight_requests.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Watchers metrics discovery | Discovery watchers by kind. |
DEPENDENT | kubernetes.api.apiserver_registered_watchers.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Workqueue metrics discovery | Discovery workqueue metrics by name. |
DEPENDENT | kubernetes.api.workqueue.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
| Group | Name | Description | Type | Key and additional info |
|---|---|---|---|---|
| Kubernetes API | Kubernetes API: Audit events, total | Accumulated number audit events generated and sent to the audit backend. |
DEPENDENT | kubernetes.api.audit_event_total Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: Virtual memory, bytes | Virtual memory size in bytes. |
DEPENDENT | kubernetes.api.process_virtual_memory_bytes Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: Resident memory, bytes | Resident memory size in bytes. |
DEPENDENT | kubernetes.api.process_resident_memory_bytes Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: CPU | Total user and system CPU usage ratio. |
DEPENDENT | kubernetes.api.cpu.util Preprocessing: - PROMETHEUS_PATTERN: - CHANGE_PER_SECOND - MULTIPLIER: |
| Kubernetes API | Kubernetes API: Goroutines | Number of goroutines that currently exist. |
DEPENDENT | kubernetes.api.go_goroutines Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: Go threads | Number of OS threads created. |
DEPENDENT | kubernetes.api.go_threads Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: Fds open | Number of open file descriptors. |
DEPENDENT | kubernetes.api.open_fds Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: Fds max | Maximum allowed open file descriptors. |
DEPENDENT | kubernetes.api.max_fds Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: gRPCs client started, rate | Total number of RPCs started per second. |
DEPENDENT | kubernetes.api.grpc_client_started.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: gRPCs messages ressived, rate | Total number of gRPC stream messages received per second. |
DEPENDENT | kubernetes.api.grpc_client_msg_received.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: gRPCs messages sent, rate | Total number of gRPC stream messages sent per second. |
DEPENDENT | kubernetes.api.grpc_client_msg_sent.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: Request terminations, rate | Number of requests which apiserver terminated in self-defense per second. |
DEPENDENT | kubernetes.api.apiserver_request_terminations Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: TLS handshake errors, rate | Number of requests dropped with 'TLS handshake error from' error per second. |
DEPENDENT | kubernetes.api.apiserver_tls_handshake_errors_total.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: API server requests: 5xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
DEPENDENT | kubernetes.api.apiserver_request_total_500.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: API server requests: 4xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
DEPENDENT | kubernetes.api.apiserver_request_total_400.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: API server requests: 3xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
DEPENDENT | kubernetes.api.apiserver_request_total_300.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: API server requests: 0 | Counter of apiserver requests broken out for each HTTP response code. |
DEPENDENT | kubernetes.api.apiserver_request_total_0.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: API server requests: 2xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
DEPENDENT | kubernetes.api.apiserver_request_total_200.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: HTTP requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
DEPENDENT | kubernetes.api.rest_client_requests_total_500.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: HTTP requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
DEPENDENT | kubernetes.api.rest_client_requests_total_400.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: HTTP requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
DEPENDENT | kubernetes.api.rest_client_requests_total_300.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: HTTP requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
DEPENDENT | kubernetes.api.rest_client_requests_total_200.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: Long-running ["{#VERB}"] requests ["{#RESOURCE}"]: {#SCOPE} | Gauge of all active long-running apiserver requests broken out by verb, resource and scope. Not all requests are tracked this way. |
DEPENDENT | kubernetes.api.longrunning_gauge["{#RESOURCE}","{#SCOPE}","{#VERB}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: ["{#VERB}"] Requests bucket: {#LE} | Response latency distribution in seconds for each verb. |
DEPENDENT | kubernetes.api.request_duration_seconds_bucket[{#LE},"{#VERB}"] Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes API | Kubernetes API: ["{#VERB}"] Requests, p90 | 90 percentile of response latency distribution in seconds for each verb. |
CALCULATED | kubernetes.api.request_duration_seconds_p90["{#VERB}"] Expression: bucket_percentile(//kubernetes.api.request_duration_seconds_bucket[*,"{#VERB}"],5m,90) |
| Kubernetes API | Kubernetes API: ["{#VERB}"] Requests, p95 | 95 percentile of response latency distribution in seconds for each verb. |
CALCULATED | kubernetes.api.request_duration_seconds_p95["{#VERB}"] Expression: bucket_percentile(//kubernetes.api.request_duration_seconds_bucket[*,"{#VERB}"],5m,95) |
| Kubernetes API | Kubernetes API: ["{#VERB}"] Requests, p99 | 99 percentile of response latency distribution in seconds for each verb. |
CALCULATED | kubernetes.api.request_duration_seconds_p99["{#VERB}"] Expression: bucket_percentile(//kubernetes.api.request_duration_seconds_bucket[*,"{#VERB}"],5m,99) |
| Kubernetes API | Kubernetes API: ["{#VERB}"] Requests, p50 | 50 percentile of response latency distribution in seconds for each verb. |
CALCULATED | kubernetes.api.request_duration_seconds_p50["{#VERB}"] Expression: bucket_percentile(//kubernetes.api.request_duration_seconds_bucket[*,"{#VERB}"],5m,50) |
| Kubernetes API | Kubernetes API: Requests current: {#KIND} | Maximal number of currently used inflight request limit of this apiserver per request kind in last second. |
DEPENDENT | kubernetes.api.current_inflight_requests["{#KIND}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: gRPCs completed: {#GRPC_CODE}, rate | Total number of RPCs completed by the client regardless of success or failure per second. |
DEPENDENT | kubernetes.api.grpc_client_handled_total.rate["{#GRPC_CODE}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: Authentication attempts: {#RESULT}, rate | Authentication attempts by result per second. |
DEPENDENT | kubernetes.api.authentication_attempts.rate["{#RESULT}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: Authenticated requests: {#NAME}, rate | Counter of authenticated requests broken out by username per second. |
DEPENDENT | kubernetes.api.authenticated_user_requests.rate["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: Watchers: {#KIND} | Number of currently registered watchers for a given resource. |
DEPENDENT | kubernetes.api.apiserver_registered_watchers["{#KIND}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: etcd objects: {#RESOURCE} | Number of stored objects at the time of last check split by kind. |
DEPENDENT | kubernetes.api.etcd_object_counts["{#RESOURCE}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: ["{#NAME}"] Workqueue depth | Current depth of workqueue. |
DEPENDENT | kubernetes.api.workqueue_depth["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: ["{#NAME}"] Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
DEPENDENT | kubernetes.api.workqueue_adds_total.rate["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes API | Kubernetes API: Certificate expiration seconds bucket, {#LE} | Distribution of the remaining lifetime on the certificate used to authenticate a request. |
DEPENDENT | kubernetes.api.client_certificate_expiration_seconds_bucket[{#LE}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes API | Kubernetes API: Client certificate expiration, p1 | 1 percentile of the remaining lifetime on the certificate used to authenticate a request. |
CALCULATED | kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.api.client_certificate_expiration_seconds_bucket[*],5m,1) |
| Zabbix raw items | Kubernetes API: Get API instance metrics | Get raw metrics from API instance /metrics endpoint. |
HTTP_AGENT | kubernetes.api.get_metrics Preprocessing: - CHECK_NOT_SUPPORTED ⛔️ON_FAIL: |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API: Too many server errors | "Kubernetes API server is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.apiserver_request_total_500.rate,5m)>{$KUBE.API.HTTP.SERVER.ERROR} |
WARNING | |
| Kubernetes API: Too many client errors | "Kubernetes API client is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.rest_client_requests_total_500.rate,5m)>{$KUBE.API.HTTP.CLIENT.ERROR} |
WARNING | |
| Kubernetes API: Kubernetes client certificate is expiring | A client certificate used to authenticate to the apiserver is expiring in {$KUBE.API.CERT.EXPIRATION} days. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < {$KUBE.API.CERT.EXPIRATION}*24*60*60 |
WARNING | Depends on: - Kubernetes API: Kubernetes client certificate expires soon |
| Kubernetes API: Kubernetes client certificate expires soon | A client certificate used to authenticate to the apiserver is expiring in less than 24.0 hours. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < 24*60*60 |
WARNING |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template or ask for help with it at ZABBIX forums.
The template to monitor Kubernetes API server that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes API server by HTTP - collects metrics by HTTP agent from API server /metrics endpoint.
Zabbix version: 6.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.API.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. Some metrics may not be collected depending on your Kubernetes API server instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.SERVER.URL} | Kubernetes API server metrics endpoint URL. |
https://localhost:6443/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.API.CERT.EXPIRATION} | Number of days for alert of client certificate used for trigger. |
7 |
| {$KUBE.API.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.API.HTTP.SERVER.ERROR} | Maximum number of HTTP server requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Get API instance metrics | Get raw metrics from API instance /metrics endpoint. |
HTTP agent | kubernetes.api.get_metrics Preprocessing
|
| Kubernetes API: Audit events, total | Accumulated number audit events generated and sent to the audit backend. |
Dependent item | kubernetes.api.audit_event_total Preprocessing
|
| Kubernetes API: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.api.process_virtual_memory_bytes Preprocessing
|
| Kubernetes API: Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.api.process_resident_memory_bytes Preprocessing
|
| Kubernetes API: CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.api.cpu.util Preprocessing
|
| Kubernetes API: Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.api.go_goroutines Preprocessing
|
| Kubernetes API: Go threads | Number of OS threads created. |
Dependent item | kubernetes.api.go_threads Preprocessing
|
| Kubernetes API: Fds open | Number of open file descriptors. |
Dependent item | kubernetes.api.open_fds Preprocessing
|
| Kubernetes API: Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.api.max_fds Preprocessing
|
| Kubernetes API: gRPCs client started, rate | Total number of RPCs started per second. |
Dependent item | kubernetes.api.grpc_client_started.rate Preprocessing
|
| Kubernetes API: gRPCs messages received, rate | Total number of gRPC stream messages received per second. |
Dependent item | kubernetes.api.grpc_client_msg_received.rate Preprocessing
|
| Kubernetes API: gRPCs messages sent, rate | Total number of gRPC stream messages sent per second. |
Dependent item | kubernetes.api.grpc_client_msg_sent.rate Preprocessing
|
| Kubernetes API: Request terminations, rate | Number of requests which apiserver terminated in self-defense per second. |
Dependent item | kubernetes.api.apiserver_request_terminations Preprocessing
|
| Kubernetes API: TLS handshake errors, rate | Number of requests dropped with 'TLS handshake error from' error per second. |
Dependent item | kubernetes.api.apiserver_tls_handshake_errors_total.rate Preprocessing
|
| Kubernetes API: API server requests: 5xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_500.rate Preprocessing
|
| Kubernetes API: API server requests: 4xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_400.rate Preprocessing
|
| Kubernetes API: API server requests: 3xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_300.rate Preprocessing
|
| Kubernetes API: API server requests: 0 | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_0.rate Preprocessing
|
| Kubernetes API: API server requests: 2xx, rate | Counter of apiserver requests broken out for each HTTP response code. |
Dependent item | kubernetes.api.apiserver_request_total_200.rate Preprocessing
|
| Kubernetes API: HTTP requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_500.rate Preprocessing
|
| Kubernetes API: HTTP requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_400.rate Preprocessing
|
| Kubernetes API: HTTP requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_300.rate Preprocessing
|
| Kubernetes API: HTTP requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.api.rest_client_requests_total_200.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API: Too many server errors | "Kubernetes API server is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.apiserver_request_total_500.rate,5m)>{$KUBE.API.HTTP.SERVER.ERROR} |
Warning | |
| Kubernetes API: Too many client errors | "Kubernetes API client is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes API server by HTTP/kubernetes.api.rest_client_requests_total_500.rate,5m)>{$KUBE.API.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Long-running requests | Discovery of long-running requests by verb, resource and scope. |
Dependent item | kubernetes.api.longrunning_gauge.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Long-running ["{#VERB}"] requests ["{#RESOURCE}"]: {#SCOPE} | Gauge of all active long-running apiserver requests broken out by verb, resource and scope. Not all requests are tracked this way. |
Dependent item | kubernetes.api.longrunning_gauge["{#RESOURCE}","{#SCOPE}","{#VERB}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Request duration histogram | Discovery raw data and percentile items of request duration. |
Dependent item | kubernetes.api.requests_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: ["{#VERB}"] Requests bucket: {#LE} | Response latency distribution in seconds for each verb. |
Dependent item | kubernetes.api.request_duration_seconds_bucket[{#LE},"{#VERB}"] Preprocessing
|
| Kubernetes API: ["{#VERB}"] Requests, p90 | 90 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p90["{#VERB}"] |
| Kubernetes API: ["{#VERB}"] Requests, p95 | 95 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p95["{#VERB}"] |
| Kubernetes API: ["{#VERB}"] Requests, p99 | 99 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p99["{#VERB}"] |
| Kubernetes API: ["{#VERB}"] Requests, p50 | 50 percentile of response latency distribution in seconds for each verb. |
Calculated | kubernetes.api.request_duration_seconds_p50["{#VERB}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Requests inflight discovery | Discovery requests inflight by kind. |
Dependent item | kubernetes.api.inflight_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Requests current: {#KIND} | Maximal number of currently used inflight request limit of this apiserver per request kind in last second. |
Dependent item | kubernetes.api.current_inflight_requests["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| gRPC completed requests discovery | Discovery grpc completed requests by grpc code. |
Dependent item | kubernetes.api.grpc_client_handled.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: gRPCs completed: {#GRPC_CODE}, rate | Total number of RPCs completed by the client regardless of success or failure per second. |
Dependent item | kubernetes.api.grpc_client_handled_total.rate["{#GRPC_CODE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication attempts discovery | Discovery authentication attempts by result. |
Dependent item | kubernetes.api.authentication_attempts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Authentication attempts: {#RESULT}, rate | Authentication attempts by result per second. |
Dependent item | kubernetes.api.authentication_attempts.rate["{#RESULT}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Authentication requests discovery | Discovery authentication attempts by name. |
Dependent item | kubernetes.api.authenticated_user_requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Authenticated requests: {#NAME}, rate | Counter of authenticated requests broken out by username per second. |
Dependent item | kubernetes.api.authenticated_user_requests.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Watchers metrics discovery | Discovery watchers by kind. |
Dependent item | kubernetes.api.apiserver_registered_watchers.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Watchers: {#KIND} | Number of currently registered watchers for a given resource. |
Dependent item | kubernetes.api.apiserver_registered_watchers["{#KIND}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Etcd objects metrics discovery | Discovery etcd objects by resource. |
Dependent item | kubernetes.api.etcd_object_counts.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: etcd objects: {#RESOURCE} | Number of stored objects at the time of last check split by kind. |
Dependent item | kubernetes.api.etcd_object_counts["{#RESOURCE}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Discovery workqueue metrics by name. |
Dependent item | kubernetes.api.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: ["{#NAME}"] Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.api.workqueue_depth["{#NAME}"] Preprocessing
|
| Kubernetes API: ["{#NAME}"] Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.api.workqueue_adds_total.rate["{#NAME}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Client certificate expiration histogram | Discovery raw data of client certificate expiration |
Dependent item | kubernetes.api.certificate_expiration.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes API: Certificate expiration seconds bucket, {#LE} | Distribution of the remaining lifetime on the certificate used to authenticate a request. |
Dependent item | kubernetes.api.client_certificate_expiration_seconds_bucket[{#LE}] Preprocessing
|
| Kubernetes API: Client certificate expiration, p1 | 1 percentile of the remaining lifetime on the certificate used to authenticate a request. |
Calculated | kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}] |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes API: Kubernetes client certificate is expiring | A client certificate used to authenticate to the apiserver is expiring in {$KUBE.API.CERT.EXPIRATION} days. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < {$KUBE.API.CERT.EXPIRATION}*24*60*60 |
Warning | Depends on:
|
| Kubernetes API: Kubernetes client certificate expires soon | A client certificate used to authenticate to the apiserver is expiring in less than 24.0 hours. |
last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) > 0 and last(/Kubernetes API server by HTTP/kubernetes.api.client_certificate_expiration_p1[{#SINGLETON}]) < 24*60*60 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP - collects metrics by HTTP agent from Controller manager /metrics endpoint.
Zabbix version: 7.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.CONTROLLER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: You might need to set the --binding-address option for Controller Manager to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
Note: Some metrics may not be collected depending on your Kubernetes Controller manager instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | API Authorization Token |
|
| {$KUBE.CONTROLLER.SERVER.URL} | Kubernetes Controller manager metrics endpoint URL. |
https://localhost:10257/metrics |
| {$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller: Get Controller metrics | Get raw metrics from Controller instance /metrics endpoint. |
HTTP agent | kubernetes.controller.get_metrics Preprocessing
|
| Leader election status | Gauge of if the reporting system is master of the relevant lease, 0 indicates backup, 1 indicates master. |
Dependent item | kubernetes.controller.leader_election_master_status Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.controller.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.controller.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.controller.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.controller.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.controller.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.controller.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.controller.max_fds Preprocessing
|
| REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_200.rate Preprocessing
|
| REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_300.rate Preprocessing
|
| REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_400.rate Preprocessing
|
| REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_500.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Controller manager: Too many HTTP client errors | "Kubernetes Controller manager is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Controller manager by HTTP/kubernetes.controller.client_http_requests_500.rate,5m)>{$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Dependent item | kubernetes.controller.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#NAME}"]: Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_adds_total["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.controller.workqueue_depth["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue unfinished work, sec | How many seconds of work has done that is in progress and hasn't been observed by work_duration. Large values indicate stuck threads. One can deduce the number of stuck threads by observing the rate at which this increases. |
Dependent item | kubernetes.controller.workqueue_unfinished_work_seconds["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue retries, rate | Total number of retries handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_retries_total["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue longest running processor, sec | How many seconds has the longest running processor for workqueue been running. |
Dependent item | kubernetes.controller.workqueue_longest_running_processor_seconds["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue work duration, p90 | 90 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p90["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, p95 | 95 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p95["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, p99 | 99 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p99["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, 50p | 50 percentiles of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p50["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p90 | 90 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p90["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p95 | 95 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p95["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p99 | 99 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p99["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, 50p | 50 percentile of how long in seconds an item stays in workqueue before being requested. If there are no requests for 5 minute, item value will be discarded. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p50["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue duration seconds bucket, {#LE} | How long in seconds processing an item from workqueue takes. |
Dependent item | kubernetes.controller.duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Queue duration seconds bucket, {#LE} | How long in seconds an item stays in workqueue before being requested. |
Dependent item | kubernetes.controller.queue_duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP - collects metrics by HTTP agent from Controller manager /metrics endpoint.
Zabbix version: 7.2 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.CONTROLLER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. You might need to set the --binding-address option for Controller Manager to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
NOTE. Some metrics may not be collected depending on your Kubernetes Controller manager instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.CONTROLLER.SERVER.URL} | Kubernetes Controller manager metrics endpoint URL. |
https://localhost:10257/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token |
|
| {$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller: Get Controller metrics | Get raw metrics from Controller instance /metrics endpoint. |
HTTP agent | kubernetes.controller.get_metrics Preprocessing
|
| Leader election status | Gauge of if the reporting system is master of the relevant lease, 0 indicates backup, 1 indicates master. |
Dependent item | kubernetes.controller.leader_election_master_status Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.controller.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.controller.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.controller.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.controller.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.controller.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.controller.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.controller.max_fds Preprocessing
|
| REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_200.rate Preprocessing
|
| REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_300.rate Preprocessing
|
| REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_400.rate Preprocessing
|
| REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_500.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Controller manager: Too many HTTP client errors | "Kubernetes Controller manager is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Controller manager by HTTP/kubernetes.controller.client_http_requests_500.rate,5m)>{$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Dependent item | kubernetes.controller.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#NAME}"]: Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_adds_total["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.controller.workqueue_depth["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue unfinished work, sec | How many seconds of work has done that is in progress and hasn't been observed by work_duration. Large values indicate stuck threads. One can deduce the number of stuck threads by observing the rate at which this increases. |
Dependent item | kubernetes.controller.workqueue_unfinished_work_seconds["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue retries, rate | Total number of retries handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_retries_total["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue longest running processor, sec | How many seconds has the longest running processor for workqueue been running. |
Dependent item | kubernetes.controller.workqueue_longest_running_processor_seconds["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue work duration, p90 | 90 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p90["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, p95 | 95 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p95["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, p99 | 99 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p99["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, 50p | 50 percentiles of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p50["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p90 | 90 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p90["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p95 | 95 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p95["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p99 | 99 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p99["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, 50p | 50 percentile of how long in seconds an item stays in workqueue before being requested. If there are no requests for 5 minute, item value will be discarded. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p50["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue duration seconds bucket, {#LE} | How long in seconds processing an item from workqueue takes. |
Dependent item | kubernetes.controller.duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Queue duration seconds bucket, {#LE} | How long in seconds an item stays in workqueue before being requested. |
Dependent item | kubernetes.controller.queue_duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP - collects metrics by HTTP agent from Controller manager /metrics endpoint.
Zabbix version: 7.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.CONTROLLER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: You might need to set the --binding-address option for Controller Manager to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
Note: Some metrics may not be collected depending on your Kubernetes Controller manager instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.CONTROLLER.SERVER.URL} | Kubernetes Controller manager metrics endpoint URL. |
https://localhost:10257/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token |
|
| {$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller: Get Controller metrics | Get raw metrics from Controller instance /metrics endpoint. |
HTTP agent | kubernetes.controller.get_metrics Preprocessing
|
| Leader election status | Gauge of if the reporting system is master of the relevant lease, 0 indicates backup, 1 indicates master. |
Dependent item | kubernetes.controller.leader_election_master_status Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.controller.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.controller.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.controller.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.controller.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.controller.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.controller.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.controller.max_fds Preprocessing
|
| REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_200.rate Preprocessing
|
| REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_300.rate Preprocessing
|
| REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_400.rate Preprocessing
|
| REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_500.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Controller manager: Too many HTTP client errors | "Kubernetes Controller manager is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Controller manager by HTTP/kubernetes.controller.client_http_requests_500.rate,5m)>{$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Dependent item | kubernetes.controller.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#NAME}"]: Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_adds_total["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.controller.workqueue_depth["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue unfinished work, sec | How many seconds of work has done that is in progress and hasn't been observed by work_duration. Large values indicate stuck threads. One can deduce the number of stuck threads by observing the rate at which this increases. |
Dependent item | kubernetes.controller.workqueue_unfinished_work_seconds["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue retries, rate | Total number of retries handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_retries_total["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue longest running processor, sec | How many seconds has the longest running processor for workqueue been running. |
Dependent item | kubernetes.controller.workqueue_longest_running_processor_seconds["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue work duration, p90 | 90 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p90["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, p95 | 95 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p95["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, p99 | 99 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p99["{#NAME}"] |
| ["{#NAME}"]: Workqueue work duration, 50p | 50 percentiles of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p50["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p90 | 90 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p90["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p95 | 95 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p95["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, p99 | 99 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p99["{#NAME}"] |
| ["{#NAME}"]: Workqueue queue duration, 50p | 50 percentile of how long in seconds an item stays in workqueue before being requested. If there are no requests for 5 minute, item value will be discarded. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p50["{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Workqueue duration seconds bucket, {#LE} | How long in seconds processing an item from workqueue takes. |
Dependent item | kubernetes.controller.duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
| ["{#NAME}"]: Queue duration seconds bucket, {#LE} | How long in seconds an item stays in workqueue before being requested. |
Dependent item | kubernetes.controller.queue_duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP - collects metrics by HTTP agent from Controller manager /metrics endpoint.
Zabbix version: 6.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.CONTROLLER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. You might need to set the --binding-address option for Controller Manager to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
NOTE. Some metrics may not be collected depending on your Kubernetes Controller manager instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.CONTROLLER.SERVER.URL} | Kubernetes Controller manager metrics endpoint URL. |
https://localhost:10257/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token |
|
| {$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller: Get Controller metrics | Get raw metrics from Controller instance /metrics endpoint. |
HTTP agent | kubernetes.controller.get_metrics Preprocessing
|
| Kubernetes Controller Manager: Leader election status | Gauge of if the reporting system is master of the relevant lease, 0 indicates backup, 1 indicates master. |
Dependent item | kubernetes.controller.leader_election_master_status Preprocessing
|
| Kubernetes Controller Manager: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.controller.process_virtual_memory_bytes Preprocessing
|
| Kubernetes Controller Manager: Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.controller.process_resident_memory_bytes Preprocessing
|
| Kubernetes Controller Manager: CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.controller.cpu.util Preprocessing
|
| Kubernetes Controller Manager: Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.controller.go_goroutines Preprocessing
|
| Kubernetes Controller Manager: Go threads | Number of OS threads created. |
Dependent item | kubernetes.controller.go_threads Preprocessing
|
| Kubernetes Controller Manager: Fds open | Number of open file descriptors. |
Dependent item | kubernetes.controller.open_fds Preprocessing
|
| Kubernetes Controller Manager: Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.controller.max_fds Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_200.rate Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_300.rate Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_400.rate Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_500.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Controller Manager: Too many HTTP client errors | "Kubernetes Controller manager is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Controller manager by HTTP/kubernetes.controller.client_http_requests_500.rate,5m)>{$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Dependent item | kubernetes.controller.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_adds_total["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.controller.workqueue_depth["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue unfinished work, sec | How many seconds of work has done that is in progress and hasn't been observed by work_duration. Large values indicate stuck threads. One can deduce the number of stuck threads by observing the rate at which this increases. |
Dependent item | kubernetes.controller.workqueue_unfinished_work_seconds["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue retries, rate | Total number of retries handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_retries_total["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue longest running processor, sec | How many seconds has the longest running processor for workqueue been running. |
Dependent item | kubernetes.controller.workqueue_longest_running_processor_seconds["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p90 | 90 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p90["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p95 | 95 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p95["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p99 | 99 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p99["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, 50p | 50 percentiles of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p50["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p90 | 90 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p90["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p95 | 95 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p95["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p99 | 99 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p99["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, 50p | 50 percentile of how long in seconds an item stays in workqueue before being requested. If there are no requests for 5 minute, item value will be discarded. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p50["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue duration seconds bucket, {#LE} | How long in seconds processing an item from workqueue takes. |
Dependent item | kubernetes.controller.duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Queue duration seconds bucket, {#LE} | How long in seconds an item stays in workqueue before being requested. |
Dependent item | kubernetes.controller.queue_duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
For Zabbix version: 6.2 and higher
The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts.
Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP — collects metrics by HTTP agent from Controller manager /metrics endpoint.
This template was tested on:
See Zabbix template operation for basic instructions.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.CONTROLLER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values. NOTE. Some metrics may not be collected depending on your Kubernetes Controller manager instance version and configuration.
No specific Zabbix configuration is required.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | API Authorization Token |
`` |
| {$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger |
2 |
| {$KUBE.CONTROLLER.SERVER.URL} | Instance URL |
http://localhost:10252/metrics |
There are no template links in this template.
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | DEPENDENT | kubernetes.controller.workqueue.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| Group | Name | Description | Type | Key and additional info |
|---|---|---|---|---|
| Kubernetes Controller | Kubernetes Controller Manager: Leader election status | Gauge of if the reporting system is master of the relevant lease, 0 indicates backup, 1 indicates master. |
DEPENDENT | kubernetes.controller.leader_election_master_status Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: Virtual memory, bytes | Virtual memory size in bytes. |
DEPENDENT | kubernetes.controller.process_virtual_memory_bytes Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: Resident memory, bytes | Resident memory size in bytes. |
DEPENDENT | kubernetes.controller.process_resident_memory_bytes Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: CPU | Total user and system CPU usage ratio. |
DEPENDENT | kubernetes.controller.cpu.util Preprocessing: - PROMETHEUS_PATTERN: - CHANGE_PER_SECOND - MULTIPLIER: |
| Kubernetes Controller | Kubernetes Controller Manager: Goroutines | Number of goroutines that currently exist. |
DEPENDENT | kubernetes.controller.go_goroutines Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: Go threads | Number of OS threads created. |
DEPENDENT | kubernetes.controller.go_threads Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: Fds open | Number of open file descriptors. |
DEPENDENT | kubernetes.controller.open_fds Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: Fds max | Maximum allowed open file descriptors. |
DEPENDENT | kubernetes.controller.max_fds Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
DEPENDENT | kubernetes.controller.client_http_requests_200.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Controller | Kubernetes Controller Manager: REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
DEPENDENT | kubernetes.controller.client_http_requests_300.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Controller | Kubernetes Controller Manager: REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
DEPENDENT | kubernetes.controller.client_http_requests_400.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Controller | Kubernetes Controller Manager: REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
DEPENDENT | kubernetes.controller.client_http_requests_500.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
DEPENDENT | kubernetes.controller.workqueue_adds_total["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue depth | Current depth of workqueue. |
DEPENDENT | kubernetes.controller.workqueue_depth["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue unfinished work, sec | How many seconds of work has done that is in progress and hasn't been observed by work_duration. Large values indicate stuck threads. One can deduce the number of stuck threads by observing the rate at which this increases. |
DEPENDENT | kubernetes.controller.workqueue_unfinished_work_seconds["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue retries, rate | Total number of retries handled by workqueue per second. |
DEPENDENT | kubernetes.controller.workqueue_retries_total["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue longest running processor, sec | How many seconds has the longest running processor for workqueue been running. |
DEPENDENT | kubernetes.controller.workqueue_longest_running_processor_seconds["{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p90 | 90 percentile of how long in seconds processing an item from workqueue takes, by queue. |
CALCULATED | kubernetes.controller.workqueue_work_duration_seconds_p90["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.duration_seconds_bucket[*,"{#NAME}"],5m,90) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p95 | 95 percentile of how long in seconds processing an item from workqueue takes, by queue. |
CALCULATED | kubernetes.controller.workqueue_work_duration_seconds_p95["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.duration_seconds_bucket[*,"{#NAME}"],5m,95) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p99 | 99 percentile of how long in seconds processing an item from workqueue takes, by queue. |
CALCULATED | kubernetes.controller.workqueue_work_duration_seconds_p99["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.duration_seconds_bucket[*,"{#NAME}"],5m,99) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, 50p | 50 percentiles of how long in seconds processing an item from workqueue takes, by queue. |
CALCULATED | kubernetes.controller.workqueue_work_duration_seconds_p50["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.duration_seconds_bucket[*,"{#NAME}"],5m,50) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p90 | 90 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
CALCULATED | kubernetes.controller.workqueue_queue_duration_seconds_p90["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.queue_duration_seconds_bucket[*,"{#NAME}"],5m,90) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p95 | 95 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
CALCULATED | kubernetes.controller.workqueue_queue_duration_seconds_p95["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.queue_duration_seconds_bucket[*,"{#NAME}"],5m,95) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p99 | 99 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
CALCULATED | kubernetes.controller.workqueue_queue_duration_seconds_p99["{#NAME}"] Expression: bucket_percentile(//kubernetes.controller.queue_duration_seconds_bucket[*,"{#NAME}"],5m,99) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, 50p | 50 percentile of how long in seconds an item stays in workqueue before being requested. If there are no requests for 5 minute, item value will be discarded. |
CALCULATED | kubernetes.controller.workqueue_queue_duration_seconds_p50["{#NAME}"] Preprocessing: - CHECK_NOT_SUPPORTED ⛔️ON_FAIL: Expression: bucket_percentile(//kubernetes.controller.queue_duration_seconds_bucket[*,"{#NAME}"],5m,50) |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Workqueue duration seconds bucket, {#LE} | How long in seconds processing an item from workqueue takes. |
DEPENDENT | kubernetes.controller.duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Controller | Kubernetes Controller Manager: ["{#NAME}"]: Queue duration seconds bucket, {#LE} | How long in seconds an item stays in workqueue before being requested. |
DEPENDENT | kubernetes.controller.queue_duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Zabbix raw items | Kubernetes Controller: Get Controller metrics | Get raw metrics from Controller instance /metrics endpoint. |
HTTP_AGENT | kubernetes.controller.get_metrics Preprocessing: - CHECK_NOT_SUPPORTED ⛔️ON_FAIL: |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Controller Manager: Too many HTTP client errors | "Kubernetes Controller manager is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Controller manager by HTTP/kubernetes.controller.client_http_requests_500.rate,5m)>{$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} |
WARNING |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template or ask for help with it at ZABBIX forums.
The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP - collects metrics by HTTP agent from Controller manager /metrics endpoint.
Zabbix version: 6.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.CONTROLLER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. You might need to set the --binding-address option for Controller Manager to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
NOTE. Some metrics may not be collected depending on your Kubernetes Controller manager instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.CONTROLLER.SERVER.URL} | Kubernetes Controller manager metrics endpoint URL. |
https://localhost:10257/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token |
|
| {$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller: Get Controller metrics | Get raw metrics from Controller instance /metrics endpoint. |
HTTP agent | kubernetes.controller.get_metrics Preprocessing
|
| Kubernetes Controller Manager: Leader election status | Gauge of if the reporting system is master of the relevant lease, 0 indicates backup, 1 indicates master. |
Dependent item | kubernetes.controller.leader_election_master_status Preprocessing
|
| Kubernetes Controller Manager: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.controller.process_virtual_memory_bytes Preprocessing
|
| Kubernetes Controller Manager: Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.controller.process_resident_memory_bytes Preprocessing
|
| Kubernetes Controller Manager: CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.controller.cpu.util Preprocessing
|
| Kubernetes Controller Manager: Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.controller.go_goroutines Preprocessing
|
| Kubernetes Controller Manager: Go threads | Number of OS threads created. |
Dependent item | kubernetes.controller.go_threads Preprocessing
|
| Kubernetes Controller Manager: Fds open | Number of open file descriptors. |
Dependent item | kubernetes.controller.open_fds Preprocessing
|
| Kubernetes Controller Manager: Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.controller.max_fds Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_200.rate Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_300.rate Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_400.rate Preprocessing
|
| Kubernetes Controller Manager: REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.controller.client_http_requests_500.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Controller Manager: Too many HTTP client errors | "Kubernetes Controller manager is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Controller manager by HTTP/kubernetes.controller.client_http_requests_500.rate,5m)>{$KUBE.CONTROLLER.HTTP.CLIENT.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Workqueue metrics discovery | Dependent item | kubernetes.controller.workqueue.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue adds total, rate | Total number of adds handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_adds_total["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue depth | Current depth of workqueue. |
Dependent item | kubernetes.controller.workqueue_depth["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue unfinished work, sec | How many seconds of work has done that is in progress and hasn't been observed by work_duration. Large values indicate stuck threads. One can deduce the number of stuck threads by observing the rate at which this increases. |
Dependent item | kubernetes.controller.workqueue_unfinished_work_seconds["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue retries, rate | Total number of retries handled by workqueue per second. |
Dependent item | kubernetes.controller.workqueue_retries_total["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue longest running processor, sec | How many seconds has the longest running processor for workqueue been running. |
Dependent item | kubernetes.controller.workqueue_longest_running_processor_seconds["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p90 | 90 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p90["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p95 | 95 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p95["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, p99 | 99 percentile of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p99["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue work duration, 50p | 50 percentiles of how long in seconds processing an item from workqueue takes, by queue. |
Calculated | kubernetes.controller.workqueue_work_duration_seconds_p50["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p90 | 90 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p90["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p95 | 95 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p95["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, p99 | 99 percentile of how long in seconds an item stays in workqueue before being requested, by queue. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p99["{#NAME}"] |
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue queue duration, 50p | 50 percentile of how long in seconds an item stays in workqueue before being requested. If there are no requests for 5 minute, item value will be discarded. |
Calculated | kubernetes.controller.workqueue_queue_duration_seconds_p50["{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Workqueue duration seconds bucket, {#LE} | How long in seconds processing an item from workqueue takes. |
Dependent item | kubernetes.controller.duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
| Kubernetes Controller Manager: ["{#NAME}"]: Queue duration seconds bucket, {#LE} | How long in seconds an item stays in workqueue before being requested. |
Dependent item | kubernetes.controller.queue_duration_seconds_bucket[{#LE},"{#NAME}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Kubelet by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Kubelet by HTTP - collects metrics by HTTP agent from Kubelet /metrics endpoint.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
Note: Some metrics may not be collected depending on your Kubernetes instance version and configuration.
Zabbix version: 7.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
Note: Some metrics may not be collected depending on your Kubernetes instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.KUBELET.URL} | Kubernetes Kubelet instance URL. |
https://localhost:10250 |
| {$KUBE.KUBELET.METRIC.ENDPOINT} | Kubelet /metrics endpoint. |
/metrics |
| {$KUBE.KUBELET.CADVISOR.ENDPOINT} | cAdvisor metrics from Kubelet /metrics/cadvisor endpoint. |
/metrics/cadvisor |
| {$KUBE.KUBELET.PODS.ENDPOINT} | Kubelet /pods endpoint. |
/pods |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get kubelet metrics | Collecting raw Kubelet metrics from /metrics endpoint. |
HTTP agent | kube.kubelet.metrics |
| Get cadvisor metrics | Collecting raw Kubelet metrics from /metrics/cadvisor endpoint. |
HTTP agent | kube.cadvisor.metrics |
| Get pods | Collecting raw Kubelet metrics from /pods endpoint. |
HTTP agent | kube.pods |
| Pods running | The number of running pods. |
Dependent item | kube.kubelet.pods.running Preprocessing
|
| Containers started | The number of started containers. |
Dependent item | kube.kubelet.containers.started Preprocessing
|
| Containers ready | The number of ready containers. |
Dependent item | kube.kubelet.containers.ready Preprocessing
|
| Containers last state terminated | The number of containers that were previously terminated. |
Dependent item | kube.kublet.containers.terminated Preprocessing
|
| Containers restarts | The number of times the container has been restarted. |
Dependent item | kube.kubelet.containers.restarts Preprocessing
|
| CPU cores, total | The number of cores in this machine (available until kubernetes v1.18). |
Dependent item | kube.kubelet.cpu.cores Preprocessing
|
| Machine memory, bytes | Resident memory size in bytes. |
Dependent item | kube.kubelet.machine.memory Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kube.kubelet.virtual.memory Preprocessing
|
| File descriptors, max | Maximum number of open file descriptors. |
Dependent item | kube.kubelet.process_max_fds Preprocessing
|
| File descriptors, open | Number of open file descriptors. |
Dependent item | kube.kubelet.process_open_fds Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Runtime operations discovery | Dependent item | kube.kubelet.runtime_operations_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| [{#OP_TYPE}] Runtime operations bucket: {#LE} | Duration in seconds of runtime operations. Broken down by operation type. |
Dependent item | kube.kublet.runtime_ops_duration_seconds_bucket[{#LE},"{#OP_TYPE}"] Preprocessing
|
| [{#OP_TYPE}] Runtime operations total, rate | Cumulative number of runtime operations by operation type. |
Dependent item | kube.kublet.runtime_ops_total.rate["{#OP_TYPE}"] Preprocessing
|
| [{#OP_TYPE}] Operations, p90 | 90 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p90["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p95 | 95 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p95["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p99 | 99 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p99["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p50 | 50 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p50["{#OP_TYPE}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pods discovery | Dependent item | kube.kubelet.pods.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Load average, 10s | Pods cpu load average over the last 10 seconds. |
Dependent item | kube.pod.container_cpu_load_average_10s[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: System seconds, total | System cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_system_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Usage seconds, total | Consumed cpu time. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_usage_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: User seconds, total | User cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_user_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| REST client requests discovery | Dependent item | kube.kubelet.rest.requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Host [{#HOST}] Request method [{#METHOD}] Code:[{#CODE}] | Number of HTTP requests, partitioned by status code, method, and host. |
Dependent item | kube.kubelet.rest.requests["{#CODE}", "{#HOST}", "{#METHOD}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Container memory discovery | Dependent item | kube.kubelet.container.memory.cache.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory page cache | Number of bytes of page cache memory. |
Dependent item | kube.kubelet.container.memory.cache["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory max usage | Maximum memory usage recorded in bytes. |
Dependent item | kube.kubelet.container.memory.max_usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: RSS | Size of RSS in bytes. |
Dependent item | kube.kubelet.container.memory.rss["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Swap | Container swap usage in bytes. |
Dependent item | kube.kubelet.container.memory.swap["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Usage | Current memory usage in bytes, including all memory regardless of when it was accessed. |
Dependent item | kube.kubelet.container.memory.usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Working set | Current working set in bytes. |
Dependent item | kube.kubelet.container.memory.working_set["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Kubelet by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Kubelet by HTTP - collects metrics by HTTP agent from Kubelet /metrics endpoint.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
Zabbix version: 7.2 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.KUBELET.URL} | Kubernetes Kubelet instance URL. |
https://localhost:10250 |
| {$KUBE.KUBELET.METRIC.ENDPOINT} | Kubelet /metrics endpoint. |
/metrics |
| {$KUBE.KUBELET.CADVISOR.ENDPOINT} | cAdvisor metrics from Kubelet /metrics/cadvisor endpoint. |
/metrics/cadvisor |
| {$KUBE.KUBELET.PODS.ENDPOINT} | Kubelet /pods endpoint. |
/pods |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get kubelet metrics | Collecting raw Kubelet metrics from /metrics endpoint. |
HTTP agent | kube.kubelet.metrics |
| Get cadvisor metrics | Collecting raw Kubelet metrics from /metrics/cadvisor endpoint. |
HTTP agent | kube.cadvisor.metrics |
| Get pods | Collecting raw Kubelet metrics from /pods endpoint. |
HTTP agent | kube.pods |
| Pods running | The number of running pods. |
Dependent item | kube.kubelet.pods.running Preprocessing
|
| Containers started | The number of started containers. |
Dependent item | kube.kubelet.containers.started Preprocessing
|
| Containers ready | The number of ready containers. |
Dependent item | kube.kubelet.containers.ready Preprocessing
|
| Containers last state terminated | The number of containers that were previously terminated. |
Dependent item | kube.kublet.containers.terminated Preprocessing
|
| Containers restarts | The number of times the container has been restarted. |
Dependent item | kube.kubelet.containers.restarts Preprocessing
|
| CPU cores, total | The number of cores in this machine (available until kubernetes v1.18). |
Dependent item | kube.kubelet.cpu.cores Preprocessing
|
| Machine memory, bytes | Resident memory size in bytes. |
Dependent item | kube.kubelet.machine.memory Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kube.kubelet.virtual.memory Preprocessing
|
| File descriptors, max | Maximum number of open file descriptors. |
Dependent item | kube.kubelet.process_max_fds Preprocessing
|
| File descriptors, open | Number of open file descriptors. |
Dependent item | kube.kubelet.process_open_fds Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Runtime operations discovery | Dependent item | kube.kubelet.runtime_operations_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| [{#OP_TYPE}] Runtime operations bucket: {#LE} | Duration in seconds of runtime operations. Broken down by operation type. |
Dependent item | kube.kublet.runtime_ops_duration_seconds_bucket[{#LE},"{#OP_TYPE}"] Preprocessing
|
| [{#OP_TYPE}] Runtime operations total, rate | Cumulative number of runtime operations by operation type. |
Dependent item | kube.kublet.runtime_ops_total.rate["{#OP_TYPE}"] Preprocessing
|
| [{#OP_TYPE}] Operations, p90 | 90 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p90["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p95 | 95 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p95["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p99 | 99 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p99["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p50 | 50 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p50["{#OP_TYPE}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pods discovery | Dependent item | kube.kubelet.pods.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Load average, 10s | Pods cpu load average over the last 10 seconds. |
Dependent item | kube.pod.container_cpu_load_average_10s[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: System seconds, total | System cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_system_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Usage seconds, total | Consumed cpu time. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_usage_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: User seconds, total | User cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_user_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| REST client requests discovery | Dependent item | kube.kubelet.rest.requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Host [{#HOST}] Request method [{#METHOD}] Code:[{#CODE}] | Number of HTTP requests, partitioned by status code, method, and host. |
Dependent item | kube.kubelet.rest.requests["{#CODE}", "{#HOST}", "{#METHOD}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Container memory discovery | Dependent item | kube.kubelet.container.memory.cache.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory page cache | Number of bytes of page cache memory. |
Dependent item | kube.kubelet.container.memory.cache["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory max usage | Maximum memory usage recorded in bytes. |
Dependent item | kube.kubelet.container.memory.max_usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: RSS | Size of RSS in bytes. |
Dependent item | kube.kubelet.container.memory.rss["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Swap | Container swap usage in bytes. |
Dependent item | kube.kubelet.container.memory.swap["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Usage | Current memory usage in bytes, including all memory regardless of when it was accessed. |
Dependent item | kube.kubelet.container.memory.usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Working set | Current working set in bytes. |
Dependent item | kube.kubelet.container.memory.working_set["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Kubelet by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Kubelet by HTTP - collects metrics by HTTP agent from Kubelet /metrics endpoint.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
Note: Some metrics may not be collected depending on your Kubernetes instance version and configuration.
Zabbix version: 7.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
Note: Some metrics may not be collected depending on your Kubernetes instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.KUBELET.URL} | Kubernetes Kubelet instance URL. |
https://localhost:10250 |
| {$KUBE.KUBELET.METRIC.ENDPOINT} | Kubelet /metrics endpoint. |
/metrics |
| {$KUBE.KUBELET.CADVISOR.ENDPOINT} | cAdvisor metrics from Kubelet /metrics/cadvisor endpoint. |
/metrics/cadvisor |
| {$KUBE.KUBELET.PODS.ENDPOINT} | Kubelet /pods endpoint. |
/pods |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get kubelet metrics | Collecting raw Kubelet metrics from /metrics endpoint. |
HTTP agent | kube.kubelet.metrics |
| Get cadvisor metrics | Collecting raw Kubelet metrics from /metrics/cadvisor endpoint. |
HTTP agent | kube.cadvisor.metrics |
| Get pods | Collecting raw Kubelet metrics from /pods endpoint. |
HTTP agent | kube.pods |
| Pods running | The number of running pods. |
Dependent item | kube.kubelet.pods.running Preprocessing
|
| Containers started | The number of started containers. |
Dependent item | kube.kubelet.containers.started Preprocessing
|
| Containers ready | The number of ready containers. |
Dependent item | kube.kubelet.containers.ready Preprocessing
|
| Containers last state terminated | The number of containers that were previously terminated. |
Dependent item | kube.kublet.containers.terminated Preprocessing
|
| Containers restarts | The number of times the container has been restarted. |
Dependent item | kube.kubelet.containers.restarts Preprocessing
|
| CPU cores, total | The number of cores in this machine (available until kubernetes v1.18). |
Dependent item | kube.kubelet.cpu.cores Preprocessing
|
| Machine memory, bytes | Resident memory size in bytes. |
Dependent item | kube.kubelet.machine.memory Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kube.kubelet.virtual.memory Preprocessing
|
| File descriptors, max | Maximum number of open file descriptors. |
Dependent item | kube.kubelet.process_max_fds Preprocessing
|
| File descriptors, open | Number of open file descriptors. |
Dependent item | kube.kubelet.process_open_fds Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Runtime operations discovery | Dependent item | kube.kubelet.runtime_operations_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| [{#OP_TYPE}] Runtime operations bucket: {#LE} | Duration in seconds of runtime operations. Broken down by operation type. |
Dependent item | kube.kublet.runtime_ops_duration_seconds_bucket[{#LE},"{#OP_TYPE}"] Preprocessing
|
| [{#OP_TYPE}] Runtime operations total, rate | Cumulative number of runtime operations by operation type. |
Dependent item | kube.kublet.runtime_ops_total.rate["{#OP_TYPE}"] Preprocessing
|
| [{#OP_TYPE}] Operations, p90 | 90 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p90["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p95 | 95 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p95["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p99 | 99 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p99["{#OP_TYPE}"] |
| [{#OP_TYPE}] Operations, p50 | 50 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p50["{#OP_TYPE}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pods discovery | Dependent item | kube.kubelet.pods.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Load average, 10s | Pods cpu load average over the last 10 seconds. |
Dependent item | kube.pod.container_cpu_load_average_10s[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: System seconds, total | System cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_system_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Usage seconds, total | Consumed cpu time. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_usage_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: User seconds, total | User cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_user_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| REST client requests discovery | Dependent item | kube.kubelet.rest.requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Host [{#HOST}] Request method [{#METHOD}] Code:[{#CODE}] | Number of HTTP requests, partitioned by status code, method, and host. |
Dependent item | kube.kubelet.rest.requests["{#CODE}", "{#HOST}", "{#METHOD}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Container memory discovery | Dependent item | kube.kubelet.container.memory.cache.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory page cache | Number of bytes of page cache memory. |
Dependent item | kube.kubelet.container.memory.cache["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory max usage | Maximum memory usage recorded in bytes. |
Dependent item | kube.kubelet.container.memory.max_usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: RSS | Size of RSS in bytes. |
Dependent item | kube.kubelet.container.memory.rss["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Swap | Container swap usage in bytes. |
Dependent item | kube.kubelet.container.memory.swap["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Usage | Current memory usage in bytes, including all memory regardless of when it was accessed. |
Dependent item | kube.kubelet.container.memory.usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Working set | Current working set in bytes. |
Dependent item | kube.kubelet.container.memory.working_set["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Kubelet by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Kubelet by HTTP - collects metrics by HTTP agent from Kubelet /metrics endpoint.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
Zabbix version: 6.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.KUBELET.URL} | Kubernetes Kubelet instance URL. |
https://localhost:10250 |
| {$KUBE.KUBELET.METRIC.ENDPOINT} | Kubelet /metrics endpoint. |
/metrics |
| {$KUBE.KUBELET.CADVISOR.ENDPOINT} | cAdvisor metrics from Kubelet /metrics/cadvisor endpoint. |
/metrics/cadvisor |
| {$KUBE.KUBELET.PODS.ENDPOINT} | Kubelet /pods endpoint. |
/pods |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Get kubelet metrics | Collecting raw Kubelet metrics from /metrics endpoint. |
HTTP agent | kube.kubelet.metrics |
| Kubernetes: Get cadvisor metrics | Collecting raw Kubelet metrics from /metrics/cadvisor endpoint. |
HTTP agent | kube.cadvisor.metrics |
| Kubernetes: Get pods | Collecting raw Kubelet metrics from /pods endpoint. |
HTTP agent | kube.pods |
| Kubernetes: Pods running | The number of running pods. |
Dependent item | kube.kubelet.pods.running Preprocessing
|
| Kubernetes: Containers running | The number of running containers. |
Dependent item | kube.kubelet.containers.running Preprocessing
|
| Kubernetes: Containers last state terminated | The number of containers that were previously terminated. |
Dependent item | kube.kublet.containers.terminated Preprocessing
|
| Kubernetes: Containers restarts | The number of times the container has been restarted. |
Dependent item | kube.kubelet.containers.restarts Preprocessing
|
| Kubernetes: CPU cores, total | The number of cores in this machine (available until kubernetes v1.18). |
Dependent item | kube.kubelet.cpu.cores Preprocessing
|
| Kubernetes: Machine memory, bytes | Resident memory size in bytes. |
Dependent item | kube.kubelet.machine.memory Preprocessing
|
| Kubernetes: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kube.kubelet.virtual.memory Preprocessing
|
| Kubernetes: File descriptors, max | Maximum number of open file descriptors. |
Dependent item | kube.kubelet.process_max_fds Preprocessing
|
| Kubernetes: File descriptors, open | Number of open file descriptors. |
Dependent item | kube.kubelet.process_open_fds Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Runtime operations discovery | Dependent item | kube.kubelet.runtime_operations_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: [{#OP_TYPE}] Runtime operations bucket: {#LE} | Duration in seconds of runtime operations. Broken down by operation type. |
Dependent item | kube.kublet.runtime_ops_duration_seconds_bucket[{#LE},"{#OP_TYPE}"] Preprocessing
|
| Kubernetes: [{#OP_TYPE}] Runtime operations total, rate | Cumulative number of runtime operations by operation type. |
Dependent item | kube.kublet.runtime_ops_total.rate["{#OP_TYPE}"] Preprocessing
|
| Kubernetes: [{#OP_TYPE}] Operations, p90 | 90 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p90["{#OP_TYPE}"] |
| Kubernetes: [{#OP_TYPE}] Operations, p95 | 95 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p95["{#OP_TYPE}"] |
| Kubernetes: [{#OP_TYPE}] Operations, p99 | 99 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p99["{#OP_TYPE}"] |
| Kubernetes: [{#OP_TYPE}] Operations, p50 | 50 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p50["{#OP_TYPE}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pods discovery | Dependent item | kube.kubelet.pods.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Load average, 10s | Pods cpu load average over the last 10 seconds. |
Dependent item | kube.pod.container_cpu_load_average_10s[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: System seconds, total | System cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_system_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Usage seconds, total | Consumed cpu time. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_usage_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: User seconds, total | User cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_user_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| REST client requests discovery | Dependent item | kube.kubelet.rest.requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Host [{#HOST}] Request method [{#METHOD}] Code:[{#CODE}] | Number of HTTP requests, partitioned by status code, method, and host. |
Dependent item | kube.kubelet.rest.requests["{#CODE}", "{#HOST}", "{#METHOD}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Container memory discovery | Dependent item | kube.kubelet.container.memory.cache.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory page cache | Number of bytes of page cache memory. |
Dependent item | kube.kubelet.container.memory.cache["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory max usage | Maximum memory usage recorded in bytes. |
Dependent item | kube.kubelet.container.memory.max_usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: RSS | Size of RSS in bytes. |
Dependent item | kube.kubelet.container.memory.rss["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Swap | Container swap usage in bytes. |
Dependent item | kube.kubelet.container.memory.swap["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Usage | Current memory usage in bytes, including all memory regardless of when it was accessed. |
Dependent item | kube.kubelet.container.memory.usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Working set | Current working set in bytes. |
Dependent item | kube.kubelet.container.memory.working_set["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
For Zabbix version: 6.2 and higher. The template to monitor Kubernetes Controller manager by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Controller manager by HTTP — collects metrics by HTTP agent from Controller manager /metrics endpoint.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}. NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
This template was tested on:
See Zabbix template operation for basic instructions.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}. NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
No specific Zabbix configuration is required.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token |
`` |
| {$KUBE.KUBELET.CADVISOR.ENDPOINT} | cAdvisor metrics from Kubelet /metrics/cadvisor endpoint |
/metrics/cadvisor |
| {$KUBE.KUBELET.METRIC.ENDPOINT} | Kubelet /metrics endpoint |
/metrics |
| {$KUBE.KUBELET.PODS.ENDPOINT} | Kubelet /pods endpoint |
/pods |
| {$KUBE.KUBELET.URL} | Instance URL |
https://localhost:10250 |
There are no template links in this template.
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Container memory discovery | DEPENDENT | kube.kubelet.container.memory.cache.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
|
| Pods discovery | DEPENDENT | kube.kubelet.pods.discovery Preprocessing: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
|
| REST client requests discovery | DEPENDENT | kube.kubelet.rest.requests.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: |
|
| Runtime operations discovery | DEPENDENT | kube.kubelet.runtime_operations_bucket.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| Group | Name | Description | Type | Key and additional info |
|---|---|---|---|---|
| Kubernetes | Kubernetes: Get kubelet metrics | Collecting raw Kubelet metrics from /metrics endpoint. |
HTTP_AGENT | kube.kubelet.metrics |
| Kubernetes | Kubernetes: Get cadvisor metrics | Collecting raw Kubelet metrics from /metrics/cadvisor endpoint. |
HTTP_AGENT | kube.cadvisor.metrics |
| Kubernetes | Kubernetes: Get pods | Collecting raw Kubelet metrics from /pods endpoint. |
HTTP_AGENT | kube.pods |
| Kubernetes | Kubernetes: Pods running | The number of running pods. |
DEPENDENT | kube.kubelet.pods.running Preprocessing: - JSONPATH: |
| Kubernetes | Kubernetes: Containers running | The number of running containers. |
DEPENDENT | kube.kubelet.containers.running Preprocessing: - JSONPATH: |
| Kubernetes | Kubernetes: Containers last state terminated | The number of containers that were previously terminated. |
DEPENDENT | kube.kublet.containers.terminated Preprocessing: - JSONPATH: |
| Kubernetes | Kubernetes: Containers restarts | The number of times the container has been restarted. |
DEPENDENT | kube.kubelet.containers.restarts Preprocessing: - JSONPATH: |
| Kubernetes | Kubernetes: CPU cores, total | The number of cores in this machine (available until kubernetes v1.18). |
DEPENDENT | kube.kubelet.cpu.cores Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes | Kubernetes: Machine memory, bytes | Resident memory size in bytes. |
DEPENDENT | kube.kubelet.machine.memory Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes | Kubernetes: Virtual memory, bytes | Virtual memory size in bytes. |
DEPENDENT | kube.kubelet.virtual.memory Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes | Kubernetes: File descriptors, max | Maximum number of open file descriptors. |
DEPENDENT | kube.kubelet.process_max_fds Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes | Kubernetes: File descriptors, open | Number of open file descriptors. |
DEPENDENT | kube.kubelet.process_open_fds Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes | Kubernetes: [{#OP_TYPE}] Runtime operations bucket: {#LE} | Duration in seconds of runtime operations. Broken down by operation type. |
DEPENDENT | kube.kublet.runtime_ops_duration_seconds_bucket[{#LE},"{#OP_TYPE}"] Preprocessing: - PROMETHEUS_PATTERN: |
| Kubernetes | Kubernetes: [{#OP_TYPE}] Runtime operations total, rate | Cumulative number of runtime operations by operation type. |
DEPENDENT | kube.kublet.runtime_ops_total.rate["{#OP_TYPE}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes | Kubernetes: [{#OP_TYPE}] Operations, p90 | 90 percentile of operation latency distribution in seconds for each verb. |
CALCULATED | kube.kublet.runtime_ops_duration_seconds_p90["{#OP_TYPE}"] Expression: bucket_percentile(//kube.kublet.runtime_ops_duration_seconds_bucket[*,"{#OP_TYPE}"],5m,90) |
| Kubernetes | Kubernetes: [{#OP_TYPE}] Operations, p95 | 95 percentile of operation latency distribution in seconds for each verb. |
CALCULATED | kube.kublet.runtime_ops_duration_seconds_p95["{#OP_TYPE}"] Expression: bucket_percentile(//kube.kublet.runtime_ops_duration_seconds_bucket[*,"{#OP_TYPE}"],5m,95) |
| Kubernetes | Kubernetes: [{#OP_TYPE}] Operations, p99 | 99 percentile of operation latency distribution in seconds for each verb. |
CALCULATED | kube.kublet.runtime_ops_duration_seconds_p99["{#OP_TYPE}"] Expression: bucket_percentile(//kube.kublet.runtime_ops_duration_seconds_bucket[*,"{#OP_TYPE}"],5m,99) |
| Kubernetes | Kubernetes: [{#OP_TYPE}] Operations, p50 | 50 percentile of operation latency distribution in seconds for each verb. |
CALCULATED | kube.kublet.runtime_ops_duration_seconds_p50["{#OP_TYPE}"] Expression: bucket_percentile(//kube.kublet.runtime_ops_duration_seconds_bucket[*,"{#OP_TYPE}"],5m,50) |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Load average, 10s | Pods cpu load average over the last 10 seconds. |
DEPENDENT | kube.pod.container_cpu_load_average_10s[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: System seconds, total | The number of cores used for system time. |
DEPENDENT | kube.pod.container_cpu_system_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: User seconds, total | The number of cores used for user time. |
DEPENDENT | kube.pod.container_cpu_user_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Host [{#HOST}] Request method [{#METHOD}] Code:[{#CODE}] | Number of HTTP requests, partitioned by status code, method, and host. |
DEPENDENT | kube.kubelet.rest.requests["{#CODE}", "{#HOST}", "{#METHOD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory page cache | Number of bytes of page cache memory. |
DEPENDENT | kube.kubelet.container.memory.cache["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory max usage | Maximum memory usage recorded in bytes. |
DEPENDENT | kube.kubelet.container.memory.max_usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: RSS | Size of RSS in bytes. |
DEPENDENT | kube.kubelet.container.memory.rss["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Swap | Container swap usage in bytes. |
DEPENDENT | kube.kubelet.container.memory.swap["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Usage | Current memory usage in bytes, including all memory regardless of when it was accessed. |
DEPENDENT | kube.kubelet.container.memory.usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Working set | Current working set in bytes. |
DEPENDENT | kube.kubelet.container.memory.working_set["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing: - PROMETHEUS_PATTERN: - DISCARD_UNCHANGED_HEARTBEAT: |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|
Please report any issues with the template at https://support.zabbix.com.
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums.
The template to monitor Kubernetes Kubelet by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Kubelet by HTTP - collects metrics by HTTP agent from Kubelet /metrics endpoint.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
Zabbix version: 6.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.KUBELET.URL}, {$KUBE.API.TOKEN}.
NOTE. Some metrics may not be collected depending on your Kubernetes instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.KUBELET.URL} | Kubernetes Kubelet instance URL. |
https://localhost:10250 |
| {$KUBE.KUBELET.METRIC.ENDPOINT} | Kubelet /metrics endpoint. |
/metrics |
| {$KUBE.KUBELET.CADVISOR.ENDPOINT} | cAdvisor metrics from Kubelet /metrics/cadvisor endpoint. |
/metrics/cadvisor |
| {$KUBE.KUBELET.PODS.ENDPOINT} | Kubelet /pods endpoint. |
/pods |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Get kubelet metrics | Collecting raw Kubelet metrics from /metrics endpoint. |
HTTP agent | kube.kubelet.metrics |
| Kubernetes: Get cadvisor metrics | Collecting raw Kubelet metrics from /metrics/cadvisor endpoint. |
HTTP agent | kube.cadvisor.metrics |
| Kubernetes: Get pods | Collecting raw Kubelet metrics from /pods endpoint. |
HTTP agent | kube.pods |
| Kubernetes: Pods running | The number of running pods. |
Dependent item | kube.kubelet.pods.running Preprocessing
|
| Kubernetes: Containers started | The number of started containers. |
Dependent item | kube.kubelet.containers.started Preprocessing
|
| Kubernetes: Containers ready | The number of ready containers. |
Dependent item | kube.kubelet.containers.ready Preprocessing
|
| Kubernetes: Containers last state terminated | The number of containers that were previously terminated. |
Dependent item | kube.kublet.containers.terminated Preprocessing
|
| Kubernetes: Containers restarts | The number of times the container has been restarted. |
Dependent item | kube.kubelet.containers.restarts Preprocessing
|
| Kubernetes: CPU cores, total | The number of cores in this machine (available until kubernetes v1.18). |
Dependent item | kube.kubelet.cpu.cores Preprocessing
|
| Kubernetes: Machine memory, bytes | Resident memory size in bytes. |
Dependent item | kube.kubelet.machine.memory Preprocessing
|
| Kubernetes: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kube.kubelet.virtual.memory Preprocessing
|
| Kubernetes: File descriptors, max | Maximum number of open file descriptors. |
Dependent item | kube.kubelet.process_max_fds Preprocessing
|
| Kubernetes: File descriptors, open | Number of open file descriptors. |
Dependent item | kube.kubelet.process_open_fds Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Runtime operations discovery | Dependent item | kube.kubelet.runtime_operations_bucket.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: [{#OP_TYPE}] Runtime operations bucket: {#LE} | Duration in seconds of runtime operations. Broken down by operation type. |
Dependent item | kube.kublet.runtime_ops_duration_seconds_bucket[{#LE},"{#OP_TYPE}"] Preprocessing
|
| Kubernetes: [{#OP_TYPE}] Runtime operations total, rate | Cumulative number of runtime operations by operation type. |
Dependent item | kube.kublet.runtime_ops_total.rate["{#OP_TYPE}"] Preprocessing
|
| Kubernetes: [{#OP_TYPE}] Operations, p90 | 90 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p90["{#OP_TYPE}"] |
| Kubernetes: [{#OP_TYPE}] Operations, p95 | 95 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p95["{#OP_TYPE}"] |
| Kubernetes: [{#OP_TYPE}] Operations, p99 | 99 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p99["{#OP_TYPE}"] |
| Kubernetes: [{#OP_TYPE}] Operations, p50 | 50 percentile of operation latency distribution in seconds for each verb. |
Calculated | kube.kublet.runtime_ops_duration_seconds_p50["{#OP_TYPE}"] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pods discovery | Dependent item | kube.kubelet.pods.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Load average, 10s | Pods cpu load average over the last 10 seconds. |
Dependent item | kube.pod.container_cpu_load_average_10s[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: System seconds, total | System cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_system_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: Usage seconds, total | Consumed cpu time. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_usage_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] CPU: User seconds, total | User cpu time consumed. It is calculated from the cumulative value using the |
Dependent item | kube.pod.container_cpu_user_seconds_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| REST client requests discovery | Dependent item | kube.kubelet.rest.requests.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Host [{#HOST}] Request method [{#METHOD}] Code:[{#CODE}] | Number of HTTP requests, partitioned by status code, method, and host. |
Dependent item | kube.kubelet.rest.requests["{#CODE}", "{#HOST}", "{#METHOD}"] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Container memory discovery | Dependent item | kube.kubelet.container.memory.cache.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory page cache | Number of bytes of page cache memory. |
Dependent item | kube.kubelet.container.memory.cache["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Memory max usage | Maximum memory usage recorded in bytes. |
Dependent item | kube.kubelet.container.memory.max_usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: RSS | Size of RSS in bytes. |
Dependent item | kube.kubelet.container.memory.rss["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Swap | Container swap usage in bytes. |
Dependent item | kube.kubelet.container.memory.swap["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Usage | Current memory usage in bytes, including all memory regardless of when it was accessed. |
Dependent item | kube.kubelet.container.memory.usage["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#POD}] Container [{#CONTAINER}]: Working set | Current working set in bytes. |
Dependent item | kube.kubelet.container.memory.working_set["{#CONTAINER}", "{#NAMESPACE}", "{#POD}"] Preprocessing
|
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes nodes that work without any external scripts. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API. Install the Zabbix Helm Chart (https://git.zabbix.com/projects/ZT/repos/kubernetes-helm/browse?at=refs%2Fheads%2Frelease/7.4) in your Kubernetes cluster.
Change the values according to the environment in the file $HOME/zabbix_values.yaml.
For example:
Enables use of Zabbix proxy
enabled: false
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the service account name. If a different release name is used.
kubectl get serviceaccounts -n monitoring
Get the generated service account token using the command:
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set up the macros to filter the metrics of discovered nodes.
Zabbix version: 7.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command:
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chart-agent.
Set up the macros to filter the metrics of discovered nodes and host creation based on host prototypes:
Set up macros to filter pod metrics by namespace:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable pods.
You can use the {$KUBE.NODE.FILTER.LABELS}, {$KUBE.POD.FILTER.LABELS}, {$KUBE.NODE.FILTER.ANNOTATIONS} and {$KUBE.POD.FILTER.ANNOTATIONS} macros for advanced filtering of nodes and pods by labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
Note: The discovered nodes will be created as separate hosts in Zabbix with the Linux template automatically assigned to them.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.NODES.ENDPOINT.NAME} | Kubernetes nodes endpoint name. See "kubectl -n monitoring get ep". |
zabbix-zabbix-helm-chart-agent |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.ROLE.MATCHES} | Filter of discoverable nodes by role. |
.* |
| {$KUBE.LLD.FILTER.NODE.ROLE.NOT_MATCHES} | Filter to exclude discovered node by role. |
CHANGE_IF_NEEDED |
| {$KUBE.NODE.FILTER.ANNOTATIONS} | Annotations to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.NODE.FILTER.LABELS} | Labels to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.ANNOTATIONS} | Annotations to filter pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.LABELS} | Labels to filter Pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.POD.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace. |
.* |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace. |
CHANGE_IF_NEEDED |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get nodes | Collecting and processing cluster nodes data via Kubernetes API. |
Script | kube.nodes |
| Get nodes check | Data collection check. |
Dependent item | kube.nodes.check Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Failed to get nodes | length(last(/Kubernetes nodes by HTTP/kube.nodes.check))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: Get data | Collecting and processing cluster by node [{#NAME}] data via Kubernetes API. |
Dependent item | kube.node.get[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: External IP | Typically the IP address of the node that is externally routable (available from outside the cluster). |
Dependent item | kube.node.addresses.external_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: Internal IP | Typically the IP address of the node that is routable only within the cluster. |
Dependent item | kube.node.addresses.internal_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: CPU | Allocatable CPU. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Memory | Allocatable Memory. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.allocatable.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: CPU | CPU resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Memory | Memory resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.capacity.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Disk pressure | True if pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
Dependent item | kube.node.conditions.diskpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Memory pressure | True if pressure exists on the node memory - that is, if the node memory is low; otherwise False. |
Dependent item | kube.node.conditions.memorypressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Network unavailable | True if the network for the node is not correctly configured, otherwise False. |
Dependent item | kube.node.conditions.networkunavailable[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: PID pressure | True if pressure exists on the processes - that is, if there are too many processes on the node; otherwise False. |
Dependent item | kube.node.conditions.pidpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Ready | True if the node is healthy and ready to accept pods, False if the node is not healthy and is not accepting pods, and Unknown if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
Dependent item | kube.node.conditions.ready[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Architecture | Node architecture. |
Dependent item | kube.node.info.architecture[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Container runtime | Container runtime. https://kubernetes.io/docs/setup/production-environment/container-runtimes/ |
Dependent item | kube.node.info.containerruntime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kernel version | Node kernel version. |
Dependent item | kube.node.info.kernelversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kubelet version | Version of Kubelet. |
Dependent item | kube.node.info.kubeletversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: KubeProxy version | Version of KubeProxy. |
Dependent item | kube.node.info.kubeproxyversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Operating system | Node operating system. |
Dependent item | kube.node.info.operatingsystem[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: OS image | Node OS image. |
Dependent item | kube.node.info.osversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Roles | Node roles. |
Dependent item | kube.node.info.roles[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: CPU | Node CPU limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: Memory | Node Memory limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: CPU | Node CPU requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: Memory | Node Memory requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Uptime | Node uptime. |
Dependent item | kube.node.uptime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Used: Pods | Current number of pods on the node. |
Dependent item | kube.node.used.pods[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the disk size | True - pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.diskpressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the node memory | True - pressure exists on the node memory - that is, if the node memory is low; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.memorypressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Network is not correctly configured | True - the network for the node is not correctly configured, otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.networkunavailable[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the processes | True - pressure exists on the processes - that is, if there are too many processes on the node; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.pidpressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Is not in Ready state | False - if the node is not healthy and is not accepting pods. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.ready[{#NAME}])<>1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 1 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 1 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.8 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.8 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] has been restarted | Uptime is less than 10 minutes. |
last(/Kubernetes nodes by HTTP/kube.node.uptime[{#NAME}])<10 |
Info | |
| Kubernetes nodes: Node [{#NAME}] Used: Kubelet too many pods | Kubelet is running at capacity. |
last(/Kubernetes nodes by HTTP/kube.node.used.pods[{#NAME}])/ last(/Kubernetes nodes by HTTP/kube.node.capacity.pods[{#NAME}]) > 0.9 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Get data | Collecting and processing cluster by node [{#NODE}] data via Kubernetes API. |
Dependent item | kube.pod.get[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Containers ready | All containers in the Pod are ready. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.containers_ready[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Initialized | All init containers have started successfully. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.initialized[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Ready | The Pod is able to serve requests and should be added to the load balancing pools of all matching Services. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.ready[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Scheduled | The Pod has been scheduled to a node. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.scheduled[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Containers: Restarts | The number of times the container has been restarted, currently based on the number of dead containers that have not yet been removed. Note that this is calculated from dead containers. But those containers are subject to garbage collection. |
Dependent item | kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Status: Phase | The phase of a Pod is a simple, high-level summary of where the Pod is in its lifecycle. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle#pod-phase |
Dependent item | kube.pod.status.phase[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Uptime | Pod uptime. |
Dependent item | kube.pod.uptime[{#NAMESPACE}/{#POD}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}])-min(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}],15m))>1 |
Warning | |
| Kubernetes nodes: Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Status: Kubernetes Pod not healthy | Pod has been in a non-ready state for longer than 10 minutes. |
count(/Kubernetes nodes by HTTP/kube.pod.status.phase[{#NAMESPACE}/{#POD}],10m, "regexp","^(1|4|5)$")>=9 |
High |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes nodes that work without any external scripts. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API. Install the Zabbix Helm Chart (https://git.zabbix.com/projects/ZT/repos/kubernetes-helm/browse?at=refs%2Fheads%2Frelease%2F7.2) in your Kubernetes cluster.
Change the values according to the environment in the file $HOME/zabbix_values.yaml.
For example:
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set up the macros to filter the metrics of discovered nodes
Zabbix version: 7.2 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chrt-agent.
Set up the macros to filter the metrics of discovered nodes and host creation based on host prototypes:
Set up macros to filter pod metrics by namespace:
Note, If you have a large cluster, it is highly recommended to set a filter for discoverable pods.
You can use the {$KUBE.NODE.FILTER.LABELS}, {$KUBE.POD.FILTER.LABELS}, {$KUBE.NODE.FILTER.ANNOTATIONS} and {$KUBE.POD.FILTER.ANNOTATIONS} macros for advanced filtering of nodes and pods by labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
Note, the discovered nodes will be created as separate hosts in Zabbix with the Linux template automatically assigned to them.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.NODES.ENDPOINT.NAME} | Kubernetes nodes endpoint name. See "kubectl -n monitoring get ep". |
zabbix-zabbix-helm-chrt-agent |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.ROLE.MATCHES} | Filter of discoverable nodes by role. |
.* |
| {$KUBE.LLD.FILTER.NODE.ROLE.NOT_MATCHES} | Filter to exclude discovered node by role. |
CHANGE_IF_NEEDED |
| {$KUBE.NODE.FILTER.ANNOTATIONS} | Annotations to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.NODE.FILTER.LABELS} | Labels to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.ANNOTATIONS} | Annotations to filter pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.LABELS} | Labels to filter Pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.POD.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace. |
.* |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace. |
CHANGE_IF_NEEDED |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get nodes | Collecting and processing cluster nodes data via Kubernetes API. |
Script | kube.nodes |
| Get nodes check | Data collection check. |
Dependent item | kube.nodes.check Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Failed to get nodes | length(last(/Kubernetes nodes by HTTP/kube.nodes.check))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: Get data | Collecting and processing cluster by node [{#NAME}] data via Kubernetes API. |
Dependent item | kube.node.get[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: External IP | Typically the IP address of the node that is externally routable (available from outside the cluster). |
Dependent item | kube.node.addresses.external_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: Internal IP | Typically the IP address of the node that is routable only within the cluster. |
Dependent item | kube.node.addresses.internal_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: CPU | Allocatable CPU. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Memory | Allocatable Memory. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.allocatable.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: CPU | CPU resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Memory | Memory resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.capacity.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Disk pressure | True if pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
Dependent item | kube.node.conditions.diskpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Memory pressure | True if pressure exists on the node memory - that is, if the node memory is low; otherwise False. |
Dependent item | kube.node.conditions.memorypressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Network unavailable | True if the network for the node is not correctly configured, otherwise False. |
Dependent item | kube.node.conditions.networkunavailable[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: PID pressure | True if pressure exists on the processes - that is, if there are too many processes on the node; otherwise False. |
Dependent item | kube.node.conditions.pidpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Ready | True if the node is healthy and ready to accept pods, False if the node is not healthy and is not accepting pods, and Unknown if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
Dependent item | kube.node.conditions.ready[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Architecture | Node architecture. |
Dependent item | kube.node.info.architecture[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Container runtime | Container runtime. https://kubernetes.io/docs/setup/production-environment/container-runtimes/ |
Dependent item | kube.node.info.containerruntime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kernel version | Node kernel version. |
Dependent item | kube.node.info.kernelversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kubelet version | Version of Kubelet. |
Dependent item | kube.node.info.kubeletversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: KubeProxy version | Version of KubeProxy. |
Dependent item | kube.node.info.kubeproxyversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Operating system | Node operating system. |
Dependent item | kube.node.info.operatingsystem[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: OS image | Node OS image. |
Dependent item | kube.node.info.osversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Roles | Node roles. |
Dependent item | kube.node.info.roles[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: CPU | Node CPU limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: Memory | Node Memory limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: CPU | Node CPU requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: Memory | Node Memory requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Uptime | Node uptime. |
Dependent item | kube.node.uptime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Used: Pods | Current number of pods on the node. |
Dependent item | kube.node.used.pods[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the disk size | True - pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.diskpressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the node memory | True - pressure exists on the node memory - that is, if the node memory is low; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.memorypressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Network is not correctly configured | True - the network for the node is not correctly configured, otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.networkunavailable[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the processes | True - pressure exists on the processes - that is, if there are too many processes on the node; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.pidpressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Is not in Ready state | False - if the node is not healthy and is not accepting pods. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.ready[{#NAME}])<>1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 1 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 1 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.8 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.8 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] has been restarted | Uptime is less than 10 minutes. |
last(/Kubernetes nodes by HTTP/kube.node.uptime[{#NAME}])<10 |
Info | |
| Kubernetes nodes: Node [{#NAME}] Used: Kubelet too many pods | Kubelet is running at capacity. |
last(/Kubernetes nodes by HTTP/kube.node.used.pods[{#NAME}])/ last(/Kubernetes nodes by HTTP/kube.node.capacity.pods[{#NAME}]) > 0.9 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Get data | Collecting and processing cluster by node [{#NODE}] data via Kubernetes API. |
Dependent item | kube.pod.get[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Containers ready | All containers in the Pod are ready. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.containers_ready[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Initialized | All init containers have started successfully. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.initialized[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Ready | The Pod is able to serve requests and should be added to the load balancing pools of all matching Services. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.ready[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Scheduled | The Pod has been scheduled to a node. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.scheduled[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Containers: Restarts | The number of times the container has been restarted, currently based on the number of dead containers that have not yet been removed. Note that this is calculated from dead containers. But those containers are subject to garbage collection. |
Dependent item | kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Status: Phase | The phase of a Pod is a simple, high-level summary of where the Pod is in its lifecycle. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle#pod-phase |
Dependent item | kube.pod.status.phase[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Uptime | Pod uptime. |
Dependent item | kube.pod.uptime[{#NAMESPACE}/{#POD}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}])-min(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}],15m))>1 |
Warning | |
| Kubernetes nodes: Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Status: Kubernetes Pod not healthy | Pod has been in a non-ready state for longer than 10 minutes. |
count(/Kubernetes nodes by HTTP/kube.pod.status.phase[{#NAMESPACE}/{#POD}],10m, "regexp","^(1|4|5)$")>=9 |
High |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes nodes that work without any external scripts. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API. Install the Zabbix Helm Chart (https://git.zabbix.com/projects/ZT/repos/kubernetes-helm/browse?at=refs%2Fheads%2Frelease%2F7.0) in your Kubernetes cluster.
Change the values according to the environment in the file $HOME/zabbix_values.yaml.
For example:
Enables use of Zabbix proxy
enabled: false
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the service account name. If a different release name is used.
kubectl get serviceaccounts -n monitoring
Get the generated service account token using the command:
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set up the macros to filter the metrics of discovered nodes.
Zabbix version: 7.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command:
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chart-agent.
Set up the macros to filter the metrics of discovered nodes and host creation based on host prototypes:
Set up macros to filter pod metrics by namespace:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable pods.
You can use the {$KUBE.NODE.FILTER.LABELS}, {$KUBE.POD.FILTER.LABELS}, {$KUBE.NODE.FILTER.ANNOTATIONS} and {$KUBE.POD.FILTER.ANNOTATIONS} macros for advanced filtering of nodes and pods by labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
Note: The discovered nodes will be created as separate hosts in Zabbix with the Linux template automatically assigned to them.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.NODES.ENDPOINT.NAME} | Kubernetes nodes endpoint name. See "kubectl -n monitoring get ep". |
zabbix-zabbix-helm-chart-agent |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.ROLE.MATCHES} | Filter of discoverable nodes by role. |
.* |
| {$KUBE.LLD.FILTER.NODE.ROLE.NOT_MATCHES} | Filter to exclude discovered node by role. |
CHANGE_IF_NEEDED |
| {$KUBE.NODE.FILTER.ANNOTATIONS} | Annotations to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.NODE.FILTER.LABELS} | Labels to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.ANNOTATIONS} | Annotations to filter pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.LABELS} | Labels to filter Pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.POD.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace. |
.* |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace. |
CHANGE_IF_NEEDED |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get nodes | Collecting and processing cluster nodes data via Kubernetes API. |
Script | kube.nodes |
| Get nodes check | Data collection check. |
Dependent item | kube.nodes.check Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Failed to get nodes | length(last(/Kubernetes nodes by HTTP/kube.nodes.check))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: Get data | Collecting and processing cluster by node [{#NAME}] data via Kubernetes API. |
Dependent item | kube.node.get[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: External IP | Typically the IP address of the node that is externally routable (available from outside the cluster). |
Dependent item | kube.node.addresses.external_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: Internal IP | Typically the IP address of the node that is routable only within the cluster. |
Dependent item | kube.node.addresses.internal_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: CPU | Allocatable CPU. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Memory | Allocatable Memory. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.allocatable.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: CPU | CPU resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Memory | Memory resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.capacity.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Disk pressure | True if pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
Dependent item | kube.node.conditions.diskpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Memory pressure | True if pressure exists on the node memory - that is, if the node memory is low; otherwise False. |
Dependent item | kube.node.conditions.memorypressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Network unavailable | True if the network for the node is not correctly configured, otherwise False. |
Dependent item | kube.node.conditions.networkunavailable[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: PID pressure | True if pressure exists on the processes - that is, if there are too many processes on the node; otherwise False. |
Dependent item | kube.node.conditions.pidpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Ready | True if the node is healthy and ready to accept pods, False if the node is not healthy and is not accepting pods, and Unknown if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
Dependent item | kube.node.conditions.ready[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Architecture | Node architecture. |
Dependent item | kube.node.info.architecture[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Container runtime | Container runtime. https://kubernetes.io/docs/setup/production-environment/container-runtimes/ |
Dependent item | kube.node.info.containerruntime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kernel version | Node kernel version. |
Dependent item | kube.node.info.kernelversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kubelet version | Version of Kubelet. |
Dependent item | kube.node.info.kubeletversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: KubeProxy version | Version of KubeProxy. |
Dependent item | kube.node.info.kubeproxyversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Operating system | Node operating system. |
Dependent item | kube.node.info.operatingsystem[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: OS image | Node OS image. |
Dependent item | kube.node.info.osversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Roles | Node roles. |
Dependent item | kube.node.info.roles[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: CPU | Node CPU limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: Memory | Node Memory limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: CPU | Node CPU requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: Memory | Node Memory requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Uptime | Node uptime. |
Dependent item | kube.node.uptime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Used: Pods | Current number of pods on the node. |
Dependent item | kube.node.used.pods[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the disk size | True - pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.diskpressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the node memory | True - pressure exists on the node memory - that is, if the node memory is low; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.memorypressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Network is not correctly configured | True - the network for the node is not correctly configured, otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.networkunavailable[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Pressure exists on the processes | True - pressure exists on the processes - that is, if there are too many processes on the node; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.pidpressure[{#NAME}])=1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Conditions: Is not in Ready state | False - if the node is not healthy and is not accepting pods. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.ready[{#NAME}])<>1 |
Warning | |
| Kubernetes nodes: Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 1 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 1 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.8 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Kubernetes nodes: Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.8 |
Average | ||
| Kubernetes nodes: Node [{#NAME}] has been restarted | Uptime is less than 10 minutes. |
last(/Kubernetes nodes by HTTP/kube.node.uptime[{#NAME}])<10 |
Info | |
| Kubernetes nodes: Node [{#NAME}] Used: Kubelet too many pods | Kubelet is running at capacity. |
last(/Kubernetes nodes by HTTP/kube.node.used.pods[{#NAME}])/ last(/Kubernetes nodes by HTTP/kube.node.capacity.pods[{#NAME}]) > 0.9 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Get data | Collecting and processing cluster by node [{#NODE}] data via Kubernetes API. |
Dependent item | kube.pod.get[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Containers ready | All containers in the Pod are ready. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.containers_ready[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Initialized | All init containers have started successfully. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.initialized[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Ready | The Pod is able to serve requests and should be added to the load balancing pools of all matching Services. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.ready[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Conditions: Scheduled | The Pod has been scheduled to a node. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.scheduled[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Containers: Restarts | The number of times the container has been restarted, currently based on the number of dead containers that have not yet been removed. Note that this is calculated from dead containers. But those containers are subject to garbage collection. |
Dependent item | kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Status: Phase | The phase of a Pod is a simple, high-level summary of where the Pod is in its lifecycle. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle#pod-phase |
Dependent item | kube.pod.status.phase[{#NAMESPACE}/{#POD}] Preprocessing
|
| Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Uptime | Pod uptime. |
Dependent item | kube.pod.uptime[{#NAMESPACE}/{#POD}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes nodes: Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}])-min(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#NAMESPACE}/{#POD}],15m))>1 |
Warning | |
| Kubernetes nodes: Node [{#NODE}] Namespace [{#NAMESPACE}] Pod [{#POD}] Status: Kubernetes Pod not healthy | Pod has been in a non-ready state for longer than 10 minutes. |
count(/Kubernetes nodes by HTTP/kube.pod.status.phase[{#NAMESPACE}/{#POD}],10m, "regexp","^(1|4|5)$")>=9 |
High |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes nodes that work without any external scripts.
It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Install the Zabbix Helm Chart (https://git.zabbix.com/projects/ZT/repos/kubernetes-helm/browse?at=refs%2Fheads%2Frelease%2F6.4) in your Kubernetes cluster.
Change the values according to the environment in the file $HOME/zabbix_values.yaml.
For example:
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set up the macros to filter the metrics of discovered nodes
Zabbix version: 6.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chrt-agent.
Set up the macros to filter the metrics of discovered nodes and host creation based on host prototypes:
Set up macros to filter pod metrics by namespace:
Note, If you have a large cluster, it is highly recommended to set a filter for discoverable pods.
You can use the {$KUBE.NODE.FILTER.LABELS}, {$KUBE.POD.FILTER.LABELS}, {$KUBE.NODE.FILTER.ANNOTATIONS} and {$KUBE.POD.FILTER.ANNOTATIONS} macros for advanced filtering of nodes and pods by labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
Note, the discovered nodes will be created as separate hosts in Zabbix with the Linux template automatically assigned to them.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.NODES.ENDPOINT.NAME} | Kubernetes nodes endpoint name. See "kubectl -n monitoring get ep". |
zabbix-zabbix-helm-chrt-agent |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.ROLE.MATCHES} | Filter of discoverable nodes by role. |
.* |
| {$KUBE.LLD.FILTER.NODE.ROLE.NOT_MATCHES} | Filter to exclude discovered node by role. |
CHANGE_IF_NEEDED |
| {$KUBE.NODE.FILTER.ANNOTATIONS} | Annotations to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.NODE.FILTER.LABELS} | Labels to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.ANNOTATIONS} | Annotations to filter pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.LABELS} | Labels to filter Pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.POD.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace. |
.* |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace. |
CHANGE_IF_NEEDED |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Get nodes | Collecting and processing cluster nodes data via Kubernetes API. |
Script | kube.nodes |
| Get nodes check | Data collection check. |
Dependent item | kube.nodes.check Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Failed to get nodes | length(last(/Kubernetes nodes by HTTP/kube.nodes.check))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: Get data | Collecting and processing cluster by node [{#NAME}] data via Kubernetes API. |
Dependent item | kube.node.get[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: External IP | Typically the IP address of the node that is externally routable (available from outside the cluster). |
Dependent item | kube.node.addresses.external_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: Internal IP | Typically the IP address of the node that is routable only within the cluster. |
Dependent item | kube.node.addresses.internal_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: CPU | Allocatable CPU. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Memory | Allocatable Memory. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.allocatable.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: CPU | CPU resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Memory | Memory resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.capacity.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Disk pressure | True if pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
Dependent item | kube.node.conditions.diskpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Memory pressure | True if pressure exists on the node memory - that is, if the node memory is low; otherwise False. |
Dependent item | kube.node.conditions.memorypressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Network unavailable | True if the network for the node is not correctly configured, otherwise False. |
Dependent item | kube.node.conditions.networkunavailable[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: PID pressure | True if pressure exists on the processes - that is, if there are too many processes on the node; otherwise False. |
Dependent item | kube.node.conditions.pidpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Ready | True if the node is healthy and ready to accept pods, False if the node is not healthy and is not accepting pods, and Unknown if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
Dependent item | kube.node.conditions.ready[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Architecture | Node architecture. |
Dependent item | kube.node.info.architecture[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Container runtime | Container runtime. https://kubernetes.io/docs/setup/production-environment/container-runtimes/ |
Dependent item | kube.node.info.containerruntime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kernel version | Node kernel version. |
Dependent item | kube.node.info.kernelversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kubelet version | Version of Kubelet. |
Dependent item | kube.node.info.kubeletversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: KubeProxy version | Version of KubeProxy. |
Dependent item | kube.node.info.kubeproxyversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Operating system | Node operating system. |
Dependent item | kube.node.info.operatingsystem[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: OS image | Node OS image. |
Dependent item | kube.node.info.osversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Roles | Node roles. |
Dependent item | kube.node.info.roles[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: CPU | Node CPU limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: Memory | Node Memory limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: CPU | Node CPU requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: Memory | Node Memory requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Uptime | Node uptime. |
Dependent item | kube.node.uptime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Used: Pods | Current number of pods on the node. |
Dependent item | kube.node.used.pods[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Node [{#NAME}] Conditions: Pressure exists on the disk size | True - pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.diskpressure[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Pressure exists on the node memory | True - pressure exists on the node memory - that is, if the node memory is low; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.memorypressure[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Network is not correctly configured | True - the network for the node is not correctly configured, otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.networkunavailable[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Pressure exists on the processes | True - pressure exists on the processes - that is, if there are too many processes on the node; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.pidpressure[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Is not in Ready state | False - if the node is not healthy and is not accepting pods. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.ready[{#NAME}])<>1 |
Warning | |
| Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 1 |
Average | ||
| Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 1 |
Average | ||
| Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.8 |
Average | ||
| Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.8 |
Average | ||
| Node [{#NAME}]: Has been restarted | Uptime is less than 10 minutes. |
last(/Kubernetes nodes by HTTP/kube.node.uptime[{#NAME}])<10 |
Info | |
| Node [{#NAME}] Used: Kubelet too many pods | Kubelet is running at capacity. |
last(/Kubernetes nodes by HTTP/kube.node.used.pods[{#NAME}])/ last(/Kubernetes nodes by HTTP/kube.node.capacity.pods[{#NAME}]) > 0.9 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NODE}] Pod [{#POD}]: Get data | Collecting and processing cluster by node [{#NODE}] data via Kubernetes API. |
Dependent item | kube.pod.get[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Containers ready | All containers in the Pod are ready. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.containers_ready[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Initialized | All init containers have started successfully. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.initialized[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Ready | The Pod is able to serve requests and should be added to the load balancing pools of all matching Services. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.ready[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Scheduled | The Pod has been scheduled to a node. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.scheduled[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Containers: Restarts | The number of times the container has been restarted, currently based on the number of dead containers that have not yet been removed. Note that this is calculated from dead containers. But those containers are subject to garbage collection. |
Dependent item | kube.pod.containers.restartcount[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Status: Phase | The phase of a Pod is a simple, high-level summary of where the Pod is in its lifecycle. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle#pod-phase |
Dependent item | kube.pod.status.phase[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Uptime | Pod uptime. |
Dependent item | kube.pod.uptime[{#POD}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Node [{#NODE}] Pod [{#POD}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#POD}])-min(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#POD}],15m))>1 |
Warning | |
| Node [{#NODE}] Pod [{#POD}] Status: Kubernetes Pod not healthy | Pod has been in a non-ready state for longer than 10 minutes. |
count(/Kubernetes nodes by HTTP/kube.pod.status.phase[{#POD}],10m, "regexp","^(1|4|5)$")>=9 |
High |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
For Zabbix version: 6.2 and higher.
The template to monitor Kubernetes nodes that work without any external scripts.
It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Install the Zabbix Helm Chart (https://git.zabbix.com/projects/ZT/repos/kubernetes-helm/browse?at=refs%2Fheads%2Frelease%2F6.2) in your Kubernetes cluster.
Set the {$KUBE.API.ENDPOINT.URL} such as <scheme>://<host>:<port>/api.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chrt-agent.
Set up the macros to filter the metrics of discovered nodes
This template was tested on:
See Zabbix template operation for basic instructions.
Install the Zabbix Helm Chart in your Kubernetes cluster.
Set the {$KUBE.API.ENDPOINT.URL} such as <scheme>://<host>:<port>/api.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chrt-agent.
Set up the macros to filter the metrics of discovered nodes:
Set up the macros to filter host creation based on host prototypes:
Set up macros to filter pod metrics by namespace:
Note, If you have a large cluster, it is highly recommended to set a filter for discoverable pods.
You can use {$KUBE.NODE.FILTER.LABELS}, {$KUBE.POD.FILTER.LABELS}, {$KUBE.NODE.FILTER.ANNOTATIONS} and {$KUBE.POD.FILTER.ANNOTATIONS} macros for advanced filtering nodes and pods by labels and annotations. Macro values are specified separated by commas and must have the key/value form with support for regular expressions in the value.
For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the nodes 5-25 without the "ingress" role will be discovered.
See documentation for details:
Note, the discovered nodes will be created as separate hosts in Zabbix with the Linux template automatically assigned to them.
No specific Zabbix configuration is required.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.ENDPOINT.URL} | Kubernetes API endpoint URL in the format |
https://localhost:6443/api |
| {$KUBE.API.TOKEN} | Service account bearer token |
`` |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.ROLE.MATCHES} | Filter of discoverable nodes by role |
.* |
| {$KUBE.LLD.FILTER.NODE.ROLE.NOT_MATCHES} | Filter to exclude discovered node by role |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE_HOST.MATCHES} | Filter of discoverable cluster nodes |
.* |
| {$KUBE.LLD.FILTER.NODE_HOST.NOT_MATCHES} | Filter to exclude discovered cluster nodes |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE_HOST.ROLE.MATCHES} | Filter of discoverable nodes hosts by role |
.* |
| {$KUBE.LLD.FILTER.NODE_HOST.ROLE.NOT_MATCHES} | Filter to exclude discovered cluster nodes by role |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace |
.* |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace |
CHANGE_IF_NEEDED |
| {$KUBE.NODE.FILTER.ANNOTATIONS} | Annotations to filter nodes (regex in values are supported) |
`` |
| {$KUBE.NODE.FILTER.LABELS} | Labels to filter nodes (regex in values are supported) |
`` |
| {$KUBE.NODES.ENDPOINT.NAME} | Kubernetes nodes endpoint name. See kubectl -n monitoring get ep |
zabbix-zabbix-helm-chrt-agent |
| {$KUBE.POD.FILTER.ANNOTATIONS} | Annotations to filter pods (regex in values are supported) |
`` |
| {$KUBE.POD.FILTER.LABELS} | Labels to filter Pods (regex in values are supported) |
`` |
There are no template links in this template.
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Cluster node discovery | - |
DEPENDENT | kube.node_host.discovery Filter: AND- {#NAME} MATCHES_REGEX - {#NAME} NOT_MATCHES_REGEX - {#ROLES} MATCHES_REGEX - {#ROLES} NOT_MATCHES_REGEX |
| Node discovery | - |
DEPENDENT | kube.node.discovery Filter: AND- {#NAME} MATCHES_REGEX - {#NAME} NOT_MATCHES_REGEX - {#ROLES} MATCHES_REGEX - {#ROLES} NOT_MATCHES_REGEX |
| Pod discovery | - |
DEPENDENT | kube.pod.discovery Preprocessing: - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NODE} MATCHES_REGEX - {#NODE} NOT_MATCHES_REGEX - {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Group | Name | Description | Type | Key and additional info |
|---|---|---|---|---|
| Kubernetes | Kubernetes: Get nodes | Collecting and processing cluster nodes data via Kubernetes API. |
SCRIPT | kube.nodes Expression: The text is too long. Please see the template. |
| Kubernetes | Get nodes check | Data collection check. |
DEPENDENT | kube.nodes.check Preprocessing: - JSONPATH: ⛔️ON_FAIL: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node LLD | Generation of data for node discovery rules. |
DEPENDENT | kube.nodes.lld Preprocessing: - JAVASCRIPT: `function parseFilters(filter) { var pairs = {}; filter.split(/\s*,\s*/).forEach(function (kv) { if (/([\w.-]+/[\w.-]+):\s*.+/.test(kv)) { var pair = kv.split(/\s*:\s*/); pairs[pair[0]] = pair[1]; } }); return pairs; } function filter(name, data, filters) { var filtered = true; if (typeof data === 'object') { Object.keys(filters).some(function (filter) { var exclude = filter.match(/^!(.+)/); if (filter in data |
| Kubernetes | Node [{#NAME}]: Get data | Collecting and processing cluster by node [{#NAME}] data via Kubernetes API. |
DEPENDENT | kube.node.get[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Addresses: External IP | Typically the IP address of the node that is externally routable (available from outside the cluster). |
DEPENDENT | kube.node.addresses.external_ip[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Addresses: Internal IP | Typically the IP address of the node that is routable only within the cluster. |
DEPENDENT | kube.node.addresses.internal_ip[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Allocatable: CPU | Allocatable CPU. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
DEPENDENT | kube.node.allocatable.cpu[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Allocatable: Memory | Allocatable Memory. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
DEPENDENT | kube.node.allocatable.memory[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Allocatable: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
DEPENDENT | kube.node.allocatable.pods[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Capacity: CPU | CPU resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
DEPENDENT | kube.node.capacity.cpu[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Capacity: Memory | Memory resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
DEPENDENT | kube.node.capacity.memory[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Capacity: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
DEPENDENT | kube.node.capacity.pods[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Conditions: Disk pressure | True if pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
DEPENDENT | kube.node.conditions.diskpressure[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NAME}] Conditions: Memory pressure | True if pressure exists on the node memory - that is, if the node memory is low; otherwise False. |
DEPENDENT | kube.node.conditions.memorypressure[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NAME}] Conditions: Network unavailable | True if the network for the node is not correctly configured, otherwise False. |
DEPENDENT | kube.node.conditions.networkunavailable[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NAME}] Conditions: PID pressure | True if pressure exists on the processes - that is, if there are too many processes on the node; otherwise False. |
DEPENDENT | kube.node.conditions.pidpressure[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NAME}] Conditions: Ready | True if the node is healthy and ready to accept pods, False if the node is not healthy and is not accepting pods, and Unknown if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
DEPENDENT | kube.node.conditions.ready[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NAME}] Info: Architecture | Node architecture. |
DEPENDENT | kube.node.info.architecture[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: Container runtime | Container runtime. https://kubernetes.io/docs/setup/production-environment/container-runtimes/ |
DEPENDENT | kube.node.info.containerruntime[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: Kernel version | Node kernel version. |
DEPENDENT | kube.node.info.kernelversion[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: Kubelet version | Version of Kubelet. |
DEPENDENT | kube.node.info.kubeletversion[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: KubeProxy version | Version of KubeProxy. |
DEPENDENT | kube.node.info.kubeproxyversion[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: Operating system | Node operating system. |
DEPENDENT | kube.node.info.operatingsystem[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: OS image | Node OS image. |
DEPENDENT | kube.node.info.osversion[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Info: Roles | Node roles. |
DEPENDENT | kube.node.info.roles[{#NAME}] Preprocessing: - JSONPATH: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Node [{#NAME}] Limits: CPU | Node CPU limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
DEPENDENT | kube.node.limits.cpu[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Limits: Memory | Node Memory limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
DEPENDENT | kube.node.limits.memory[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Requests: CPU | Node CPU requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
DEPENDENT | kube.node.requests.cpu[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Requests: Memory | Node Memory requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
DEPENDENT | kube.node.requests.memory[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NAME}] Uptime | Node uptime. |
DEPENDENT | kube.node.uptime[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: |
| Kubernetes | Node [{#NAME}] Used: Pods | Current number of pods on the node. |
DEPENDENT | kube.node.used.pods[{#NAME}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NODE}] Pod [{#POD}]: Get data | Collecting and processing cluster by node [{#NODE}] data via Kubernetes API. |
DEPENDENT | kube.pod.get[{#POD}] Preprocessing: - JSONPATH: |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Conditions: Containers ready | All containers in the Pod are ready. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
DEPENDENT | kube.pod.conditions.containers_ready[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Conditions: Initialized | All init containers have started successfully. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
DEPENDENT | kube.pod.conditions.initialized[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Conditions: Ready | The Pod is able to serve requests and should be added to the load balancing pools of all matching Services. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
DEPENDENT | kube.pod.conditions.ready[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Conditions: Scheduled | The Pod has been scheduled to a node. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
DEPENDENT | kube.pod.conditions.scheduled[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['True', 'False', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Containers: Restarts | The number of times the container has been restarted, currently based on the number of dead containers that have not yet been removed. Note that this is calculated from dead containers. But those containers are subject to garbage collection. |
DEPENDENT | kube.pod.containers.restartcount[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Status: Phase | The phase of a Pod is a simple, high-level summary of where the Pod is in its lifecycle. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle#pod-phase |
DEPENDENT | kube.pod.status.phase[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: `return ['Pending', 'Running', 'Succeeded', 'Failed', 'Unknown'].indexOf(value) + 1 |
| Kubernetes | Node [{#NODE}] Pod [{#POD}] Uptime | Pod uptime. |
DEPENDENT | kube.pod.uptime[{#POD}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: - JAVASCRIPT: |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Failed to get nodes | - |
length(last(/Kubernetes nodes by HTTP/kube.nodes.check))>0 |
WARNING | |
| Node [{#NAME}] Conditions: Pressure exists on the disk size | True - pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.diskpressure[{#NAME}])=1 |
WARNING | |
| Node [{#NAME}] Conditions: Pressure exists on the node memory | True - pressure exists on the node memory - that is, if the node memory is low; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.memorypressure[{#NAME}])=1 |
WARNING | |
| Node [{#NAME}] Conditions: Network is not correctly configured | True - the network for the node is not correctly configured, otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.networkunavailable[{#NAME}])=1 |
WARNING | |
| Node [{#NAME}] Conditions: Pressure exists on the processes | True - pressure exists on the processes - that is, if there are too many processes on the node; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.pidpressure[{#NAME}])=1 |
WARNING | |
| Node [{#NAME}] Conditions: Is not in Ready state | False - if the node is not healthy and is not accepting pods. Unknown - if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
last(/Kubernetes nodes by HTTP/kube.node.conditions.ready[{#NAME}])<>1 |
WARNING | |
| Node [{#NAME}] Limits: Total CPU limits are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.9 |
WARNING | Depends on: - Node [{#NAME}] Limits: Total CPU limits are too high |
| Node [{#NAME}] Limits: Total CPU limits are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 1 |
AVERAGE | |
| Node [{#NAME}] Limits: Total memory limits are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.9 |
WARNING | Depends on: - Node [{#NAME}] Limits: Total memory limits are too high |
| Node [{#NAME}] Limits: Total memory limits are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 1 |
AVERAGE | |
| Node [{#NAME}] Requests: Total CPU requests are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.5 |
WARNING | Depends on: - Node [{#NAME}] Requests: Total CPU requests are too high |
| Node [{#NAME}] Requests: Total CPU requests are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.8 |
AVERAGE | |
| Node [{#NAME}] Requests: Total memory requests are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.5 |
WARNING | Depends on: - Node [{#NAME}] Requests: Total memory requests are too high |
| Node [{#NAME}] Requests: Total memory requests are too high | - |
last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.8 |
AVERAGE | |
| Node [{#NAME}]: Has been restarted | Uptime is less than 10 minutes |
last(/Kubernetes nodes by HTTP/kube.node.uptime[{#NAME}])<10 |
INFO | |
| Node [{#NAME}] Used: Kubelet too many pods | Kubelet is running at capacity. |
last(/Kubernetes nodes by HTTP/kube.node.used.pods[{#NAME}])/ last(/Kubernetes nodes by HTTP/kube.node.capacity.pods[{#NAME}]) > 0.9 |
WARNING | |
| Node [{#NODE}] Pod [{#POD}]: Pod is crash looping | Pos restarts more than 2 times in the last 3 minutes. |
(last(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#POD}])-min(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#POD}],3m))>2 |
WARNING | |
| Node [{#NODE}] Pod [{#POD}] Status: Kubernetes Pod not healthy | Pod has been in a non-ready state for longer than 10 minutes. |
`count(/Kubernetes nodes by HTTP/kube.pod.status.phase[{#POD}],10m, "regexp","^(1 | 4 | 5)$")>=9` |
Please report any issues with the template at https://support.zabbix.com.
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums.
The template to monitor Kubernetes nodes that work without any external scripts.
It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Install the Zabbix Helm Chart (https://git.zabbix.com/projects/ZT/repos/kubernetes-helm/browse?at=refs%2Fheads%2Frelease%2F6.0) in your Kubernetes cluster.
Change the values according to the environment in the file $HOME/zabbix_values.yaml.
For example:
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set up the macros to filter the metrics of discovered nodes
Zabbix version: 6.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.NODES.ENDPOINT.NAME} with Zabbix agent's endpoint name. See kubectl -n monitoring get ep. Default: zabbix-zabbix-helm-chrt-agent.
Set up the macros to filter the metrics of discovered nodes and host creation based on host prototypes:
Set up macros to filter pod metrics by namespace:
Note, If you have a large cluster, it is highly recommended to set a filter for discoverable pods.
You can use the {$KUBE.NODE.FILTER.LABELS}, {$KUBE.POD.FILTER.LABELS}, {$KUBE.NODE.FILTER.ANNOTATIONS} and {$KUBE.POD.FILTER.ANNOTATIONS} macros for advanced filtering of nodes and pods by labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
Note, the discovered nodes will be created as separate hosts in Zabbix with the Linux template automatically assigned to them.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.NODES.ENDPOINT.NAME} | Kubernetes nodes endpoint name. See "kubectl -n monitoring get ep". |
zabbix-zabbix-helm-chrt-agent |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.ROLE.MATCHES} | Filter of discoverable nodes by role. |
.* |
| {$KUBE.LLD.FILTER.NODE.ROLE.NOT_MATCHES} | Filter to exclude discovered node by role. |
CHANGE_IF_NEEDED |
| {$KUBE.NODE.FILTER.ANNOTATIONS} | Annotations to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.NODE.FILTER.LABELS} | Labels to filter nodes (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.ANNOTATIONS} | Annotations to filter pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.POD.FILTER.LABELS} | Labels to filter Pods (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.POD.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace. |
.* |
| {$KUBE.LLD.FILTER.POD.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace. |
CHANGE_IF_NEEDED |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Get nodes | Collecting and processing cluster nodes data via Kubernetes API. |
Script | kube.nodes |
| Get nodes check | Data collection check. |
Dependent item | kube.nodes.check Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Failed to get nodes | length(last(/Kubernetes nodes by HTTP/kube.nodes.check))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: Get data | Collecting and processing cluster by node [{#NAME}] data via Kubernetes API. |
Dependent item | kube.node.get[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: External IP | Typically the IP address of the node that is externally routable (available from outside the cluster). |
Dependent item | kube.node.addresses.external_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Addresses: Internal IP | Typically the IP address of the node that is routable only within the cluster. |
Dependent item | kube.node.addresses.internal_ip[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: CPU | Allocatable CPU. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Memory | Allocatable Memory. 'Allocatable' on a Kubernetes node is defined as the amount of compute resources that are available for pods. The scheduler does not over-subscribe 'Allocatable'. 'CPU', 'memory' and 'ephemeral-storage' are supported as of now. |
Dependent item | kube.node.allocatable.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Allocatable: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.allocatable.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: CPU | CPU resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Memory | Memory resource capacity. https://kubernetes.io/docs/concepts/architecture/nodes/#capacity |
Dependent item | kube.node.capacity.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Capacity: Pods | https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/ |
Dependent item | kube.node.capacity.pods[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Disk pressure | True if pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
Dependent item | kube.node.conditions.diskpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Memory pressure | True if pressure exists on the node memory - that is, if the node memory is low; otherwise False. |
Dependent item | kube.node.conditions.memorypressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Network unavailable | True if the network for the node is not correctly configured, otherwise False. |
Dependent item | kube.node.conditions.networkunavailable[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: PID pressure | True if pressure exists on the processes - that is, if there are too many processes on the node; otherwise False. |
Dependent item | kube.node.conditions.pidpressure[{#NAME}] Preprocessing
|
| Node [{#NAME}] Conditions: Ready | True if the node is healthy and ready to accept pods, False if the node is not healthy and is not accepting pods, and Unknown if the node controller has not heard from the node in the last node-monitor-grace-period (default is 40 seconds). |
Dependent item | kube.node.conditions.ready[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Architecture | Node architecture. |
Dependent item | kube.node.info.architecture[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Container runtime | Container runtime. https://kubernetes.io/docs/setup/production-environment/container-runtimes/ |
Dependent item | kube.node.info.containerruntime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kernel version | Node kernel version. |
Dependent item | kube.node.info.kernelversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Kubelet version | Version of Kubelet. |
Dependent item | kube.node.info.kubeletversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: KubeProxy version | Version of KubeProxy. |
Dependent item | kube.node.info.kubeproxyversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Operating system | Node operating system. |
Dependent item | kube.node.info.operatingsystem[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: OS image | Node OS image. |
Dependent item | kube.node.info.osversion[{#NAME}] Preprocessing
|
| Node [{#NAME}] Info: Roles | Node roles. |
Dependent item | kube.node.info.roles[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: CPU | Node CPU limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Limits: Memory | Node Memory limits. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.limits.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: CPU | Node CPU requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.cpu[{#NAME}] Preprocessing
|
| Node [{#NAME}] Requests: Memory | Node Memory requests. https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ |
Dependent item | kube.node.requests.memory[{#NAME}] Preprocessing
|
| Node [{#NAME}] Uptime | Node uptime. |
Dependent item | kube.node.uptime[{#NAME}] Preprocessing
|
| Node [{#NAME}] Used: Pods | Current number of pods on the node. |
Dependent item | kube.node.used.pods[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Node [{#NAME}] Conditions: Pressure exists on the disk size | True - pressure exists on the disk size - that is, if the disk capacity is low; otherwise False. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.diskpressure[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Pressure exists on the node memory | True - pressure exists on the node memory - that is, if the node memory is low; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.memorypressure[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Network is not correctly configured | True - the network for the node is not correctly configured, otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.networkunavailable[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Pressure exists on the processes | True - pressure exists on the processes - that is, if there are too many processes on the node; otherwise False |
last(/Kubernetes nodes by HTTP/kube.node.conditions.pidpressure[{#NAME}])=1 |
Warning | |
| Node [{#NAME}] Conditions: Is not in Ready state | False - if the node is not healthy and is not accepting pods. |
last(/Kubernetes nodes by HTTP/kube.node.conditions.ready[{#NAME}])<>1 |
Warning | |
| Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Node [{#NAME}] Limits: Total CPU limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 1 |
Average | ||
| Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.9 |
Warning | Depends on:
|
|
| Node [{#NAME}] Limits: Total memory limits are too high | last(/Kubernetes nodes by HTTP/kube.node.limits.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 1 |
Average | ||
| Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Node [{#NAME}] Requests: Total CPU requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.cpu[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.cpu[{#NAME}]) > 0.8 |
Average | ||
| Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.5 |
Warning | Depends on:
|
|
| Node [{#NAME}] Requests: Total memory requests are too high | last(/Kubernetes nodes by HTTP/kube.node.requests.memory[{#NAME}]) / last(/Kubernetes nodes by HTTP/kube.node.allocatable.memory[{#NAME}]) > 0.8 |
Average | ||
| Node [{#NAME}]: Has been restarted | Uptime is less than 10 minutes. |
last(/Kubernetes nodes by HTTP/kube.node.uptime[{#NAME}])<10 |
Info | |
| Node [{#NAME}] Used: Kubelet too many pods | Kubelet is running at capacity. |
last(/Kubernetes nodes by HTTP/kube.node.used.pods[{#NAME}])/ last(/Kubernetes nodes by HTTP/kube.node.capacity.pods[{#NAME}]) > 0.9 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NODE}] Pod [{#POD}]: Get data | Collecting and processing cluster by node [{#NODE}] data via Kubernetes API. |
Dependent item | kube.pod.get[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Containers ready | All containers in the Pod are ready. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.containers_ready[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Initialized | All init containers have started successfully. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.initialized[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Ready | The Pod is able to serve requests and should be added to the load balancing pools of all matching Services. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.ready[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Conditions: Scheduled | The Pod has been scheduled to a node. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions |
Dependent item | kube.pod.conditions.scheduled[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Containers: Restarts | The number of times the container has been restarted, currently based on the number of dead containers that have not yet been removed. Note that this is calculated from dead containers. But those containers are subject to garbage collection. |
Dependent item | kube.pod.containers.restartcount[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Status: Phase | The phase of a Pod is a simple, high-level summary of where the Pod is in its lifecycle. https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle#pod-phase |
Dependent item | kube.pod.status.phase[{#POD}] Preprocessing
|
| Node [{#NODE}] Pod [{#POD}] Uptime | Pod uptime. |
Dependent item | kube.pod.uptime[{#POD}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Node [{#NODE}] Pod [{#POD}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#POD}])-min(/Kubernetes nodes by HTTP/kube.pod.containers.restartcount[{#POD}],15m))>1 |
Warning | |
| Node [{#NODE}] Pod [{#POD}] Status: Kubernetes Pod not healthy | Pod has been in a non-ready state for longer than 10 minutes. |
count(/Kubernetes nodes by HTTP/kube.pod.status.phase[{#POD}],10m, "regexp","^(1|4|5)$")>=9 |
High |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Scheduler by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Scheduler by HTTP - collects metrics by HTTP agent from Scheduler /metrics endpoint.
Zabbix version: 7.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.SCHEDULER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: You might need to set the --binding-address option for Scheduler to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
Note: Some metrics may not be collected depending on your Kubernetes Scheduler instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.SCHEDULER.SERVER.URL} | Kubernetes Scheduler metrics endpoint URL. |
https://localhost:10259/metrics |
| {$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.SCHEDULER.UNSCHEDULABLE} | Maximum number of scheduling failures with 'unschedulable' used for trigger. |
2 |
| {$KUBE.SCHEDULER.ERROR} | Maximum number of scheduling failures with 'error' used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get Scheduler metrics | Get raw metrics from Scheduler instance /metrics endpoint. |
HTTP agent | kubernetes.scheduler.get_metrics Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.scheduler.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.scheduler.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.scheduler.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.scheduler.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.scheduler.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.scheduler.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.scheduler.max_fds Preprocessing
|
| REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_200.rate Preprocessing
|
| REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_300.rate Preprocessing
|
| REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_400.rate Preprocessing
|
| REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_500.rate Preprocessing
|
| Schedule attempts: scheduled | Number of attempts to schedule pods with result "scheduled" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.scheduled.rate Preprocessing
|
| Schedule attempts: unschedulable | Number of attempts to schedule pods with result "unschedulable" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate Preprocessing
|
| Schedule attempts: error | Number of attempts to schedule pods with result "error" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.error.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler: Too many REST Client errors | "Kubernetes Scheduler REST Client requests is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.client_http_requests_500.rate,5m)>{$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} |
Warning | |
| Kubernetes Scheduler: Too many unschedulable pods | Number of attempts to schedule pods with 'unschedulable' result is too high. 'unschedulable' means a pod could not be scheduled. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate,5m)>{$KUBE.SCHEDULER.UNSCHEDULABLE} |
Warning | |
| Kubernetes Scheduler: Too many schedule attempts with errors | Number of attempts to schedule pods with 'error' result is too high. 'error' means an internal scheduler problem. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.error.rate,5m)>{$KUBE.SCHEDULER.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm histogram | Discovery raw data of scheduling algorithm latency. |
Dependent item | kubernetes.scheduler.scheduling_algorithm.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm duration bucket, {#LE} | Scheduling algorithm latency in seconds. |
Dependent item | kubernetes.scheduler.scheduling_algorithm_duration[{#LE}] Preprocessing
|
| Scheduling algorithm duration, p90 | 90 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p90[{#SINGLETON}] |
| Scheduling algorithm duration, p95 | 95 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p95[{#SINGLETON}] |
| Scheduling algorithm duration, p99 | 99 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p99[{#SINGLETON}] |
| Scheduling algorithm duration, p50 | 50 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding histogram | Discovery raw data of binding latency. |
Dependent item | kubernetes.scheduler.binding.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding duration bucket, {#LE} | Binding latency in seconds. |
Dependent item | kubernetes.scheduler.binding_duration[{#LE}] Preprocessing
|
| Binding duration, p90 | 90 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p90[{#SINGLETON}] |
| Binding duration, p95 | 99 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p95[{#SINGLETON}] |
| Binding duration, p99 | 95 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p99[{#SINGLETON}] |
| Binding duration, p50 | 50 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| e2e scheduling histogram | Discovery raw data and percentile items of e2e scheduling latency. |
Dependent item | kubernetes.controller.e2e_scheduling.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#RESULT}"]: e2e scheduling seconds bucket, {#LE} | E2e scheduling latency in seconds (scheduling algorithm + binding) |
Dependent item | kubernetes.scheduler.e2e_scheduling_bucket[{#LE},"{#RESULT}"] Preprocessing
|
| ["{#RESULT}"]: e2e scheduling, p50 | 50 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p50["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p90 | 90 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p90["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p95 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p95["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p99 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p99["{#RESULT}"] |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Scheduler by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Scheduler by HTTP - collects metrics by HTTP agent from Scheduler /metrics endpoint.
Zabbix version: 7.2 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.SCHEDULER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. You might need to set the --binding-address option for Scheduler to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
NOTE. Some metrics may not be collected depending on your Kubernetes Scheduler instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.SCHEDULER.SERVER.URL} | Kubernetes Scheduler metrics endpoint URL. |
https://localhost:10259/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.SCHEDULER.UNSCHEDULABLE} | Maximum number of scheduling failures with 'unschedulable' used for trigger. |
2 |
| {$KUBE.SCHEDULER.ERROR} | Maximum number of scheduling failures with 'error' used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get Scheduler metrics | Get raw metrics from Scheduler instance /metrics endpoint. |
HTTP agent | kubernetes.scheduler.get_metrics Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.scheduler.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.scheduler.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.scheduler.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.scheduler.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.scheduler.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.scheduler.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.scheduler.max_fds Preprocessing
|
| REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_200.rate Preprocessing
|
| REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_300.rate Preprocessing
|
| REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_400.rate Preprocessing
|
| REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_500.rate Preprocessing
|
| Schedule attempts: scheduled | Number of attempts to schedule pods with result "scheduled" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.scheduled.rate Preprocessing
|
| Schedule attempts: unschedulable | Number of attempts to schedule pods with result "unschedulable" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate Preprocessing
|
| Schedule attempts: error | Number of attempts to schedule pods with result "error" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.error.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler: Too many REST Client errors | "Kubernetes Scheduler REST Client requests is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.client_http_requests_500.rate,5m)>{$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} |
Warning | |
| Kubernetes Scheduler: Too many unschedulable pods | Number of attempts to schedule pods with 'unschedulable' result is too high. 'unschedulable' means a pod could not be scheduled. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate,5m)>{$KUBE.SCHEDULER.UNSCHEDULABLE} |
Warning | |
| Kubernetes Scheduler: Too many schedule attempts with errors | Number of attempts to schedule pods with 'error' result is too high. 'error' means an internal scheduler problem. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.error.rate,5m)>{$KUBE.SCHEDULER.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm histogram | Discovery raw data of scheduling algorithm latency. |
Dependent item | kubernetes.scheduler.scheduling_algorithm.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm duration bucket, {#LE} | Scheduling algorithm latency in seconds. |
Dependent item | kubernetes.scheduler.scheduling_algorithm_duration[{#LE}] Preprocessing
|
| Scheduling algorithm duration, p90 | 90 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p90[{#SINGLETON}] |
| Scheduling algorithm duration, p95 | 95 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p95[{#SINGLETON}] |
| Scheduling algorithm duration, p99 | 99 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p99[{#SINGLETON}] |
| Scheduling algorithm duration, p50 | 50 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding histogram | Discovery raw data of binding latency. |
Dependent item | kubernetes.scheduler.binding.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding duration bucket, {#LE} | Binding latency in seconds. |
Dependent item | kubernetes.scheduler.binding_duration[{#LE}] Preprocessing
|
| Binding duration, p90 | 90 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p90[{#SINGLETON}] |
| Binding duration, p95 | 99 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p95[{#SINGLETON}] |
| Binding duration, p99 | 95 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p99[{#SINGLETON}] |
| Binding duration, p50 | 50 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| e2e scheduling histogram | Discovery raw data and percentile items of e2e scheduling latency. |
Dependent item | kubernetes.controller.e2e_scheduling.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#RESULT}"]: e2e scheduling seconds bucket, {#LE} | E2e scheduling latency in seconds (scheduling algorithm + binding) |
Dependent item | kubernetes.scheduler.e2e_scheduling_bucket[{#LE},"{#RESULT}"] Preprocessing
|
| ["{#RESULT}"]: e2e scheduling, p50 | 50 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p50["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p90 | 90 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p90["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p95 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p95["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p99 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p99["{#RESULT}"] |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Scheduler by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Scheduler by HTTP - collects metrics by HTTP agent from Scheduler /metrics endpoint.
Zabbix version: 7.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.SCHEDULER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: You might need to set the --binding-address option for Scheduler to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
Note: Some metrics may not be collected depending on your Kubernetes Scheduler instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.SCHEDULER.SERVER.URL} | Kubernetes Scheduler metrics endpoint URL. |
https://localhost:10259/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.SCHEDULER.UNSCHEDULABLE} | Maximum number of scheduling failures with 'unschedulable' used for trigger. |
2 |
| {$KUBE.SCHEDULER.ERROR} | Maximum number of scheduling failures with 'error' used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get Scheduler metrics | Get raw metrics from Scheduler instance /metrics endpoint. |
HTTP agent | kubernetes.scheduler.get_metrics Preprocessing
|
| Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.scheduler.process_virtual_memory_bytes Preprocessing
|
| Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.scheduler.process_resident_memory_bytes Preprocessing
|
| CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.scheduler.cpu.util Preprocessing
|
| Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.scheduler.go_goroutines Preprocessing
|
| Go threads | Number of OS threads created. |
Dependent item | kubernetes.scheduler.go_threads Preprocessing
|
| Fds open | Number of open file descriptors. |
Dependent item | kubernetes.scheduler.open_fds Preprocessing
|
| Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.scheduler.max_fds Preprocessing
|
| REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_200.rate Preprocessing
|
| REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_300.rate Preprocessing
|
| REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_400.rate Preprocessing
|
| REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_500.rate Preprocessing
|
| Schedule attempts: scheduled | Number of attempts to schedule pods with result "scheduled" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.scheduled.rate Preprocessing
|
| Schedule attempts: unschedulable | Number of attempts to schedule pods with result "unschedulable" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate Preprocessing
|
| Schedule attempts: error | Number of attempts to schedule pods with result "error" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.error.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler: Too many REST Client errors | "Kubernetes Scheduler REST Client requests is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.client_http_requests_500.rate,5m)>{$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} |
Warning | |
| Kubernetes Scheduler: Too many unschedulable pods | Number of attempts to schedule pods with 'unschedulable' result is too high. 'unschedulable' means a pod could not be scheduled. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate,5m)>{$KUBE.SCHEDULER.UNSCHEDULABLE} |
Warning | |
| Kubernetes Scheduler: Too many schedule attempts with errors | Number of attempts to schedule pods with 'error' result is too high. 'error' means an internal scheduler problem. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.error.rate,5m)>{$KUBE.SCHEDULER.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm histogram | Discovery raw data of scheduling algorithm latency. |
Dependent item | kubernetes.scheduler.scheduling_algorithm.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm duration bucket, {#LE} | Scheduling algorithm latency in seconds. |
Dependent item | kubernetes.scheduler.scheduling_algorithm_duration[{#LE}] Preprocessing
|
| Scheduling algorithm duration, p90 | 90 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p90[{#SINGLETON}] |
| Scheduling algorithm duration, p95 | 95 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p95[{#SINGLETON}] |
| Scheduling algorithm duration, p99 | 99 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p99[{#SINGLETON}] |
| Scheduling algorithm duration, p50 | 50 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding histogram | Discovery raw data of binding latency. |
Dependent item | kubernetes.scheduler.binding.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding duration bucket, {#LE} | Binding latency in seconds. |
Dependent item | kubernetes.scheduler.binding_duration[{#LE}] Preprocessing
|
| Binding duration, p90 | 90 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p90[{#SINGLETON}] |
| Binding duration, p95 | 99 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p95[{#SINGLETON}] |
| Binding duration, p99 | 95 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p99[{#SINGLETON}] |
| Binding duration, p50 | 50 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| e2e scheduling histogram | Discovery raw data and percentile items of e2e scheduling latency. |
Dependent item | kubernetes.controller.e2e_scheduling.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ["{#RESULT}"]: e2e scheduling seconds bucket, {#LE} | E2e scheduling latency in seconds (scheduling algorithm + binding) |
Dependent item | kubernetes.scheduler.e2e_scheduling_bucket[{#LE},"{#RESULT}"] Preprocessing
|
| ["{#RESULT}"]: e2e scheduling, p50 | 50 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p50["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p90 | 90 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p90["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p95 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p95["{#RESULT}"] |
| ["{#RESULT}"]: e2e scheduling, p99 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p99["{#RESULT}"] |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes Scheduler by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Scheduler by HTTP - collects metrics by HTTP agent from Scheduler /metrics endpoint.
Zabbix version: 6.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.SCHEDULER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. You might need to set the --binding-address option for Scheduler to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
NOTE. Some metrics may not be collected depending on your Kubernetes Scheduler instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.SCHEDULER.SERVER.URL} | Kubernetes Scheduler metrics endpoint URL. |
https://localhost:10259/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.SCHEDULER.UNSCHEDULABLE} | Maximum number of scheduling failures with 'unschedulable' used for trigger. |
2 |
| {$KUBE.SCHEDULER.ERROR} | Maximum number of scheduling failures with 'error' used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: Get Scheduler metrics | Get raw metrics from Scheduler instance /metrics endpoint. |
HTTP agent | kubernetes.scheduler.get_metrics Preprocessing
|
| Kubernetes Scheduler: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.scheduler.process_virtual_memory_bytes Preprocessing
|
| Kubernetes Scheduler: Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.scheduler.process_resident_memory_bytes Preprocessing
|
| Kubernetes Scheduler: CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.scheduler.cpu.util Preprocessing
|
| Kubernetes Scheduler: Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.scheduler.go_goroutines Preprocessing
|
| Kubernetes Scheduler: Go threads | Number of OS threads created. |
Dependent item | kubernetes.scheduler.go_threads Preprocessing
|
| Kubernetes Scheduler: Fds open | Number of open file descriptors. |
Dependent item | kubernetes.scheduler.open_fds Preprocessing
|
| Kubernetes Scheduler: Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.scheduler.max_fds Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_200.rate Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_300.rate Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_400.rate Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_500.rate Preprocessing
|
| Kubernetes Scheduler: Schedule attempts: scheduled | Number of attempts to schedule pods with result "scheduled" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.scheduled.rate Preprocessing
|
| Kubernetes Scheduler: Schedule attempts: unschedulable | Number of attempts to schedule pods with result "unschedulable" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate Preprocessing
|
| Kubernetes Scheduler: Schedule attempts: error | Number of attempts to schedule pods with result "error" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.error.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler: Too many REST Client errors | "Kubernetes Scheduler REST Client requests is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.client_http_requests_500.rate,5m)>{$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} |
Warning | |
| Kubernetes Scheduler: Too many unschedulable pods | Number of attempts to schedule pods with 'unschedulable' result is too high. 'unschedulable' means a pod could not be scheduled. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate,5m)>{$KUBE.SCHEDULER.UNSCHEDULABLE} |
Warning | |
| Kubernetes Scheduler: Too many schedule attempts with errors | Number of attempts to schedule pods with 'error' result is too high. 'error' means an internal scheduler problem. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.error.rate,5m)>{$KUBE.SCHEDULER.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm histogram | Discovery raw data of scheduling algorithm latency. |
Dependent item | kubernetes.scheduler.scheduling_algorithm.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: Scheduling algorithm duration bucket, {#LE} | Scheduling algorithm latency in seconds. |
Dependent item | kubernetes.scheduler.scheduling_algorithm_duration[{#LE}] Preprocessing
|
| Kubernetes Scheduler: Scheduling algorithm duration, p90 | 90 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p90[{#SINGLETON}] |
| Kubernetes Scheduler: Scheduling algorithm duration, p95 | 95 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p95[{#SINGLETON}] |
| Kubernetes Scheduler: Scheduling algorithm duration, p99 | 99 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p99[{#SINGLETON}] |
| Kubernetes Scheduler: Scheduling algorithm duration, p50 | 50 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding histogram | Discovery raw data of binding latency. |
Dependent item | kubernetes.scheduler.binding.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: Binding duration bucket, {#LE} | Binding latency in seconds. |
Dependent item | kubernetes.scheduler.binding_duration[{#LE}] Preprocessing
|
| Kubernetes Scheduler: Binding duration, p90 | 90 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p90[{#SINGLETON}] |
| Kubernetes Scheduler: Binding duration, p95 | 99 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p95[{#SINGLETON}] |
| Kubernetes Scheduler: Binding duration, p99 | 95 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p99[{#SINGLETON}] |
| Kubernetes Scheduler: Binding duration, p50 | 50 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| e2e scheduling histogram | Discovery raw data and percentile items of e2e scheduling latency. |
Dependent item | kubernetes.controller.e2e_scheduling.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling seconds bucket, {#LE} | E2e scheduling latency in seconds (scheduling algorithm + binding) |
Dependent item | kubernetes.scheduler.e2e_scheduling_bucket[{#LE},"{#RESULT}"] Preprocessing
|
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p50 | 50 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p50["{#RESULT}"] |
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p90 | 90 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p90["{#RESULT}"] |
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p95 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p95["{#RESULT}"] |
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p99 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p99["{#RESULT}"] |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
For Zabbix version: 6.2 and higher
The template to monitor Kubernetes Scheduler by Zabbix that works without any external scripts.
Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Scheduler by HTTP — collects metrics by HTTP agent from Scheduler /metrics endpoint.
This template was tested on:
See Zabbix template operation for basic instructions.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.SCHEDULER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values. NOTE. Some metrics may not be collected depending on your Kubernetes Scheduler instance version and configuration.
No specific Zabbix configuration is required.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | API Authorization Token |
`` |
| {$KUBE.SCHEDULER.ERROR} | Maximum number of scheduling failures with 'error' used for trigger |
2 |
| {$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger |
2 |
| {$KUBE.SCHEDULER.SERVER.URL} | Instance URL |
http://localhost:10251/metrics |
| {$KUBE.SCHEDULER.UNSCHEDULABLE} | Maximum number of scheduling failures with 'unschedulable' used for trigger |
2 |
There are no template links in this template.
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding histogram | Discovery raw data of binding latency. |
DEPENDENT | kubernetes.scheduler.binding.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| e2e scheduling histogram | Discovery raw data and percentile items of e2e scheduling latency. |
DEPENDENT | kubernetes.controller.e2e_scheduling.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| Scheduling algorithm histogram | Discovery raw data of scheduling algorithm latency. |
DEPENDENT | kubernetes.scheduler.scheduling_algorithm.discovery Preprocessing: - PROMETHEUS_TO_JSON: - JAVASCRIPT: - DISCARD_UNCHANGED_HEARTBEAT: Overrides: bucket item total item |
| Group | Name | Description | Type | Key and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler | Kubernetes Scheduler: Virtual memory, bytes | Virtual memory size in bytes. |
DEPENDENT | kubernetes.scheduler.process_virtual_memory_bytes Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: Resident memory, bytes | Resident memory size in bytes. |
DEPENDENT | kubernetes.scheduler.process_resident_memory_bytes Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: CPU | Total user and system CPU usage ratio. |
DEPENDENT | kubernetes.scheduler.cpu.util Preprocessing: - PROMETHEUS_PATTERN: - CHANGE_PER_SECOND - MULTIPLIER: |
| Kubernetes Scheduler | Kubernetes Scheduler: Goroutines | Number of goroutines that currently exist. |
DEPENDENT | kubernetes.scheduler.go_goroutines Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: Go threads | Number of OS threads created. |
DEPENDENT | kubernetes.scheduler.go_threads Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: Fds open | Number of open file descriptors. |
DEPENDENT | kubernetes.scheduler.open_fds Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: Fds max | Maximum allowed open file descriptors. |
DEPENDENT | kubernetes.scheduler.max_fds Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
DEPENDENT | kubernetes.scheduler.client_http_requests_200.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
DEPENDENT | kubernetes.scheduler.client_http_requests_300.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
DEPENDENT | kubernetes.scheduler.client_http_requests_400.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
DEPENDENT | kubernetes.scheduler.client_http_requests_500.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: Schedule attempts: scheduled | Number of attempts to schedule pods with result "scheduled" per second. |
DEPENDENT | kubernetes.scheduler.scheduler_schedule_attempts.scheduled.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: Schedule attempts: unschedulable | Number of attempts to schedule pods with result "unschedulable" per second. |
DEPENDENT | kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: Schedule attempts: error | Number of attempts to schedule pods with result "error" per second. |
DEPENDENT | kubernetes.scheduler.scheduler_schedule_attempts.error.rate Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - CHANGE_PER_SECOND |
| Kubernetes Scheduler | Kubernetes Scheduler: Scheduling algorithm duration bucket, {#LE} | Scheduling algorithm latency in seconds. |
DEPENDENT | kubernetes.scheduler.scheduling_algorithm_duration[{#LE}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: Scheduling algorithm duration, p90 | 90 percentile of scheduling algorithm latency in seconds. |
CALCULATED | kubernetes.scheduler.scheduling_algorithm_duration_p90[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.scheduling_algorithm_duration[*],5m,90) |
| Kubernetes Scheduler | Kubernetes Scheduler: Scheduling algorithm duration, p95 | 95 percentile of scheduling algorithm latency in seconds. |
CALCULATED | kubernetes.scheduler.scheduling_algorithm_duration_p95[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.scheduling_algorithm_duration[*],5m,95) |
| Kubernetes Scheduler | Kubernetes Scheduler: Scheduling algorithm duration, p99 | 99 percentile of scheduling algorithm latency in seconds. |
CALCULATED | kubernetes.scheduler.scheduling_algorithm_duration_p99[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.scheduling_algorithm_duration[*],5m,99) |
| Kubernetes Scheduler | Kubernetes Scheduler: Scheduling algorithm duration, p50 | 50 percentile of scheduling algorithm latency in seconds. |
CALCULATED | kubernetes.scheduler.scheduling_algorithm_duration_p50[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.scheduling_algorithm_duration[*],5m,50) |
| Kubernetes Scheduler | Kubernetes Scheduler: Binding duration bucket, {#LE} | Binding latency in seconds. |
DEPENDENT | kubernetes.scheduler.binding_duration[{#LE}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: Binding duration, p90 | 90 percentile of binding latency in seconds. |
CALCULATED | kubernetes.scheduler.binding_duration_p90[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.binding_duration[*],5m,90) |
| Kubernetes Scheduler | Kubernetes Scheduler: Binding duration, p95 | 99 percentile of binding latency in seconds. |
CALCULATED | kubernetes.scheduler.binding_duration_p95[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.binding_duration[*],5m,95) |
| Kubernetes Scheduler | Kubernetes Scheduler: Binding duration, p99 | 95 percentile of binding latency in seconds. |
CALCULATED | kubernetes.scheduler.binding_duration_p99[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.binding_duration[*],5m,99) |
| Kubernetes Scheduler | Kubernetes Scheduler: Binding duration, p50 | 50 percentile of binding latency in seconds. |
CALCULATED | kubernetes.scheduler.binding_duration_p50[{#SINGLETON}] Expression: bucket_percentile(//kubernetes.scheduler.binding_duration[*],5m,50) |
| Kubernetes Scheduler | Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling seconds bucket, {#LE} | E2e scheduling latency in seconds (scheduling algorithm + binding) |
DEPENDENT | kubernetes.scheduler.e2e_scheduling_bucket[{#LE},"{#RESULT}"] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes Scheduler | Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p50 | 50 percentile of e2e scheduling latency. |
CALCULATED | kubernetes.scheduler.e2e_scheduling_p50["{#RESULT}"] Expression: bucket_percentile(//kubernetes.scheduler.e2e_scheduling_bucket[*,"{#RESULT}"],5m,50) |
| Kubernetes Scheduler | Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p90 | 90 percentile of e2e scheduling latency. |
CALCULATED | kubernetes.scheduler.e2e_scheduling_p90["{#RESULT}"] Expression: bucket_percentile(//kubernetes.scheduler.e2e_scheduling_bucket[*,"{#RESULT}"],5m,90) |
| Kubernetes Scheduler | Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p95 | 95 percentile of e2e scheduling latency. |
CALCULATED | kubernetes.scheduler.e2e_scheduling_p95["{#RESULT}"] Expression: bucket_percentile(//kubernetes.scheduler.e2e_scheduling_bucket[*,"{#RESULT}"],5m,95) |
| Kubernetes Scheduler | Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p99 | 95 percentile of e2e scheduling latency. |
CALCULATED | kubernetes.scheduler.e2e_scheduling_p99["{#RESULT}"] Expression: bucket_percentile(//kubernetes.scheduler.e2e_scheduling_bucket[*,"{#RESULT}"],5m,99) |
| Zabbix raw items | Kubernetes Scheduler: Get Scheduler metrics | Get raw metrics from Scheduler instance /metrics endpoint. |
HTTP_AGENT | kubernetes.scheduler.get_metrics Preprocessing: - CHECK_NOT_SUPPORTED ⛔️ON_FAIL: |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler: Too many REST Client errors | "Kubernetes Scheduler REST Client requests is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.client_http_requests_500.rate,5m)>{$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} |
WARNING | |
| Kubernetes Scheduler: Too many unschedulable pods | "Number of attempts to schedule pods with 'unschedulable' result is too high. 'unschedulable' means a pod could not be scheduled." |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate,5m)>{$KUBE.SCHEDULER.UNSCHEDULABLE} |
WARNING | |
| Kubernetes Scheduler: Too many schedule attempts with errors | "Number of attempts to schedule pods with 'error' result is too high. 'error' means an internal scheduler problem." |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.error.rate,5m)>{$KUBE.SCHEDULER.ERROR} |
WARNING |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template or ask for help with it at ZABBIX forums.
The template to monitor Kubernetes Scheduler by Zabbix that works without any external scripts. Most of the metrics are collected in one go, thanks to Zabbix bulk data collection.
Template Kubernetes Scheduler by HTTP - collects metrics by HTTP agent from Scheduler /metrics endpoint.
Zabbix version: 6.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Internal service metrics are collected from /metrics endpoint. Template needs to use Authorization via API token.
Don't forget change macros {$KUBE.SCHEDULER.SERVER.URL}, {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. You might need to set the --binding-address option for Scheduler to the address where Zabbix proxy can reach it.
For example, for clusters created with kubeadm it can be set in the following manifest file (changes will be applied immediately):
NOTE. Some metrics may not be collected depending on your Kubernetes Scheduler instance version and configuration.
| Name | Description | Default |
|---|---|---|
| {$KUBE.SCHEDULER.SERVER.URL} | Kubernetes Scheduler metrics endpoint URL. |
https://localhost:10259/metrics |
| {$KUBE.API.TOKEN} | API Authorization Token. |
|
| {$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} | Maximum number of HTTP client requests failures used for trigger. |
2 |
| {$KUBE.SCHEDULER.UNSCHEDULABLE} | Maximum number of scheduling failures with 'unschedulable' used for trigger. |
2 |
| {$KUBE.SCHEDULER.ERROR} | Maximum number of scheduling failures with 'error' used for trigger. |
2 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: Get Scheduler metrics | Get raw metrics from Scheduler instance /metrics endpoint. |
HTTP agent | kubernetes.scheduler.get_metrics Preprocessing
|
| Kubernetes Scheduler: Virtual memory, bytes | Virtual memory size in bytes. |
Dependent item | kubernetes.scheduler.process_virtual_memory_bytes Preprocessing
|
| Kubernetes Scheduler: Resident memory, bytes | Resident memory size in bytes. |
Dependent item | kubernetes.scheduler.process_resident_memory_bytes Preprocessing
|
| Kubernetes Scheduler: CPU | Total user and system CPU usage ratio. |
Dependent item | kubernetes.scheduler.cpu.util Preprocessing
|
| Kubernetes Scheduler: Goroutines | Number of goroutines that currently exist. |
Dependent item | kubernetes.scheduler.go_goroutines Preprocessing
|
| Kubernetes Scheduler: Go threads | Number of OS threads created. |
Dependent item | kubernetes.scheduler.go_threads Preprocessing
|
| Kubernetes Scheduler: Fds open | Number of open file descriptors. |
Dependent item | kubernetes.scheduler.open_fds Preprocessing
|
| Kubernetes Scheduler: Fds max | Maximum allowed open file descriptors. |
Dependent item | kubernetes.scheduler.max_fds Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 2xx, rate | Number of HTTP requests with 2xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_200.rate Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 3xx, rate | Number of HTTP requests with 3xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_300.rate Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 4xx, rate | Number of HTTP requests with 4xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_400.rate Preprocessing
|
| Kubernetes Scheduler: REST Client requests: 5xx, rate | Number of HTTP requests with 5xx status code per second. |
Dependent item | kubernetes.scheduler.client_http_requests_500.rate Preprocessing
|
| Kubernetes Scheduler: Schedule attempts: scheduled | Number of attempts to schedule pods with result "scheduled" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.scheduled.rate Preprocessing
|
| Kubernetes Scheduler: Schedule attempts: unschedulable | Number of attempts to schedule pods with result "unschedulable" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate Preprocessing
|
| Kubernetes Scheduler: Schedule attempts: error | Number of attempts to schedule pods with result "error" per second. |
Dependent item | kubernetes.scheduler.scheduler_schedule_attempts.error.rate Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes Scheduler: Too many REST Client errors | "Kubernetes Scheduler REST Client requests is experiencing high error rate (with 5xx HTTP code). |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.client_http_requests_500.rate,5m)>{$KUBE.SCHEDULER.HTTP.CLIENT.ERROR} |
Warning | |
| Kubernetes Scheduler: Too many unschedulable pods | Number of attempts to schedule pods with 'unschedulable' result is too high. 'unschedulable' means a pod could not be scheduled. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.unschedulable.rate,5m)>{$KUBE.SCHEDULER.UNSCHEDULABLE} |
Warning | |
| Kubernetes Scheduler: Too many schedule attempts with errors | Number of attempts to schedule pods with 'error' result is too high. 'error' means an internal scheduler problem. |
min(/Kubernetes Scheduler by HTTP/kubernetes.scheduler.scheduler_schedule_attempts.error.rate,5m)>{$KUBE.SCHEDULER.ERROR} |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduling algorithm histogram | Discovery raw data of scheduling algorithm latency. |
Dependent item | kubernetes.scheduler.scheduling_algorithm.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: Scheduling algorithm duration bucket, {#LE} | Scheduling algorithm latency in seconds. |
Dependent item | kubernetes.scheduler.scheduling_algorithm_duration[{#LE}] Preprocessing
|
| Kubernetes Scheduler: Scheduling algorithm duration, p90 | 90 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p90[{#SINGLETON}] |
| Kubernetes Scheduler: Scheduling algorithm duration, p95 | 95 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p95[{#SINGLETON}] |
| Kubernetes Scheduler: Scheduling algorithm duration, p99 | 99 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p99[{#SINGLETON}] |
| Kubernetes Scheduler: Scheduling algorithm duration, p50 | 50 percentile of scheduling algorithm latency in seconds. |
Calculated | kubernetes.scheduler.scheduling_algorithm_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Binding histogram | Discovery raw data of binding latency. |
Dependent item | kubernetes.scheduler.binding.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: Binding duration bucket, {#LE} | Binding latency in seconds. |
Dependent item | kubernetes.scheduler.binding_duration[{#LE}] Preprocessing
|
| Kubernetes Scheduler: Binding duration, p90 | 90 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p90[{#SINGLETON}] |
| Kubernetes Scheduler: Binding duration, p95 | 99 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p95[{#SINGLETON}] |
| Kubernetes Scheduler: Binding duration, p99 | 95 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p99[{#SINGLETON}] |
| Kubernetes Scheduler: Binding duration, p50 | 50 percentile of binding latency in seconds. |
Calculated | kubernetes.scheduler.binding_duration_p50[{#SINGLETON}] |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| e2e scheduling histogram | Discovery raw data and percentile items of e2e scheduling latency. |
Dependent item | kubernetes.controller.e2e_scheduling.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling seconds bucket, {#LE} | E2e scheduling latency in seconds (scheduling algorithm + binding) |
Dependent item | kubernetes.scheduler.e2e_scheduling_bucket[{#LE},"{#RESULT}"] Preprocessing
|
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p50 | 50 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p50["{#RESULT}"] |
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p90 | 90 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p90["{#RESULT}"] |
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p95 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p95["{#RESULT}"] |
| Kubernetes Scheduler: ["{#RESULT}"]: e2e scheduling, p99 | 95 percentile of e2e scheduling latency. |
Calculated | kubernetes.scheduler.e2e_scheduling_p99["{#RESULT}"] |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes state. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Template Kubernetes cluster state by HTTP - collects metrics by HTTP agent from kube-state-metrics endpoint and Kubernetes API.
Don't forget to change macros {$KUBE.API.URL} and {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: Some metrics may not be collected depending on your Kubernetes version and configuration.
Zabbix version: 7.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster. Internal service metrics are collected from kube-state-metrics endpoint.
Template needs to use authorization via API token.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the service account name. If a different release name is used.
kubectl get serviceaccounts -n monitoring
Get the generated service account token using the command:
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.STATE.ENDPOINT.NAME} with Kube state metrics endpoint name. See kubectl -n monitoring get ep. Default: zabbix-kube-state-metrics.
Note: If you wish to monitor Controller Manager and Scheduler components, you might need to set the --binding-address option for them to the address where Zabbix proxy can reach them.
For example, for clusters created with kubeadm it can be set in the following manifest files (changes will be applied immediately):
Depending on your Kubernetes distribution, you might need to adjust {$KUBE.CONTROL_PLANE.TAINT} macro (for example, set it to node-role.kubernetes.io/master for OpenShift).
Note: Some metrics may not be collected depending on your Kubernetes version and configuration.
Also, see the Macros section for a list of macros used to set trigger values.
Set up the macros to filter the metrics of discovered Kubelets by node names:
Set up macros to filter metrics by namespace:
Set up macros to filter node metrics by nodename:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable namespaces.
You can use the {$KUBE.KUBELET.FILTER.LABELS} and {$KUBE.KUBELET.FILTER.ANNOTATIONS} macros for advanced filtering of kubelets by node labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the kubelets on nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
You can also set up evaluation periods for replica mismatch triggers (Deployments, ReplicaSets, StatefulSets) with the macro {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD}, which supports context and regular expressions. For example, you can create the following macros:
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:default:nginx-deployment"} = #3
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"deployment:.*:.*"} = #10 or {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"^deployment.*"} = #10
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:".*:default:.*"} = 15m
Note: that different context macros with regular expressions matching the same string can be applied in an undefined order, and simple context macros (without regular expressions) have higher priority. Read the Important notes section in Zabbix documentation for details.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.LIVEZ.ENDPOINT} | Kubernetes API livez endpoint /livez |
/livez |
| {$KUBE.API.COMPONENTSTATUSES.ENDPOINT} | Kubernetes API componentstatuses endpoint /api/v1/componentstatuses |
/api/v1/componentstatuses |
| {$KUBE.API.READYZ.ENDPOINT} | Kubernetes API readyz endpoint /readyz |
/readyz |
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.STATE.ENDPOINT.NAME} | Endpoint name for kube-state-metrics service (check with |
zabbix-kube-state-metrics |
| {$OPENSHIFT.STATE.ENDPOINT.NAME} | OpenShift state endpoint name. |
openshift-state-metrics |
| {$KUBE.API_SERVER.SCHEME} | Kubernetes API servers metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.API_SERVER.PORT} | Kubernetes API servers metrics endpoint port. Used in ControlPlane LLD. |
6443 |
| {$KUBE.CONTROL_PLANE.TAINT} | Taint that applies to control plane nodes. Change if needed. Used in ControlPlane LLD. |
node-role.kubernetes.io/control-plane |
| {$KUBE.CONTROLLER_MANAGER.SCHEME} | Kubernetes Controller manager metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.CONTROLLER_MANAGER.PORT} | Kubernetes Controller manager metrics endpoint port. Used in ControlPlane LLD. |
10257 |
| {$KUBE.SCHEDULER.SCHEME} | Kubernetes Scheduler metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.SCHEDULER.PORT} | Kubernetes Scheduler metrics endpoint port. Used in ControlPlane LLD. |
10259 |
| {$KUBE.KUBELET.SCHEME} | Kubernetes Kubelet metrics endpoint scheme. Used in Kubelet LLD. |
https |
| {$KUBE.KUBELET.PORT} | Kubernetes Kubelet metrics endpoint port. Used in Kubelet LLD. |
10250 |
| {$KUBE.LLD.FILTER.NAMESPACE.MATCHES} | Filter of discoverable metrics by namespace. |
.* |
| {$KUBE.LLD.FILTER.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered metrics by namespace. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes by nodename. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.KUBELET_NODE.MATCHES} | Filter of discoverable Kubelets by nodename. |
.* |
| {$KUBE.LLD.FILTER.KUBELET_NODE.NOT_MATCHES} | Filter to exclude discovered Kubelets by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.KUBELET.FILTER.ANNOTATIONS} | Node annotations to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.KUBELET.FILTER.LABELS} | Node labels to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.PV.MATCHES} | Filter of discoverable persistent volumes by name. |
.* |
| {$KUBE.LLD.FILTER.PV.NOT_MATCHES} | Filter to exclude discovered persistent volumes by name. |
CHANGE_IF_NEEDED |
| {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD} | The evaluation period range which is used for calculation of expressions in trigger prototypes (time period or value range). Can be used with context. |
#5 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get state metrics | Collecting Kubernetes metrics from kube-state-metrics. |
Script | kube.state.metrics |
| Control plane LLD | Generation of data for Control plane discovery rules. |
Script | kube.control_plane.lld Preprocessing
|
| Node LLD | Generation of data for Kubelet discovery rules. |
Script | kube.node.lld Preprocessing
|
| Get component statuses | HTTP agent | kube.componentstatuses Preprocessing
|
|
| Get readyz | HTTP agent | kube.readyz Preprocessing
|
|
| Get livez | HTTP agent | kube.livez Preprocessing
|
|
| Namespace count | The number of namespaces. |
Dependent item | kube.namespace.count Preprocessing
|
| CronJob count | Number of cronjobs. |
Dependent item | kube.cronjob.count Preprocessing
|
| Job count | Number of jobs (generated by cronjob + job). |
Dependent item | kube.job.count Preprocessing
|
| Endpoint count | Number of endpoints. |
Dependent item | kube.endpoint.count Preprocessing
|
| Deployment count | The number of deployments. |
Dependent item | kube.deployment.count Preprocessing
|
| Service count | The number of services. |
Dependent item | kube.service.count Preprocessing
|
| StatefulSet count | The number of statefulsets. |
Dependent item | kube.statefulset.count Preprocessing
|
| Node count | The number of nodes. |
Dependent item | kube.node.count Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| API servers discovery | Dependent item | kube.api_servers.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Controller manager nodes discovery | Dependent item | kube.controller_manager.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduler servers nodes discovery | Dependent item | kube.scheduler.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubelet discovery | Dependent item | kube.kubelet.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Daemonset discovery | Dependent item | kube.daemonset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Ready | The number of nodes that should be running the daemon pod and have one or more running and ready. |
Dependent item | kube.daemonset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Scheduled | The number of nodes that run at least one daemon pod and are supposed to. |
Dependent item | kube.daemonset.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Desired | The number of nodes that should be running the daemon pod. |
Dependent item | kube.daemonset.desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Misscheduled | The number of nodes that run a daemon pod but are not supposed to. |
Dependent item | kube.daemonset.misscheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Updated number scheduled | The total number of nodes that are running updated daemon pod. |
Dependent item | kube.daemonset.updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PVC discovery | Dependent item | kube.pvc.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase | The current status phase of the persistent volume claim. |
Dependent item | kube.pvc.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC [{#NAME}] Requested storage | The capacity of storage requested by the persistent volume claim. |
Dependent item | kube.pvc.requested.storage[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Bound, sum | The total amount of persistent volume claims in the Bound phase. |
Dependent item | kube.pvc.status_phase.bound.sum[{#NAMESPACE}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Lost, sum | The total amount of persistent volume claims in the Lost phase. |
Dependent item | kube.pvc.status_phase.lost.sum[{#NAMESPACE}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Pending, sum | The total amount of persistent volume claims in the Pending phase. |
Dependent item | kube.pvc.status_phase.pending.sum[{#NAMESPACE}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: NS [{#NAMESPACE}] PVC [{#NAME}]: PVC is pending | count(/Kubernetes cluster state by HTTP/kube.pvc.status_phase[{#NAMESPACE}/{#NAME}],2m,,5)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV discovery | Dependent item | kube.pv.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV [{#NAME}] Status phase | The current status phase of the persistent volume. |
Dependent item | kube.pv.status_phase[{#NAME}] Preprocessing
|
| PV [{#NAME}] Capacity bytes | A capacity of the persistent volume in bytes. |
Dependent item | kube.pv.capacity.bytes[{#NAME}] Preprocessing
|
| PV status phase: Pending, sum | The total amount of persistent volumes in the Pending phase. |
Dependent item | kube.pv.status_phase.pending.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Available, sum | The total amount of persistent volumes in the Available phase. |
Dependent item | kube.pv.status_phase.available.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Bound, sum | The total amount of persistent volumes in the Bound phase. |
Dependent item | kube.pv.status_phase.bound.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Released, sum | The total amount of persistent volumes in the Released phase. |
Dependent item | kube.pv.status_phase.released.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Failed, sum | The total amount of persistent volumes in the Failed phase. |
Dependent item | kube.pv.status_phase.failed.sum[{#SINGLETON}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: PV [{#NAME}]: PV has failed | count(/Kubernetes cluster state by HTTP/kube.pv.status_phase[{#NAME}],2m,,3)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Deployment discovery | Dependent item | kube.deployment.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Paused | Whether the deployment is paused and will not be processed by the deployment controller. |
Dependent item | kube.deployment.spec_paused[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas desired | Number of desired pods for a deployment. |
Dependent item | kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Rollingupdate max unavailable | Maximum number of unavailable replicas during a rolling update of a deployment. |
Dependent item | kube.deployment.rollingupdate.max_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas | The number of replicas per deployment. |
Dependent item | kube.deployment.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas available | The number of available replicas per deployment. |
Dependent item | kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas unavailable | The number of unavailable replicas per deployment. |
Dependent item | kube.deployment.replicas_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas updated | The number of updated replicas per deployment. |
Dependent item | kube.deployment.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas mismatched | The number of available replicas not matching the desired number of replicas. |
Dependent item | kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Deployment replicas mismatch | Deployment has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Endpoint discovery | Dependent item | kube.endpoint.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address available | Number of addresses available in endpoint. |
Dependent item | kube.endpoint.address_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address not ready | Number of addresses not ready in endpoint. |
Dependent item | kube.endpoint.address_not_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Age | Endpoint age (number of seconds since creation). |
Dependent item | kube.endpoint.age[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: CPU allocatable | The CPU resources of a node that are available for scheduling. |
Dependent item | kube.node.cpu_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Memory allocatable | The memory resources of a node that are available for scheduling. |
Dependent item | kube.node.memory_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Pods allocatable | The pods resources of a node that are available for scheduling. |
Dependent item | kube.node.pods_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Ephemeral storage allocatable | The allocatable ephemeral storage of a node that is available for scheduling. |
Dependent item | kube.node.ephemeral_storage_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: CPU capacity | The capacity for CPU resources of a node. |
Dependent item | kube.node.cpu_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Memory capacity | The capacity for memory resources of a node. |
Dependent item | kube.node.memory_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Ephemeral storage capacity | The ephemeral storage capacity of a node. |
Dependent item | kube.node.ephemeral_storage_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Pods capacity | The capacity for pods resources of a node. |
Dependent item | kube.node.pods_capacity[{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Pending | Pod is in pending state. |
Dependent item | kube.pod.phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Succeeded | Pod is in succeeded state. |
Dependent item | kube.pod.phase.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Failed | Pod is in failed state. |
Dependent item | kube.pod.phase.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Unknown | Pod is in unknown state. |
Dependent item | kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Running | Pod is in unknown state. |
Dependent item | kube.pod.phase.running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers terminated | Describes whether the container is currently in terminated state. |
Dependent item | kube.pod.containers_terminated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers waiting | Describes whether the container is currently in waiting state. |
Dependent item | kube.pod.containers_waiting[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers ready | Describes whether the containers readiness check succeeded. |
Dependent item | kube.pod.containers_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers restarts | The number of container restarts. |
Dependent item | kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers running | Describes whether the container is currently in running state. |
Dependent item | kube.pod.containers_running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Ready | Describes whether the pod is ready to serve requests. |
Dependent item | kube.pod.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Scheduled | Describes the status of the scheduling process for the pod. |
Dependent item | kube.pod.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Unschedulable | Describes the unschedulable status for the pod. |
Dependent item | kube.pod.unschedulable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU limits | The limit on CPU cores to be used by a container. |
Dependent item | kube.pod.containers.limits.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory limits | The limit on memory to be used by a container. |
Dependent item | kube.pod.containers.limits.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU requests | The number of requested CPU cores by a container. |
Dependent item | kube.pod.containers.requests.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory requests | The number of requested memory bytes by a container. |
Dependent item | kube.pod.containers.requests.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is not healthy | min(/Kubernetes cluster state by HTTP/kube.pod.phase.failed[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.pending[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}],10m)>0 |
High | ||
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}])-min(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}],15m))>1 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ReplicaSet discovery | Dependent item | kube.replicaset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas | The number of replicas per ReplicaSet. |
Dependent item | kube.replicaset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Desired replicas | Number of desired pods for a ReplicaSet. |
Dependent item | kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Fully labeled replicas | The number of fully labeled replicas per ReplicaSet. |
Dependent item | kube.replicaset.fully_labeled_replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Ready | The number of ready replicas per ReplicaSet. |
Dependent item | kube.replicaset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the desired number of replicas. |
Dependent item | kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] RS [{#NAME}]: ReplicaSet mismatch | ReplicaSet has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"replicaset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| StatefulSet discovery | Dependent item | kube.statefulset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas | The number of replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Desired replicas | Number of desired pods for a StatefulSet. |
Dependent item | kube.statefulset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Current replicas | The number of current replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Ready replicas | The number of ready replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Updated replicas | The number of updated replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the number of replicas. |
Dependent item | kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet is down | (last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]) / last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}]))<>1 |
High | ||
| Kubernetes cluster state: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet replicas mismatch | StatefulSet has not matched the number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"statefulset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PodDisruptionBudget discovery | Dependent item | kube.pdb.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods healthy | Current number of healthy pods. |
Dependent item | kube.pdb.pods_healthy[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods desired | Minimum desired number of healthy pods. |
Dependent item | kube.pdb.pods_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Disruptions allowed | Number of pod disruptions that are allowed. |
Dependent item | kube.pdb.disruptions_allowed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods total | Total number of pods counted by this disruption budget. |
Dependent item | kube.pdb.pods_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| CronJob discovery | Dependent item | kube.cronjob.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Suspend | Suspend flag tells the controller to suspend subsequent executions. |
Dependent item | kube.cronjob.spec_suspend[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Active | Active holds pointers to currently running jobs. |
Dependent item | kube.cronjob.status_active[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Last schedule | LastScheduleTime keeps information of when was the last time the job was successfully scheduled. |
Dependent item | kube.cronjob.last_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Next schedule | Next time the cronjob should be scheduled. The time after lastScheduleTime or after the cron job's creation time if it's never been scheduled. Use this to determine if the job is delayed. |
Dependent item | kube.cronjob.next_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.cronjob.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.cronjob.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.cronjob.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.cronjob.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Job discovery | Dependent item | kube.job.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.job.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.job.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.job.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.job.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component statuses discovery | Dependent item | kube.componentstatuses.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component [{#NAME}]: Healthy | Cluster component healthy. |
Dependent item | kube.componentstatuses.healthy[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Component [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}],#2,"ne","True")=2 and length(last(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz discovery | Dependent item | kube.readyz.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz [{#NAME}]: Healthcheck | Result of readyz healthcheck for component. |
Dependent item | kube.readyz.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Readyz [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez discovery | Dependent item | kube.livez.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez [{#NAME}]: Healthcheck | Result of livez healthcheck for component. |
Dependent item | kube.livez.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Livez [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift BuildConfig discovery | Dependent item | openshift.buildconfig.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Created | OpenShift BuildConfig Unix creation timestamp. |
Dependent item | openshift.buildconfig.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.buildconfig.generation[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Latest version | The latest version of BuildConfig. |
Dependent item | openshift.buildconfig.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Build discovery | Dependent item | openshift.build.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Created | OpenShift Build Unix creation timestamp. |
Dependent item | openshift.build.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.build.sequence.number[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Status phase | The Build phase. |
Dependent item | openshift.build.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Build [{#NAME}]: Build has failed | count(/Kubernetes cluster state by HTTP/openshift.build.status_phase[{#NAMESPACE}/{#NAME}],2m,"ge",6)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift ClusterResourceQuota discovery | Dependent item | openshift.cluster.resource.quota.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Quota [{#NAME}] Resource [{#RESOURCE}]: Type [{#TYPE}]] | Usage about resource quota. |
Dependent item | openshift.cluster.resource.quota[{#RESOURCE}/{#NAME}/{#TYPE}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Route discovery | Dependent item | openshift.route.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Route [{#NAME}]: Created | OpenShift Route Unix creation timestamp. |
Dependent item | openshift.route.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Route [{#NAME}]: Status | Information about route status. |
Dependent item | openshift.route.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Route [{#NAME}] with issue: Status is false | count(/Kubernetes cluster state by HTTP/openshift.route.status[{#NAMESPACE}/{#NAME}],2m,,0)>=2 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes state. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Template Kubernetes cluster state by HTTP - collects metrics by HTTP agent from kube-state-metrics endpoint and Kubernetes API.
Don't forget to change macros {$KUBE.API.URL} and {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Zabbix version: 7.2 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster. Internal service metrics are collected from kube-state-metrics endpoint.
Template needs to use authorization via API token.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command:
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.STATE.ENDPOINT.NAME} with Kube state metrics endpoint name. See kubectl -n monitoring get ep. Default: zabbix-kube-state-metrics.
NOTE. If you wish to monitor Controller Manager and Scheduler components, you might need to set the --binding-address option for them to the address where Zabbix proxy can reach them.
For example, for clusters created with kubeadm it can be set in the following manifest files (changes will be applied immediately):
Depending on your Kubernetes distribution, you might need to adjust {$KUBE.CONTROL_PLANE.TAINT} macro (for example, set it to node-role.kubernetes.io/master for OpenShift).
NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Also, see the Macros section for a list of macros used to set trigger values.
Set up the macros to filter the metrics of discovered Kubelets by node names:
Set up macros to filter metrics by namespace:
Set up macros to filter node metrics by nodename:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable namespaces.
You can use the {$KUBE.KUBELET.FILTER.LABELS} and {$KUBE.KUBELET.FILTER.ANNOTATIONS} macros for advanced filtering of kubelets by node labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the kubelets on nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
You can also set up evaluation periods for replica mismatch triggers (Deployments, ReplicaSets, StatefulSets) with the macro {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD}, which supports context and regular expressions. For example, you can create the following macros:
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:default:nginx-deployment"} = #3
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"deployment:.*:.*"} = #10 or {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"^deployment.*"} = #10
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:".*:default:.*"} = 15m
Note that different context macros with regular expressions matching the same string can be applied in an undefined order, and simple context macros (without regular expressions) have higher priority. Read the Important notes section in Zabbix documentation for details.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.READYZ.ENDPOINT} | Kubernetes API readyz endpoint /readyz |
/readyz |
| {$KUBE.API.LIVEZ.ENDPOINT} | Kubernetes API livez endpoint /livez |
/livez |
| {$KUBE.API.COMPONENTSTATUSES.ENDPOINT} | Kubernetes API componentstatuses endpoint /api/v1/componentstatuses |
/api/v1/componentstatuses |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.STATE.ENDPOINT.NAME} | Kubernetes state endpoint name. |
zabbix-kube-state-metrics |
| {$OPENSHIFT.STATE.ENDPOINT.NAME} | OpenShift state endpoint name. |
openshift-state-metrics |
| {$KUBE.API_SERVER.SCHEME} | Kubernetes API servers metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.API_SERVER.PORT} | Kubernetes API servers metrics endpoint port. Used in ControlPlane LLD. |
6443 |
| {$KUBE.CONTROL_PLANE.TAINT} | Taint that applies to control plane nodes. Change if needed. Used in ControlPlane LLD. |
node-role.kubernetes.io/control-plane |
| {$KUBE.CONTROLLER_MANAGER.SCHEME} | Kubernetes Controller manager metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.CONTROLLER_MANAGER.PORT} | Kubernetes Controller manager metrics endpoint port. Used in ControlPlane LLD. |
10257 |
| {$KUBE.SCHEDULER.SCHEME} | Kubernetes Scheduler metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.SCHEDULER.PORT} | Kubernetes Scheduler metrics endpoint port. Used in ControlPlane LLD. |
10259 |
| {$KUBE.KUBELET.SCHEME} | Kubernetes Kubelet metrics endpoint scheme. Used in Kubelet LLD. |
https |
| {$KUBE.KUBELET.PORT} | Kubernetes Kubelet metrics endpoint port. Used in Kubelet LLD. |
10250 |
| {$KUBE.LLD.FILTER.NAMESPACE.MATCHES} | Filter of discoverable metrics by namespace. |
.* |
| {$KUBE.LLD.FILTER.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered metrics by namespace. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes by nodename. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.KUBELET_NODE.MATCHES} | Filter of discoverable Kubelets by nodename. |
.* |
| {$KUBE.LLD.FILTER.KUBELET_NODE.NOT_MATCHES} | Filter to exclude discovered Kubelets by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.KUBELET.FILTER.ANNOTATIONS} | Node annotations to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.KUBELET.FILTER.LABELS} | Node labels to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.PV.MATCHES} | Filter of discoverable persistent volumes by name. |
.* |
| {$KUBE.LLD.FILTER.PV.NOT_MATCHES} | Filter to exclude discovered persistent volumes by name. |
CHANGE_IF_NEEDED |
| {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD} | The evaluation period range which is used for calculation of expressions in trigger prototypes (time period or value range). Can be used with context. |
#5 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get state metrics | Collecting Kubernetes metrics from kube-state-metrics. |
Script | kube.state.metrics |
| Control plane LLD | Generation of data for Control plane discovery rules. |
Script | kube.control_plane.lld Preprocessing
|
| Node LLD | Generation of data for Kubelet discovery rules. |
Script | kube.node.lld Preprocessing
|
| Get component statuses | HTTP agent | kube.componentstatuses Preprocessing
|
|
| Get readyz | HTTP agent | kube.readyz Preprocessing
|
|
| Get livez | HTTP agent | kube.livez Preprocessing
|
|
| Namespace count | The number of namespaces. |
Dependent item | kube.namespace.count Preprocessing
|
| CronJob count | Number of cronjobs. |
Dependent item | kube.cronjob.count Preprocessing
|
| Job count | Number of jobs (generated by cronjob + job). |
Dependent item | kube.job.count Preprocessing
|
| Endpoint count | Number of endpoints. |
Dependent item | kube.endpoint.count Preprocessing
|
| Deployment count | The number of deployments. |
Dependent item | kube.deployment.count Preprocessing
|
| Service count | The number of services. |
Dependent item | kube.service.count Preprocessing
|
| StatefulSet count | The number of statefulsets. |
Dependent item | kube.statefulset.count Preprocessing
|
| Node count | The number of nodes. |
Dependent item | kube.node.count Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| API servers discovery | Dependent item | kube.api_servers.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Controller manager nodes discovery | Dependent item | kube.controller_manager.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduler servers nodes discovery | Dependent item | kube.scheduler.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubelet discovery | Dependent item | kube.kubelet.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Daemonset discovery | Dependent item | kube.daemonset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Ready | The number of nodes that should be running the daemon pod and have one or more running and ready. |
Dependent item | kube.daemonset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Scheduled | The number of nodes that run at least one daemon pod and are supposed to. |
Dependent item | kube.daemonset.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Desired | The number of nodes that should be running the daemon pod. |
Dependent item | kube.daemonset.desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Misscheduled | The number of nodes that run a daemon pod but are not supposed to. |
Dependent item | kube.daemonset.misscheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Updated number scheduled | The total number of nodes that are running updated daemon pod. |
Dependent item | kube.daemonset.updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PVC discovery | Dependent item | kube.pvc.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase | The current status phase of the persistent volume claim. |
Dependent item | kube.pvc.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC [{#NAME}] Requested storage | The capacity of storage requested by the persistent volume claim. |
Dependent item | kube.pvc.requested.storage[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Bound, sum | The total amount of persistent volume claims in the Bound phase. |
Dependent item | kube.pvc.status_phase.bound.sum[{#NAMESPACE}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Lost, sum | The total amount of persistent volume claims in the Lost phase. |
Dependent item | kube.pvc.status_phase.lost.sum[{#NAMESPACE}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Pending, sum | The total amount of persistent volume claims in the Pending phase. |
Dependent item | kube.pvc.status_phase.pending.sum[{#NAMESPACE}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: NS [{#NAMESPACE}] PVC [{#NAME}]: PVC is pending | count(/Kubernetes cluster state by HTTP/kube.pvc.status_phase[{#NAMESPACE}/{#NAME}],2m,,5)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV discovery | Dependent item | kube.pv.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV [{#NAME}] Status phase | The current status phase of the persistent volume. |
Dependent item | kube.pv.status_phase[{#NAME}] Preprocessing
|
| PV [{#NAME}] Capacity bytes | A capacity of the persistent volume in bytes. |
Dependent item | kube.pv.capacity.bytes[{#NAME}] Preprocessing
|
| PV status phase: Pending, sum | The total amount of persistent volumes in the Pending phase. |
Dependent item | kube.pv.status_phase.pending.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Available, sum | The total amount of persistent volumes in the Available phase. |
Dependent item | kube.pv.status_phase.available.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Bound, sum | The total amount of persistent volumes in the Bound phase. |
Dependent item | kube.pv.status_phase.bound.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Released, sum | The total amount of persistent volumes in the Released phase. |
Dependent item | kube.pv.status_phase.released.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Failed, sum | The total amount of persistent volumes in the Failed phase. |
Dependent item | kube.pv.status_phase.failed.sum[{#SINGLETON}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: PV [{#NAME}]: PV has failed | count(/Kubernetes cluster state by HTTP/kube.pv.status_phase[{#NAME}],2m,,3)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Deployment discovery | Dependent item | kube.deployment.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Paused | Whether the deployment is paused and will not be processed by the deployment controller. |
Dependent item | kube.deployment.spec_paused[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas desired | Number of desired pods for a deployment. |
Dependent item | kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Rollingupdate max unavailable | Maximum number of unavailable replicas during a rolling update of a deployment. |
Dependent item | kube.deployment.rollingupdate.max_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas | The number of replicas per deployment. |
Dependent item | kube.deployment.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas available | The number of available replicas per deployment. |
Dependent item | kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas unavailable | The number of unavailable replicas per deployment. |
Dependent item | kube.deployment.replicas_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas updated | The number of updated replicas per deployment. |
Dependent item | kube.deployment.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas mismatched | The number of available replicas not matching the desired number of replicas. |
Dependent item | kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Deployment replicas mismatch | Deployment has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Endpoint discovery | Dependent item | kube.endpoint.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address available | Number of addresses available in endpoint. |
Dependent item | kube.endpoint.address_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address not ready | Number of addresses not ready in endpoint. |
Dependent item | kube.endpoint.address_not_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Age | Endpoint age (number of seconds since creation). |
Dependent item | kube.endpoint.age[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: CPU allocatable | The CPU resources of a node that are available for scheduling. |
Dependent item | kube.node.cpu_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Memory allocatable | The memory resources of a node that are available for scheduling. |
Dependent item | kube.node.memory_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Pods allocatable | The pods resources of a node that are available for scheduling. |
Dependent item | kube.node.pods_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Ephemeral storage allocatable | The allocatable ephemeral storage of a node that is available for scheduling. |
Dependent item | kube.node.ephemeral_storage_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: CPU capacity | The capacity for CPU resources of a node. |
Dependent item | kube.node.cpu_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Memory capacity | The capacity for memory resources of a node. |
Dependent item | kube.node.memory_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Ephemeral storage capacity | The ephemeral storage capacity of a node. |
Dependent item | kube.node.ephemeral_storage_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Pods capacity | The capacity for pods resources of a node. |
Dependent item | kube.node.pods_capacity[{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Pending | Pod is in pending state. |
Dependent item | kube.pod.phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Succeeded | Pod is in succeeded state. |
Dependent item | kube.pod.phase.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Failed | Pod is in failed state. |
Dependent item | kube.pod.phase.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Unknown | Pod is in unknown state. |
Dependent item | kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Running | Pod is in unknown state. |
Dependent item | kube.pod.phase.running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers terminated | Describes whether the container is currently in terminated state. |
Dependent item | kube.pod.containers_terminated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers waiting | Describes whether the container is currently in waiting state. |
Dependent item | kube.pod.containers_waiting[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers ready | Describes whether the containers readiness check succeeded. |
Dependent item | kube.pod.containers_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers restarts | The number of container restarts. |
Dependent item | kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers running | Describes whether the container is currently in running state. |
Dependent item | kube.pod.containers_running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Ready | Describes whether the pod is ready to serve requests. |
Dependent item | kube.pod.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Scheduled | Describes the status of the scheduling process for the pod. |
Dependent item | kube.pod.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Unschedulable | Describes the unschedulable status for the pod. |
Dependent item | kube.pod.unschedulable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU limits | The limit on CPU cores to be used by a container. |
Dependent item | kube.pod.containers.limits.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory limits | The limit on memory to be used by a container. |
Dependent item | kube.pod.containers.limits.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU requests | The number of requested CPU cores by a container. |
Dependent item | kube.pod.containers.requests.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory requests | The number of requested memory bytes by a container. |
Dependent item | kube.pod.containers.requests.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is not healthy | min(/Kubernetes cluster state by HTTP/kube.pod.phase.failed[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.pending[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}],10m)>0 |
High | ||
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}])-min(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}],15m))>1 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ReplicaSet discovery | Dependent item | kube.replicaset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas | The number of replicas per ReplicaSet. |
Dependent item | kube.replicaset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Desired replicas | Number of desired pods for a ReplicaSet. |
Dependent item | kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Fully labeled replicas | The number of fully labeled replicas per ReplicaSet. |
Dependent item | kube.replicaset.fully_labeled_replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Ready | The number of ready replicas per ReplicaSet. |
Dependent item | kube.replicaset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the desired number of replicas. |
Dependent item | kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] RS [{#NAME}]: ReplicaSet mismatch | ReplicaSet has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"replicaset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| StatefulSet discovery | Dependent item | kube.statefulset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas | The number of replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Desired replicas | Number of desired pods for a StatefulSet. |
Dependent item | kube.statefulset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Current replicas | The number of current replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Ready replicas | The number of ready replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Updated replicas | The number of updated replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the number of replicas. |
Dependent item | kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet is down | (last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]) / last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}]))<>1 |
High | ||
| Kubernetes cluster state: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet replicas mismatch | StatefulSet has not matched the number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"statefulset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PodDisruptionBudget discovery | Dependent item | kube.pdb.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods healthy | Current number of healthy pods. |
Dependent item | kube.pdb.pods_healthy[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods desired | Minimum desired number of healthy pods. |
Dependent item | kube.pdb.pods_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Disruptions allowed | Number of pod disruptions that are allowed. |
Dependent item | kube.pdb.disruptions_allowed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods total | Total number of pods counted by this disruption budget. |
Dependent item | kube.pdb.pods_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| CronJob discovery | Dependent item | kube.cronjob.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Suspend | Suspend flag tells the controller to suspend subsequent executions. |
Dependent item | kube.cronjob.spec_suspend[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Active | Active holds pointers to currently running jobs. |
Dependent item | kube.cronjob.status_active[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Last schedule | LastScheduleTime keeps information of when was the last time the job was successfully scheduled. |
Dependent item | kube.cronjob.last_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Next schedule | Next time the cronjob should be scheduled. The time after lastScheduleTime or after the cron job's creation time if it's never been scheduled. Use this to determine if the job is delayed. |
Dependent item | kube.cronjob.next_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.cronjob.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.cronjob.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.cronjob.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.cronjob.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Job discovery | Dependent item | kube.job.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.job.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.job.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.job.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.job.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component statuses discovery | Dependent item | kube.componentstatuses.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component [{#NAME}]: Healthy | Cluster component healthy. |
Dependent item | kube.componentstatuses.healthy[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Component [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}],#2,"ne","True")=2 and length(last(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz discovery | Dependent item | kube.readyz.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz [{#NAME}]: Healthcheck | Result of readyz healthcheck for component. |
Dependent item | kube.readyz.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Readyz [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez discovery | Dependent item | kube.livez.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez [{#NAME}]: Healthcheck | Result of livez healthcheck for component. |
Dependent item | kube.livez.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Livez [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift BuildConfig discovery | Dependent item | openshift.buildconfig.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Created | OpenShift BuildConfig Unix creation timestamp. |
Dependent item | openshift.buildconfig.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.buildconfig.generation[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Latest version | The latest version of BuildConfig. |
Dependent item | openshift.buildconfig.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Build discovery | Dependent item | openshift.build.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Created | OpenShift Build Unix creation timestamp. |
Dependent item | openshift.build.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.build.sequence.number[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Status phase | The Build phase. |
Dependent item | openshift.build.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Build [{#NAME}]: Build has failed | count(/Kubernetes cluster state by HTTP/openshift.build.status_phase[{#NAMESPACE}/{#NAME}],2m,"ge",6)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift ClusterResourceQuota discovery | Dependent item | openshift.cluster.resource.quota.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Quota [{#NAME}] Resource [{#RESOURCE}]: Type [{#TYPE}]] | Usage about resource quota. |
Dependent item | openshift.cluster.resource.quota[{#RESOURCE}/{#NAME}/{#TYPE}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Route discovery | Dependent item | openshift.route.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Route [{#NAME}]: Created | OpenShift Route Unix creation timestamp. |
Dependent item | openshift.route.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Route [{#NAME}]: Status | Information about route status. |
Dependent item | openshift.route.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Route [{#NAME}] with issue: Status is false | count(/Kubernetes cluster state by HTTP/openshift.route.status[{#NAMESPACE}/{#NAME}],2m,,0)>=2 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes state. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Template Kubernetes cluster state by HTTP - collects metrics by HTTP agent from kube-state-metrics endpoint and Kubernetes API.
Don't forget to change macros {$KUBE.API.URL} and {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
Note: Some metrics may not be collected depending on your Kubernetes version and configuration.
Zabbix version: 7.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster. Internal service metrics are collected from kube-state-metrics endpoint.
Template needs to use authorization via API token.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the service account name. If a different release name is used.
kubectl get serviceaccounts -n monitoring
Get the generated service account token using the command:
kubectl get secret zabbix-zabbix-helm-chart -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.STATE.ENDPOINT.NAME} with Kube state metrics endpoint name. See kubectl -n monitoring get ep. Default: zabbix-kube-state-metrics.
Note: If you wish to monitor Controller Manager and Scheduler components, you might need to set the --binding-address option for them to the address where Zabbix proxy can reach them.
For example, for clusters created with kubeadm it can be set in the following manifest files (changes will be applied immediately):
Depending on your Kubernetes distribution, you might need to adjust {$KUBE.CONTROL_PLANE.TAINT} macro (for example, set it to node-role.kubernetes.io/master for OpenShift).
Note: Some metrics may not be collected depending on your Kubernetes version and configuration.
Also, see the Macros section for a list of macros used to set trigger values.
Set up the macros to filter the metrics of discovered Kubelets by node names:
Set up macros to filter metrics by namespace:
Set up macros to filter node metrics by nodename:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable namespaces.
You can use the {$KUBE.KUBELET.FILTER.LABELS} and {$KUBE.KUBELET.FILTER.ANNOTATIONS} macros for advanced filtering of kubelets by node labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the kubelets on nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
You can also set up evaluation periods for replica mismatch triggers (Deployments, ReplicaSets, StatefulSets) with the macro {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD}, which supports context and regular expressions. For example, you can create the following macros:
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:default:nginx-deployment"} = #3
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"deployment:.*:.*"} = #10 or {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"^deployment.*"} = #10
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:".*:default:.*"} = 15m
Note: that different context macros with regular expressions matching the same string can be applied in an undefined order, and simple context macros (without regular expressions) have higher priority. Read the Important notes section in Zabbix documentation for details.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.READYZ.ENDPOINT} | Kubernetes API readyz endpoint /readyz |
/readyz |
| {$KUBE.API.LIVEZ.ENDPOINT} | Kubernetes API livez endpoint /livez |
/livez |
| {$KUBE.API.COMPONENTSTATUSES.ENDPOINT} | Kubernetes API componentstatuses endpoint /api/v1/componentstatuses |
/api/v1/componentstatuses |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.STATE.ENDPOINT.NAME} | Endpoint name for kube-state-metrics service (check with |
zabbix-kube-state-metrics |
| {$OPENSHIFT.STATE.ENDPOINT.NAME} | OpenShift state endpoint name. |
openshift-state-metrics |
| {$KUBE.API_SERVER.SCHEME} | Kubernetes API servers metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.API_SERVER.PORT} | Kubernetes API servers metrics endpoint port. Used in ControlPlane LLD. |
6443 |
| {$KUBE.CONTROL_PLANE.TAINT} | Taint that applies to control plane nodes. Change if needed. Used in ControlPlane LLD. |
node-role.kubernetes.io/control-plane |
| {$KUBE.CONTROLLER_MANAGER.SCHEME} | Kubernetes Controller manager metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.CONTROLLER_MANAGER.PORT} | Kubernetes Controller manager metrics endpoint port. Used in ControlPlane LLD. |
10257 |
| {$KUBE.SCHEDULER.SCHEME} | Kubernetes Scheduler metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.SCHEDULER.PORT} | Kubernetes Scheduler metrics endpoint port. Used in ControlPlane LLD. |
10259 |
| {$KUBE.KUBELET.SCHEME} | Kubernetes Kubelet metrics endpoint scheme. Used in Kubelet LLD. |
https |
| {$KUBE.KUBELET.PORT} | Kubernetes Kubelet metrics endpoint port. Used in Kubelet LLD. |
10250 |
| {$KUBE.LLD.FILTER.NAMESPACE.MATCHES} | Filter of discoverable metrics by namespace. |
.* |
| {$KUBE.LLD.FILTER.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered metrics by namespace. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes by nodename. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.KUBELET_NODE.MATCHES} | Filter of discoverable Kubelets by nodename. |
.* |
| {$KUBE.LLD.FILTER.KUBELET_NODE.NOT_MATCHES} | Filter to exclude discovered Kubelets by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.KUBELET.FILTER.ANNOTATIONS} | Node annotations to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.KUBELET.FILTER.LABELS} | Node labels to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.PV.MATCHES} | Filter of discoverable persistent volumes by name. |
.* |
| {$KUBE.LLD.FILTER.PV.NOT_MATCHES} | Filter to exclude discovered persistent volumes by name. |
CHANGE_IF_NEEDED |
| {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD} | The evaluation period range which is used for calculation of expressions in trigger prototypes (time period or value range). Can be used with context. |
#5 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Get state metrics | Collecting Kubernetes metrics from kube-state-metrics. |
Script | kube.state.metrics |
| Control plane LLD | Generation of data for Control plane discovery rules. |
Script | kube.control_plane.lld Preprocessing
|
| Node LLD | Generation of data for Kubelet discovery rules. |
Script | kube.node.lld Preprocessing
|
| Get component statuses | HTTP agent | kube.componentstatuses Preprocessing
|
|
| Get readyz | HTTP agent | kube.readyz Preprocessing
|
|
| Get livez | HTTP agent | kube.livez Preprocessing
|
|
| Namespace count | The number of namespaces. |
Dependent item | kube.namespace.count Preprocessing
|
| CronJob count | Number of cronjobs. |
Dependent item | kube.cronjob.count Preprocessing
|
| Job count | Number of jobs (generated by cronjob + job). |
Dependent item | kube.job.count Preprocessing
|
| Endpoint count | Number of endpoints. |
Dependent item | kube.endpoint.count Preprocessing
|
| Deployment count | The number of deployments. |
Dependent item | kube.deployment.count Preprocessing
|
| Service count | The number of services. |
Dependent item | kube.service.count Preprocessing
|
| StatefulSet count | The number of statefulsets. |
Dependent item | kube.statefulset.count Preprocessing
|
| Node count | The number of nodes. |
Dependent item | kube.node.count Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| API servers discovery | Dependent item | kube.api_servers.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Controller manager nodes discovery | Dependent item | kube.controller_manager.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduler servers nodes discovery | Dependent item | kube.scheduler.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubelet discovery | Dependent item | kube.kubelet.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Daemonset discovery | Dependent item | kube.daemonset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Ready | The number of nodes that should be running the daemon pod and have one or more running and ready. |
Dependent item | kube.daemonset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Scheduled | The number of nodes that run at least one daemon pod and are supposed to. |
Dependent item | kube.daemonset.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Desired | The number of nodes that should be running the daemon pod. |
Dependent item | kube.daemonset.desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Misscheduled | The number of nodes that run a daemon pod but are not supposed to. |
Dependent item | kube.daemonset.misscheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Updated number scheduled | The total number of nodes that are running updated daemon pod. |
Dependent item | kube.daemonset.updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PVC discovery | Dependent item | kube.pvc.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase | The current status phase of the persistent volume claim. |
Dependent item | kube.pvc.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC [{#NAME}] Requested storage | The capacity of storage requested by the persistent volume claim. |
Dependent item | kube.pvc.requested.storage[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Bound, sum | The total amount of persistent volume claims in the Bound phase. |
Dependent item | kube.pvc.status_phase.bound.sum[{#NAMESPACE}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Lost, sum | The total amount of persistent volume claims in the Lost phase. |
Dependent item | kube.pvc.status_phase.lost.sum[{#NAMESPACE}] Preprocessing
|
| Namespace [{#NAMESPACE}] PVC status phase: Pending, sum | The total amount of persistent volume claims in the Pending phase. |
Dependent item | kube.pvc.status_phase.pending.sum[{#NAMESPACE}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: NS [{#NAMESPACE}] PVC [{#NAME}]: PVC is pending | count(/Kubernetes cluster state by HTTP/kube.pvc.status_phase[{#NAMESPACE}/{#NAME}],2m,,5)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV discovery | Dependent item | kube.pv.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV [{#NAME}] Status phase | The current status phase of the persistent volume. |
Dependent item | kube.pv.status_phase[{#NAME}] Preprocessing
|
| PV [{#NAME}] Capacity bytes | A capacity of the persistent volume in bytes. |
Dependent item | kube.pv.capacity.bytes[{#NAME}] Preprocessing
|
| PV status phase: Pending, sum | The total amount of persistent volumes in the Pending phase. |
Dependent item | kube.pv.status_phase.pending.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Available, sum | The total amount of persistent volumes in the Available phase. |
Dependent item | kube.pv.status_phase.available.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Bound, sum | The total amount of persistent volumes in the Bound phase. |
Dependent item | kube.pv.status_phase.bound.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Released, sum | The total amount of persistent volumes in the Released phase. |
Dependent item | kube.pv.status_phase.released.sum[{#SINGLETON}] Preprocessing
|
| PV status phase: Failed, sum | The total amount of persistent volumes in the Failed phase. |
Dependent item | kube.pv.status_phase.failed.sum[{#SINGLETON}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: PV [{#NAME}]: PV has failed | count(/Kubernetes cluster state by HTTP/kube.pv.status_phase[{#NAME}],2m,,3)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Deployment discovery | Dependent item | kube.deployment.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Paused | Whether the deployment is paused and will not be processed by the deployment controller. |
Dependent item | kube.deployment.spec_paused[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas desired | Number of desired pods for a deployment. |
Dependent item | kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Rollingupdate max unavailable | Maximum number of unavailable replicas during a rolling update of a deployment. |
Dependent item | kube.deployment.rollingupdate.max_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas | The number of replicas per deployment. |
Dependent item | kube.deployment.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas available | The number of available replicas per deployment. |
Dependent item | kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas unavailable | The number of unavailable replicas per deployment. |
Dependent item | kube.deployment.replicas_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas updated | The number of updated replicas per deployment. |
Dependent item | kube.deployment.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas mismatched | The number of available replicas not matching the desired number of replicas. |
Dependent item | kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Deployment replicas mismatch | Deployment has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Endpoint discovery | Dependent item | kube.endpoint.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address available | Number of addresses available in endpoint. |
Dependent item | kube.endpoint.address_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address not ready | Number of addresses not ready in endpoint. |
Dependent item | kube.endpoint.address_not_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Age | Endpoint age (number of seconds since creation). |
Dependent item | kube.endpoint.age[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node [{#NAME}]: CPU allocatable | The CPU resources of a node that are available for scheduling. |
Dependent item | kube.node.cpu_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Memory allocatable | The memory resources of a node that are available for scheduling. |
Dependent item | kube.node.memory_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Pods allocatable | The pods resources of a node that are available for scheduling. |
Dependent item | kube.node.pods_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Ephemeral storage allocatable | The allocatable ephemeral storage of a node that is available for scheduling. |
Dependent item | kube.node.ephemeral_storage_allocatable[{#NAME}] Preprocessing
|
| Node [{#NAME}]: CPU capacity | The capacity for CPU resources of a node. |
Dependent item | kube.node.cpu_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Memory capacity | The capacity for memory resources of a node. |
Dependent item | kube.node.memory_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Ephemeral storage capacity | The ephemeral storage capacity of a node. |
Dependent item | kube.node.ephemeral_storage_capacity[{#NAME}] Preprocessing
|
| Node [{#NAME}]: Pods capacity | The capacity for pods resources of a node. |
Dependent item | kube.node.pods_capacity[{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Pending | Pod is in pending state. |
Dependent item | kube.pod.phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Succeeded | Pod is in succeeded state. |
Dependent item | kube.pod.phase.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Failed | Pod is in failed state. |
Dependent item | kube.pod.phase.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Unknown | Pod is in unknown state. |
Dependent item | kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Running | Pod is in unknown state. |
Dependent item | kube.pod.phase.running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers terminated | Describes whether the container is currently in terminated state. |
Dependent item | kube.pod.containers_terminated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers waiting | Describes whether the container is currently in waiting state. |
Dependent item | kube.pod.containers_waiting[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers ready | Describes whether the containers readiness check succeeded. |
Dependent item | kube.pod.containers_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers restarts | The number of container restarts. |
Dependent item | kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers running | Describes whether the container is currently in running state. |
Dependent item | kube.pod.containers_running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Ready | Describes whether the pod is ready to serve requests. |
Dependent item | kube.pod.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Scheduled | Describes the status of the scheduling process for the pod. |
Dependent item | kube.pod.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Unschedulable | Describes the unschedulable status for the pod. |
Dependent item | kube.pod.unschedulable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU limits | The limit on CPU cores to be used by a container. |
Dependent item | kube.pod.containers.limits.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory limits | The limit on memory to be used by a container. |
Dependent item | kube.pod.containers.limits.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU requests | The number of requested CPU cores by a container. |
Dependent item | kube.pod.containers.requests.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory requests | The number of requested memory bytes by a container. |
Dependent item | kube.pod.containers.requests.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is not healthy | min(/Kubernetes cluster state by HTTP/kube.pod.phase.failed[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.pending[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}],10m)>0 |
High | ||
| Kubernetes cluster state: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}])-min(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}],15m))>1 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ReplicaSet discovery | Dependent item | kube.replicaset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas | The number of replicas per ReplicaSet. |
Dependent item | kube.replicaset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Desired replicas | Number of desired pods for a ReplicaSet. |
Dependent item | kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Fully labeled replicas | The number of fully labeled replicas per ReplicaSet. |
Dependent item | kube.replicaset.fully_labeled_replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Ready | The number of ready replicas per ReplicaSet. |
Dependent item | kube.replicaset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the desired number of replicas. |
Dependent item | kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] RS [{#NAME}]: ReplicaSet mismatch | ReplicaSet has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"replicaset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| StatefulSet discovery | Dependent item | kube.statefulset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas | The number of replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Desired replicas | Number of desired pods for a StatefulSet. |
Dependent item | kube.statefulset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Current replicas | The number of current replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Ready replicas | The number of ready replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Updated replicas | The number of updated replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the number of replicas. |
Dependent item | kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet is down | (last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]) / last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}]))<>1 |
High | ||
| Kubernetes cluster state: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet replicas mismatch | StatefulSet has not matched the number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"statefulset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PodDisruptionBudget discovery | Dependent item | kube.pdb.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods healthy | Current number of healthy pods. |
Dependent item | kube.pdb.pods_healthy[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods desired | Minimum desired number of healthy pods. |
Dependent item | kube.pdb.pods_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Disruptions allowed | Number of pod disruptions that are allowed. |
Dependent item | kube.pdb.disruptions_allowed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods total | Total number of pods counted by this disruption budget. |
Dependent item | kube.pdb.pods_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| CronJob discovery | Dependent item | kube.cronjob.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Suspend | Suspend flag tells the controller to suspend subsequent executions. |
Dependent item | kube.cronjob.spec_suspend[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Active | Active holds pointers to currently running jobs. |
Dependent item | kube.cronjob.status_active[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Last schedule | LastScheduleTime keeps information of when was the last time the job was successfully scheduled. |
Dependent item | kube.cronjob.last_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Next schedule | Next time the cronjob should be scheduled. The time after lastScheduleTime or after the cron job's creation time if it's never been scheduled. Use this to determine if the job is delayed. |
Dependent item | kube.cronjob.next_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.cronjob.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.cronjob.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.cronjob.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.cronjob.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Job discovery | Dependent item | kube.job.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.job.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.job.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.job.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.job.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component statuses discovery | Dependent item | kube.componentstatuses.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component [{#NAME}]: Healthy | Cluster component healthy. |
Dependent item | kube.componentstatuses.healthy[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Component [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}],#2,"ne","True")=2 and length(last(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz discovery | Dependent item | kube.readyz.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz [{#NAME}]: Healthcheck | Result of readyz healthcheck for component. |
Dependent item | kube.readyz.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Readyz [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez discovery | Dependent item | kube.livez.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez [{#NAME}]: Healthcheck | Result of livez healthcheck for component. |
Dependent item | kube.livez.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Livez [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift BuildConfig discovery | Dependent item | openshift.buildconfig.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Created | OpenShift BuildConfig Unix creation timestamp. |
Dependent item | openshift.buildconfig.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.buildconfig.generation[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Latest version | The latest version of BuildConfig. |
Dependent item | openshift.buildconfig.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Build discovery | Dependent item | openshift.build.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Created | OpenShift Build Unix creation timestamp. |
Dependent item | openshift.build.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.build.sequence.number[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Build [{#NAME}]: Status phase | The Build phase. |
Dependent item | openshift.build.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Build [{#NAME}]: Build has failed | count(/Kubernetes cluster state by HTTP/openshift.build.status_phase[{#NAMESPACE}/{#NAME}],2m,"ge",6)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift ClusterResourceQuota discovery | Dependent item | openshift.cluster.resource.quota.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Quota [{#NAME}] Resource [{#RESOURCE}]: Type [{#TYPE}]] | Usage about resource quota. |
Dependent item | openshift.cluster.resource.quota[{#RESOURCE}/{#NAME}/{#TYPE}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Route discovery | Dependent item | openshift.route.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Namespace [{#NAMESPACE}] Route [{#NAME}]: Created | OpenShift Route Unix creation timestamp. |
Dependent item | openshift.route.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Namespace [{#NAMESPACE}] Route [{#NAME}]: Status | Information about route status. |
Dependent item | openshift.route.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes cluster state: Route [{#NAME}] with issue: Status is false | count(/Kubernetes cluster state by HTTP/openshift.route.status[{#NAMESPACE}/{#NAME}],2m,,0)>=2 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
The template to monitor Kubernetes state. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Template Kubernetes cluster state by HTTP - collects metrics by HTTP agent from kube-state-metrics endpoint and Kubernetes API.
Don't forget to change macros {$KUBE.API.URL} and {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Zabbix version: 6.4 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster. Internal service metrics are collected from kube-state-metrics endpoint.
Template needs to use authorization via API token.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command:
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.STATE.ENDPOINT.NAME} with Kube state metrics endpoint name. See kubectl -n monitoring get ep. Default: zabbix-kube-state-metrics.
NOTE. If you wish to monitor Controller Manager and Scheduler components, you might need to set the --binding-address option for them to the address where Zabbix proxy can reach them.
For example, for clusters created with kubeadm it can be set in the following manifest files (changes will be applied immediately):
Depending on your Kubernetes distribution, you might need to adjust {$KUBE.CONTROL_PLANE.TAINT} macro (for example, set it to node-role.kubernetes.io/master for OpenShift).
NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Also, see the Macros section for a list of macros used to set trigger values.
Set up the macros to filter the metrics of discovered Kubelets by node names:
Set up macros to filter metrics by namespace:
Set up macros to filter node metrics by nodename:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable namespaces.
You can use the {$KUBE.KUBELET.FILTER.LABELS} and {$KUBE.KUBELET.FILTER.ANNOTATIONS} macros for advanced filtering of kubelets by node labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the kubelets on nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
You can also set up evaluation periods for replica mismatch triggers (Deployments, ReplicaSets, StatefulSets) with the macro {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD}, which supports context and regular expressions. For example, you can create the following macros:
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:default:nginx-deployment"} = #3
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"deployment:.*:.*"} = #10 or {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"^deployment.*"} = #10
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:".*:default:.*"} = 15m
Note that different context macros with regular expressions matching the same string can be applied in an undefined order, and simple context macros (without regular expressions) have higher priority. Read the Important notes section in Zabbix documentation for details.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.READYZ.ENDPOINT} | Kubernetes API readyz endpoint /readyz |
/readyz |
| {$KUBE.API.LIVEZ.ENDPOINT} | Kubernetes API livez endpoint /livez |
/livez |
| {$KUBE.API.COMPONENTSTATUSES.ENDPOINT} | Kubernetes API componentstatuses endpoint /api/v1/componentstatuses |
/api/v1/componentstatuses |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.STATE.ENDPOINT.NAME} | Kubernetes state endpoint name. |
zabbix-kube-state-metrics |
| {$OPENSHIFT.STATE.ENDPOINT.NAME} | OpenShift state endpoint name. |
openshift-state-metrics |
| {$KUBE.API_SERVER.SCHEME} | Kubernetes API servers metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.API_SERVER.PORT} | Kubernetes API servers metrics endpoint port. Used in ControlPlane LLD. |
6443 |
| {$KUBE.CONTROL_PLANE.TAINT} | Taint that applies to control plane nodes. Change if needed. Used in ControlPlane LLD. |
node-role.kubernetes.io/control-plane |
| {$KUBE.CONTROLLER_MANAGER.SCHEME} | Kubernetes Controller manager metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.CONTROLLER_MANAGER.PORT} | Kubernetes Controller manager metrics endpoint port. Used in ControlPlane LLD. |
10257 |
| {$KUBE.SCHEDULER.SCHEME} | Kubernetes Scheduler metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.SCHEDULER.PORT} | Kubernetes Scheduler metrics endpoint port. Used in ControlPlane LLD. |
10259 |
| {$KUBE.KUBELET.SCHEME} | Kubernetes Kubelet metrics endpoint scheme. Used in Kubelet LLD. |
https |
| {$KUBE.KUBELET.PORT} | Kubernetes Kubelet metrics endpoint port. Used in Kubelet LLD. |
10250 |
| {$KUBE.LLD.FILTER.NAMESPACE.MATCHES} | Filter of discoverable metrics by namespace. |
.* |
| {$KUBE.LLD.FILTER.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered metrics by namespace. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes by nodename. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.KUBELET_NODE.MATCHES} | Filter of discoverable Kubelets by nodename. |
.* |
| {$KUBE.LLD.FILTER.KUBELET_NODE.NOT_MATCHES} | Filter to exclude discovered Kubelets by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.KUBELET.FILTER.ANNOTATIONS} | Node annotations to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.KUBELET.FILTER.LABELS} | Node labels to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.PV.MATCHES} | Filter of discoverable persistent volumes by name. |
.* |
| {$KUBE.LLD.FILTER.PV.NOT_MATCHES} | Filter to exclude discovered persistent volumes by name. |
CHANGE_IF_NEEDED |
| {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD} | The evaluation period range which is used for calculation of expressions in trigger prototypes (time period or value range). Can be used with context. |
#5 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Get state metrics | Collecting Kubernetes metrics from kube-state-metrics. |
Script | kube.state.metrics |
| Kubernetes: Control plane LLD | Generation of data for Control plane discovery rules. |
Script | kube.control_plane.lld Preprocessing
|
| Kubernetes: Node LLD | Generation of data for Kubelet discovery rules. |
Script | kube.node.lld Preprocessing
|
| Kubernetes: Get component statuses | HTTP agent | kube.componentstatuses Preprocessing
|
|
| Kubernetes: Get readyz | HTTP agent | kube.readyz Preprocessing
|
|
| Kubernetes: Get livez | HTTP agent | kube.livez Preprocessing
|
|
| Kubernetes: Namespace count | The number of namespaces. |
Dependent item | kube.namespace.count Preprocessing
|
| Kubernetes: CronJob count | Number of cronjobs. |
Dependent item | kube.cronjob.count Preprocessing
|
| Kubernetes: Job count | Number of jobs (generated by cronjob + job). |
Dependent item | kube.job.count Preprocessing
|
| Kubernetes: Endpoint count | Number of endpoints. |
Dependent item | kube.endpoint.count Preprocessing
|
| Kubernetes: Deployment count | The number of deployments. |
Dependent item | kube.deployment.count Preprocessing
|
| Kubernetes: Service count | The number of services. |
Dependent item | kube.service.count Preprocessing
|
| Kubernetes: StatefulSet count | The number of statefulsets. |
Dependent item | kube.statefulset.count Preprocessing
|
| Kubernetes: Node count | The number of nodes. |
Dependent item | kube.node.count Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| API servers discovery | Dependent item | kube.api_servers.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Controller manager nodes discovery | Dependent item | kube.controller_manager.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduler servers nodes discovery | Dependent item | kube.scheduler.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubelet discovery | Dependent item | kube.kubelet.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Daemonset discovery | Dependent item | kube.daemonset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Ready | The number of nodes that should be running the daemon pod and have one or more running and ready. |
Dependent item | kube.daemonset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Scheduled | The number of nodes that run at least one daemon pod and are supposed to. |
Dependent item | kube.daemonset.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Desired | The number of nodes that should be running the daemon pod. |
Dependent item | kube.daemonset.desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Misscheduled | The number of nodes that run a daemon pod but are not supposed to. |
Dependent item | kube.daemonset.misscheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Updated number scheduled | The total number of nodes that are running updated daemon pod. |
Dependent item | kube.daemonset.updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PVC discovery | Dependent item | kube.pvc.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase | The current status phase of the persistent volume claim. |
Dependent item | kube.pvc.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Requested storage | The capacity of storage requested by the persistent volume claim. |
Dependent item | kube.pvc.requested.storage[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC status phase: Bound, sum | The total amount of persistent volume claims in the Bound phase. |
Dependent item | kube.pvc.status_phase.bound.sum[{#NAMESPACE}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC status phase: Lost, sum | The total amount of persistent volume claims in the Lost phase. |
Dependent item | kube.pvc.status_phase.lost.sum[{#NAMESPACE}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC status phase: Pending, sum | The total amount of persistent volume claims in the Pending phase. |
Dependent item | kube.pvc.status_phase.pending.sum[{#NAMESPACE}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: NS [{#NAMESPACE}] PVC [{#NAME}]: PVC is pending | count(/Kubernetes cluster state by HTTP/kube.pvc.status_phase[{#NAMESPACE}/{#NAME}],2m,,5)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV discovery | Dependent item | kube.pv.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: PV [{#NAME}] Status phase | The current status phase of the persistent volume. |
Dependent item | kube.pv.status_phase[{#NAME}] Preprocessing
|
| Kubernetes: PV [{#NAME}] Capacity bytes | A capacity of the persistent volume in bytes. |
Dependent item | kube.pv.capacity.bytes[{#NAME}] Preprocessing
|
| Kubernetes: PV status phase: Pending, sum | The total amount of persistent volumes in the Pending phase. |
Dependent item | kube.pv.status_phase.pending.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Available, sum | The total amount of persistent volumes in the Available phase. |
Dependent item | kube.pv.status_phase.available.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Bound, sum | The total amount of persistent volumes in the Bound phase. |
Dependent item | kube.pv.status_phase.bound.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Released, sum | The total amount of persistent volumes in the Released phase. |
Dependent item | kube.pv.status_phase.released.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Failed, sum | The total amount of persistent volumes in the Failed phase. |
Dependent item | kube.pv.status_phase.failed.sum[{#SINGLETON}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: PV [{#NAME}]: PV has failed | count(/Kubernetes cluster state by HTTP/kube.pv.status_phase[{#NAME}],2m,,3)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Deployment discovery | Dependent item | kube.deployment.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Paused | Whether the deployment is paused and will not be processed by the deployment controller. |
Dependent item | kube.deployment.spec_paused[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas desired | Number of desired pods for a deployment. |
Dependent item | kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Rollingupdate max unavailable | Maximum number of unavailable replicas during a rolling update of a deployment. |
Dependent item | kube.deployment.rollingupdate.max_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas | The number of replicas per deployment. |
Dependent item | kube.deployment.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas available | The number of available replicas per deployment. |
Dependent item | kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas unavailable | The number of unavailable replicas per deployment. |
Dependent item | kube.deployment.replicas_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas updated | The number of updated replicas per deployment. |
Dependent item | kube.deployment.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas mismatched | The number of available replicas not matching the desired number of replicas. |
Dependent item | kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Deployment replicas mismatch | Deployment has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Endpoint discovery | Dependent item | kube.endpoint.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address available | Number of addresses available in endpoint. |
Dependent item | kube.endpoint.address_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address not ready | Number of addresses not ready in endpoint. |
Dependent item | kube.endpoint.address_not_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Age | Endpoint age (number of seconds since creation). |
Dependent item | kube.endpoint.age[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Node [{#NAME}]: CPU allocatable | The CPU resources of a node that are available for scheduling. |
Dependent item | kube.node.cpu_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Memory allocatable | The memory resources of a node that are available for scheduling. |
Dependent item | kube.node.memory_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Pods allocatable | The pods resources of a node that are available for scheduling. |
Dependent item | kube.node.pods_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Ephemeral storage allocatable | The allocatable ephemeral storage of a node that is available for scheduling. |
Dependent item | kube.node.ephemeral_storage_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: CPU capacity | The capacity for CPU resources of a node. |
Dependent item | kube.node.cpu_capacity[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Memory capacity | The capacity for memory resources of a node. |
Dependent item | kube.node.memory_capacity[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Ephemeral storage capacity | The ephemeral storage capacity of a node. |
Dependent item | kube.node.ephemeral_storage_capacity[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Pods capacity | The capacity for pods resources of a node. |
Dependent item | kube.node.pods_capacity[{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Pending | Pod is in pending state. |
Dependent item | kube.pod.phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Succeeded | Pod is in succeeded state. |
Dependent item | kube.pod.phase.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Failed | Pod is in failed state. |
Dependent item | kube.pod.phase.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Unknown | Pod is in unknown state. |
Dependent item | kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Running | Pod is in unknown state. |
Dependent item | kube.pod.phase.running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers terminated | Describes whether the container is currently in terminated state. |
Dependent item | kube.pod.containers_terminated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers waiting | Describes whether the container is currently in waiting state. |
Dependent item | kube.pod.containers_waiting[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers ready | Describes whether the containers readiness check succeeded. |
Dependent item | kube.pod.containers_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers restarts | The number of container restarts. |
Dependent item | kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers running | Describes whether the container is currently in running state. |
Dependent item | kube.pod.containers_running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Ready | Describes whether the pod is ready to serve requests. |
Dependent item | kube.pod.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Scheduled | Describes the status of the scheduling process for the pod. |
Dependent item | kube.pod.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Unschedulable | Describes the unschedulable status for the pod. |
Dependent item | kube.pod.unschedulable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU limits | The limit on CPU cores to be used by a container. |
Dependent item | kube.pod.containers.limits.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory limits | The limit on memory to be used by a container. |
Dependent item | kube.pod.containers.limits.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU requests | The number of requested CPU cores by a container. |
Dependent item | kube.pod.containers.requests.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory requests | The number of requested memory bytes by a container. |
Dependent item | kube.pod.containers.requests.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is not healthy | min(/Kubernetes cluster state by HTTP/kube.pod.phase.failed[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.pending[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}],10m)>0 |
High | ||
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}])-min(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}],15m))>1 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ReplicaSet discovery | Dependent item | kube.replicaset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas | The number of replicas per ReplicaSet. |
Dependent item | kube.replicaset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Desired replicas | Number of desired pods for a ReplicaSet. |
Dependent item | kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Fully labeled replicas | The number of fully labeled replicas per ReplicaSet. |
Dependent item | kube.replicaset.fully_labeled_replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Ready | The number of ready replicas per ReplicaSet. |
Dependent item | kube.replicaset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the desired number of replicas. |
Dependent item | kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] RS [{#NAME}]: ReplicaSet mismatch | ReplicaSet has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"replicaset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| StatefulSet discovery | Dependent item | kube.statefulset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas | The number of replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Desired replicas | Number of desired pods for a StatefulSet. |
Dependent item | kube.statefulset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Current replicas | The number of current replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Ready replicas | The number of ready replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Updated replicas | The number of updated replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the number of replicas. |
Dependent item | kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet is down | (last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]) / last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}]))<>1 |
High | ||
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet replicas mismatch | StatefulSet has not matched the number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"statefulset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PodDisruptionBudget discovery | Dependent item | kube.pdb.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods healthy | Current number of healthy pods. |
Dependent item | kube.pdb.pods_healthy[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods desired | Minimum desired number of healthy pods. |
Dependent item | kube.pdb.pods_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Disruptions allowed | Number of pod disruptions that are allowed. |
Dependent item | kube.pdb.disruptions_allowed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods total | Total number of pods counted by this disruption budget. |
Dependent item | kube.pdb.pods_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| CronJob discovery | Dependent item | kube.cronjob.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Suspend | Suspend flag tells the controller to suspend subsequent executions. |
Dependent item | kube.cronjob.spec_suspend[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Active | Active holds pointers to currently running jobs. |
Dependent item | kube.cronjob.status_active[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Last schedule | LastScheduleTime keeps information of when was the last time the job was successfully scheduled. |
Dependent item | kube.cronjob.last_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Next schedule | Next time the cronjob should be scheduled. The time after lastScheduleTime or after the cron job's creation time if it's never been scheduled. Use this to determine if the job is delayed. |
Dependent item | kube.cronjob.next_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.cronjob.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.cronjob.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.cronjob.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.cronjob.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Job discovery | Dependent item | kube.job.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.job.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.job.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.job.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.job.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component statuses discovery | Dependent item | kube.componentstatuses.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Component [{#NAME}]: Healthy | Cluster component healthy. |
Dependent item | kube.componentstatuses.healthy[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Component [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}],#3,,"True")<2 and length(last(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz discovery | Dependent item | kube.readyz.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Readyz [{#NAME}]: Healthcheck | Result of readyz healthcheck for component. |
Dependent item | kube.readyz.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Readyz [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}],#3,,"ok")<2 and length(last(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez discovery | Dependent item | kube.livez.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Livez [{#NAME}]: Healthcheck | Result of livez healthcheck for component. |
Dependent item | kube.livez.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Livez [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}],#3,,"ok")<2 and length(last(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift BuildConfig discovery | Dependent item | openshift.buildconfig.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Created | OpenShift BuildConfig Unix creation timestamp. |
Dependent item | openshift.buildconfig.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.buildconfig.generation[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Latest version | The latest version of BuildConfig. |
Dependent item | openshift.buildconfig.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Build discovery | Dependent item | openshift.build.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Namespace [{#NAMESPACE}] Build [{#NAME}]: Created | OpenShift Build Unix creation timestamp. |
Dependent item | openshift.build.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] Build [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.build.sequence.number[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] Build [{#NAME}]: Status phase | The Build phase. |
Dependent item | openshift.build.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| OpenShift: Build [{#NAME}]: Build has failed | count(/Kubernetes cluster state by HTTP/openshift.build.status_phase[{#NAMESPACE}/{#NAME}],2m,"ge",6)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift ClusterResourceQuota discovery | Dependent item | openshift.cluster.resource.quota.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Quota [{#NAME}] Resource [{#RESOURCE}]: Type [{#TYPE}]] | Usage about resource quota. |
Dependent item | openshift.cluster.resource.quota[{#RESOURCE}/{#NAME}/{#TYPE}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Route discovery | Dependent item | openshift.route.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Namespace [{#NAMESPACE}] Route [{#NAME}]: Created | OpenShift Route Unix creation timestamp. |
Dependent item | openshift.route.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] Route [{#NAME}]: Status | Information about route status. |
Dependent item | openshift.route.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| OpenShift: Route [{#NAME}] with issue: Status is false | count(/Kubernetes cluster state by HTTP/openshift.route.status[{#NAMESPACE}/{#NAME}],2m,,0)>=2 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
For Zabbix version: 6.2 and higher. The template to monitor Kubernetes state that work without any external scripts. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Template Kubernetes cluster state by HTTP — collects metrics by HTTP agent from kube-state-metrics endpoint and Kubernetes API.
Don't forget change macros {$KUBE.API.URL} and {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values. NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
This template was tested on:
See Zabbix template operation for basic instructions.
Install the Zabbix Helm Chart in your Kubernetes cluster. Internal service metrics are collected from kube-state-metrics endpoint.
Template needs to use Authorization via API token.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.STATE.ENDPOINT.NAME} with Kube state metrics endpoint name. See kubectl -n monitoring get ep. Default: zabbix-kube-state-metrics.
Also, see the Macros section for a list of macros used to set trigger values. NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Set up the macros to filter the metrics of discovered worker nodes:
Set up macros to filter metrics by namespace:
Set up macros to filter node metrics by nodename:
Note, If you have a large cluster, it is highly recommended to set a filter for discoverable namespaces.
No specific Zabbix configuration is required.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.COMPONENTSTATUSES.ENDPOINT} | Kubernetes API componentstatuses endpoint /api/v1/componentstatuses |
/api/v1/componentstatuses |
| {$KUBE.API.LIVEZ.ENDPOINT} | Kubernetes API livez endpoint /livez |
/livez |
| {$KUBE.API.READYZ.ENDPOINT} | Kubernetes API readyz endpoint /readyz |
/readyz |
| {$KUBE.API.TOKEN} | Service account bearer token |
`` |
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://localhost:6443 |
| {$KUBE.API_SERVER.PORT} | Kubernetes API servers metrics endpoint port. Used in ControlPlane LLD. |
6443 |
| {$KUBE.API_SERVER.SCHEME} | Kubernetes API servers metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.CONTROLLER_MANAGER.PORT} | Kubernetes Controller manager metrics endpoint port. Used in ControlPlane LLD. |
10252 |
| {$KUBE.CONTROLLER_MANAGER.SCHEME} | Kubernetes Controller manager metrics endpoint scheme. Used in ControlPlane LLD. |
http |
| {$KUBE.KUBELET.PORT} | Kubernetes Kubelet manager metrics endpoint port. Used in Kubelet LLD. |
10250 |
| {$KUBE.KUBELET.SCHEME} | Kubernetes Kubelet manager metrics endpoint scheme. Used in Kubelet LLD. |
https |
| {$KUBE.LLD.FILTER.NAMESPACE.MATCHES} | Filter of discoverable pods by namespace |
.* |
| {$KUBE.LLD.FILTER.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered pods by namespace |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes by nodename |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes by nodename |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.WORKER_NODE.MATCHES} | Filter of discoverable worker nodes by nodename |
.* |
| {$KUBE.LLD.FILTER.WORKER_NODE.NOT_MATCHES} | Filter to exclude discovered worker nodes by nodename |
CHANGE_IF_NEEDED |
| {$KUBE.SCHEDULER.PORT} | Kubernetes Scheduler manager metrics endpoint port. Used in ControlPlane LLD. |
10251 |
| {$KUBE.SCHEDULER.SCHEME} | Kubernetes Scheduler manager metrics endpoint scheme. Used in ControlPlane LLD. |
http |
| {$KUBE.STATE.ENDPOINT.NAME} | Kubernetes state endpoint name |
zabbix-kube-state-metrics |
There are no template links in this template.
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| API servers discovery | - |
DEPENDENT | kube.api_servers.discovery |
| Component statuses discovery | - |
DEPENDENT | kube.componentstatuses.discovery Preprocessing: - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT |
| Controller manager nodes discovery | - |
DEPENDENT | kube.controller_manager.discovery |
| CronJob discovery | - |
DEPENDENT | kube.cronjob.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Daemonset discovery | - |
DEPENDENT | kube.daemonset.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Deployment discovery | - |
DEPENDENT | kube.deployment.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Endpoint discovery | - |
DEPENDENT | kube.endpoint.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Job discovery | - |
DEPENDENT | kube.job.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Kubelet discovery | - |
DEPENDENT | kube.kubelet.discovery Filter: AND- {#NAME} MATCHES_REGEX - {#NAME} NOT_MATCHES_REGEX |
| Livez discovery | - |
DEPENDENT | kube.livez.discovery Preprocessing: - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT |
| Node discovery | - |
DEPENDENT | kube.node.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAME} MATCHES_REGEX - {#NAME} NOT_MATCHES_REGEX |
| Pod discovery | - |
DEPENDENT | kube.pod.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| PodDisruptionBudget discovery | - |
DEPENDENT | kube.pdb.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| PVC discovery | - |
DEPENDENT | kube.pvc.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Readyz discovery | - |
DEPENDENT | kube.readyz.discovery Preprocessing: - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT |
| Replicaset discovery | - |
DEPENDENT | kube.replicaset.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Scheduler servers nodes discovery | - |
DEPENDENT | kube.scheduler.discovery |
| Statefulset discovery | - |
DEPENDENT | kube.statefulset.discovery Preprocessing: - PROMETHEUS_TO_JSON - JAVASCRIPT - DISCARD_UNCHANGED_HEARTBEAT Filter: AND- {#NAMESPACE} MATCHES_REGEX - {#NAMESPACE} NOT_MATCHES_REGEX |
| Group | Name | Description | Type | Key and additional info |
|---|---|---|---|---|
| Kubernetes | Kubernetes: Get state metrics | Collecting Kubernetes metrics from kube-state-metrics. |
SCRIPT | kube.state.metrics Expression: The text is too long. Please see the template. |
| Kubernetes | Kubernetes: Control plane LLD | Generation of data for Control plane discovery rules. |
SCRIPT | kube.control_plane.lld Preprocessing: - DISCARD_UNCHANGED_HEARTBEAT: Expression: The text is too long. Please see the template. |
| Kubernetes | Kubernetes: Node LLD | Generation of data for Kubelet discovery rules. |
SCRIPT | kube.node.lld Preprocessing: - DISCARD_UNCHANGED_HEARTBEAT: Expression: The text is too long. Please see the template. |
| Kubernetes | Kubernetes: Get component statuses | - |
HTTP_AGENT | kube.componentstatuses Preprocessing: - CHECK_NOT_SUPPORTED ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Get readyz | - |
HTTP_AGENT | kube.readyz Preprocessing: - JAVASCRIPT: |
| Kubernetes | Kubernetes: Get livez | - |
HTTP_AGENT | kube.livez Preprocessing: - JAVASCRIPT: |
| Kubernetes | Kubernetes: Namespace count | The number of namespaces. |
DEPENDENT | kube.namespace.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: CronJob count | Number of cronjobs. |
DEPENDENT | kube.cronjob.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Job count | Number of jobs(generated by cronjob + job). |
DEPENDENT | kube.job.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Endpoint count | Number of endpoints. |
DEPENDENT | kube.endpoint.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Deployment count | The number of deployments. |
DEPENDENT | kube.deployment.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Service count | The number of services. |
DEPENDENT | kube.service.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Statefulset count | The number of statefulsets. |
DEPENDENT | kube.statefulset.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node count | The number of nodes. |
DEPENDENT | kube.node.count Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Ready | The number of nodes that should be running the daemon pod and have one or more running and ready. |
DEPENDENT | kube.daemonset.ready[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Scheduled | The number of nodes running at least one daemon pod and are supposed to. |
DEPENDENT | kube.daemonset.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Desired | The number of nodes that should be running the daemon pod. |
DEPENDENT | kube.daemonset.desired[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Misscheduled | The number of nodes running a daemon pod but are not supposed to. |
DEPENDENT | kube.daemonset.misscheduled[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Updated number scheduled | The total number of nodes that are running updated daemon pod. |
DEPENDENT | kube.daemonset.updated[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase: Available | Persistent volume claim is currently in Active phase. |
DEPENDENT | kube.pvc.status_phase.active[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase: Lost | Persistent volume claim is currently in Lost phase. |
DEPENDENT | kube.pvc.status_phase.lost[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase: Bound | Persistent volume claim is currently in Bound phase. |
DEPENDENT | kube.pvc.status_phase.bound[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase: Pending | Persistent volume claim is currently in Pending phase. |
DEPENDENT | kube.pvc.status_phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Requested storage | The capacity of storage requested by the persistent volume claim. |
DEPENDENT | kube.pvc.requested.storage[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Status phase: Pending, sum | Persistent volume claim is currently in Pending phase. |
DEPENDENT | kube.pvc.status_phase.pending.sum[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Status phase: Active, sum | Persistent volume claim is currently in Active phase. |
DEPENDENT | kube.pvc.status_phase.active.sum[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Status phase: Bound, sum | Persistent volume claim is currently in Bound phase. |
DEPENDENT | kube.pvc.status_phase.bound.sum[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Status phase: Lost, sum | Persistent volume claim is currently in Lost phase. |
DEPENDENT | kube.pvc.status_phase.lost.sum[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Paused | Whether the deployment is paused and will not be processed by the deployment controller. |
DEPENDENT | kube.deployment.spec_paused[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas desired | Number of desired pods for a deployment. |
DEPENDENT | kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Rollingupdate max unavailable | Maximum number of unavailable replicas during a rolling update of a deployment. |
DEPENDENT | kube.deployment.rollingupdate.max_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas | The number of replicas per deployment. |
DEPENDENT | kube.deployment.replicas[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas available | The number of available replicas per deployment. |
DEPENDENT | kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas unavailable | The number of unavailable replicas per deployment. |
DEPENDENT | kube.deployment.replicas_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas updated | The number of updated replicas per deployment. |
DEPENDENT | kube.deployment.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address available | Number of addresses available in endpoint. |
DEPENDENT | kube.endpoint.address_available[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address not ready | Number of addresses not ready in endpoint. |
DEPENDENT | kube.endpoint.address_not_ready[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Age | Endpoint age (number of seconds since creation). |
DEPENDENT | kube.endpoint.age[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - JAVASCRIPT: |
| Kubernetes | Kubernetes: Node [{#NAME}]: CPU allocatable | The CPU resources of a node that are available for scheduling. |
DEPENDENT | kube.node.cpu_allocatable[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: Memory allocatable | The Memory resources of a node that are available for scheduling. |
DEPENDENT | kube.node.memory_allocatable[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: Pods allocatable | The Pods resources of a node that are available for scheduling. |
DEPENDENT | kube.node.pods_allocatable[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: Ephemeral storage allocatable | The allocatable ephemeral-storage of a node that is available for scheduling. |
DEPENDENT | kube.node.ephemeral_storage_allocatable[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: CPU capacity | The capacity for CPU resources of a node. |
DEPENDENT | kube.node.cpu_capacity[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: Memory capacity | The capacity for Memory resources of a node. |
DEPENDENT | kube.node.memory_capacity[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: Ephemeral storage capacity | The ephemeral-storage capacity of a node. |
DEPENDENT | kube.node.ephemeral_storage_capacity[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Node [{#NAME}]: Pods capacity | The capacity for Pods resources of a node. |
DEPENDENT | kube.node.pods_capacity[{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Pending | Pod is in pending state. |
DEPENDENT | kube.pod.phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Succeeded | Pod is in succeeded state. |
DEPENDENT | kube.pod.phase.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Failed | Pod is in failed state. |
DEPENDENT | kube.pod.phase.failed[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Unknown | Pod is in unknown state. |
DEPENDENT | kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Running | Pod is in unknown state. |
DEPENDENT | kube.pod.phase.running[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers terminated | Describes whether the container is currently in terminated state. |
DEPENDENT | kube.pod.containers_terminated[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers waiting | Describes whether the container is currently in waiting state. |
DEPENDENT | kube.pod.containers_waiting[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers ready | Describes whether the containers readiness check succeeded. |
DEPENDENT | kube.pod.containers_ready[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers restarts | The number of container restarts. |
DEPENDENT | kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers running | Describes whether the container is currently in running state. |
DEPENDENT | kube.pod.containers_running[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Ready | Describes whether the pod is ready to serve requests. |
DEPENDENT | kube.pod.ready[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Scheduled | Describes the status of the scheduling process for the pod. |
DEPENDENT | kube.pod.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Unschedulable | Describes the unschedulable status for the pod. |
DEPENDENT | kube.pod.unschedulable[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU limits | The limit on CPU cores to be used by a container. |
DEPENDENT | kube.pod.containers.limits.cpu[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory limits | The limit on memory to be used by a container. |
DEPENDENT | kube.pod.containers.limits.memory[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU requests | The number of requested cpu cores by a container. |
DEPENDENT | kube.pod.containers.requests.cpu[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory requests | The number of requested memory bytes by a container. |
DEPENDENT | kube.pod.containers.requests.memory[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Replicaset [{#NAME}]: Replicas | The number of replicas per ReplicaSet. |
DEPENDENT | kube.replicaset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Replicaset [{#NAME}]: Desired replicas | Number of desired pods for a ReplicaSet. |
DEPENDENT | kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Replicaset [{#NAME}]: Fully labeled replicas | The number of fully labeled replicas per ReplicaSet. |
DEPENDENT | kube.replicaset.fully_labeled_replicas[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Replicaset [{#NAME}]: Ready | The number of ready replicas per ReplicaSet. |
DEPENDENT | kube.replicaset.ready[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Statefulset [{#NAME}]: Replicas | The number of replicas per StatefulSet. |
DEPENDENT | kube.statefulset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Statefulset [{#NAME}]: Desired replicas | Number of desired pods for a StatefulSet. |
DEPENDENT | kube.statefulset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Statefulset [{#NAME}]: Current replicas | The number of current replicas per StatefulSet. |
DEPENDENT | kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Statefulset [{#NAME}]: Ready replicas | The number of ready replicas per StatefulSet. |
DEPENDENT | kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Statefulset [{#NAME}]: Updated replicas | The number of updated replicas per StatefulSet. |
DEPENDENT | kube.statefulset.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods healthy | Current number of healthy pods. |
DEPENDENT | kube.pdb.pods_healthy[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods desired | Minimum desired number of healthy pods. |
DEPENDENT | kube.pdb.pods_desired[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Disruptions allowed | Number of pod disruptions that are allowed. |
DEPENDENT | kube.pdb.disruptions_allowed[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods total | Total number of pods counted by this disruption budget. |
DEPENDENT | kube.pdb.pods_total[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Suspend | Suspend flag tells the controller to suspend subsequent executions. |
DEPENDENT | kube.cronjob.spec_suspend[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - DISCARD_UNCHANGED_HEARTBEAT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Active | Active holds pointers to currently running jobs. |
DEPENDENT | kube.cronjob.status_active[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Last schedule | LastScheduleTime keeps information of when was the last time the job was successfully scheduled. |
DEPENDENT | kube.cronjob.last_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - JAVASCRIPT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Next schedule | Next time the cronjob should be scheduled. The time after lastScheduleTime, or after the cron job's creation time if it's never been scheduled. Use this to determine if the job is delayed. |
DEPENDENT | kube.cronjob.next_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: - JAVASCRIPT: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Failed | The number of pods which reached Phase Failed and the reason for failure. |
DEPENDENT | kube.cronjob.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Succeeded | The number of pods which reached Phase Succeeded. |
DEPENDENT | kube.cronjob.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion succeeded | Number of job has completed its execution. |
DEPENDENT | kube.cronjob.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion failed | Number of job has failed its execution. |
DEPENDENT | kube.cronjob.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Failed | The number of pods which reached Phase Failed and the reason for failure. |
DEPENDENT | kube.job.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Succeeded | The number of pods which reached Phase Succeeded. |
DEPENDENT | kube.job.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion succeeded | Number of job has completed its execution. |
DEPENDENT | kube.job.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion failed | Number of job has failed its execution. |
DEPENDENT | kube.job.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing: - PROMETHEUS_PATTERN: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Component [{#NAME}]: Healthy | Cluster component healthy. |
DEPENDENT | kube.componentstatuses.healthy[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Readyz [{#NAME}]: Healthcheck | Result of readyz healthcheck for component. |
DEPENDENT | kube.readyz.healthcheck[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: |
| Kubernetes | Kubernetes: Livez [{#NAME}]: Healthcheck | Result of livez healthcheck for component. |
DEPENDENT | kube.livez.healthcheck[{#NAME}] Preprocessing: - JSONPATH: ⛔️ON_FAIL: |
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: NS [{#NAMESPACE}] PVC [{#NAME}]: PVC is pending | - |
min(/Kubernetes cluster state by HTTP/kube.pvc.status_phase.pending[{#NAMESPACE}/{#NAME}],2m)>0 |
WARNING | |
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Deployment replicas mismatch | - |
(last(/Kubernetes cluster state by HTTP/kube.deployment.replicas[{#NAMESPACE}/{#NAME}])-last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}]))<>0 |
WARNING | |
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is not healthy | - |
min(/Kubernetes cluster state by HTTP/kube.pod.phase.failed[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.pending[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}],10m)>0 |
HIGH | |
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is crash looping | - |
(last(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}])-min(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}],#3))>2 |
WARNING | |
| Kubernetes: Namespace [{#NAMESPACE}] RS [{#NAME}]: ReplicasSet mismatch | - |
(last(/Kubernetes cluster state by HTTP/kube.replicaset.replicas[{#NAMESPACE}/{#NAME}])-last(/Kubernetes cluster state by HTTP/kube.replicaset.ready[{#NAMESPACE}/{#NAME}]))<>0 |
WARNING | |
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatfulSet is down | - |
(last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]) / last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}]))<>1 |
HIGH | |
| Kubernetes: Namespace [{#NAMESPACE}] RS [{#NAME}]: Statefulset replicas mismatch | - |
(last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas[{#NAMESPACE}/{#NAME}])-last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]))<>0 |
WARNING | |
| Kubernetes: Component [{#NAME}] is unhealthy | - |
count(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}],#3,,"True")<2 and length(last(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}]))>0 |
WARNING | |
| Kubernetes: Readyz [{#NAME}] is unhealthy | - |
count(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}],#3,,"ok")<2 and length(last(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}]))>0 |
WARNING | |
| Kubernetes: Livez [{#NAME}] is unhealthy | - |
count(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}],#3,,"ok")<2 and length(last(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}]))>0 |
WARNING |
Please report any issues with the template at https://support.zabbix.com.
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums.
The template to monitor Kubernetes state that work without any external scripts. It works without external scripts and uses the script item to make HTTP requests to the Kubernetes API.
Template Kubernetes cluster state by HTTP - collects metrics by HTTP agent from kube-state-metrics endpoint and Kubernetes API.
Don't forget change macros {$KUBE.API.URL} and {$KUBE.API.TOKEN}. Also, see the Macros section for a list of macros used to set trigger values.
NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Zabbix version: 6.0 and higher.
This template has been tested on:
Zabbix should be configured according to the instructions in the Templates out of the box section.
Install the Zabbix Helm Chart in your Kubernetes cluster. Internal service metrics are collected from kube-state-metrics endpoint.
Template needs to use authorization via API token.
Set the {$KUBE.API.URL} such as <scheme>://<host>:<port>.
Get the generated service account token using the command:
kubectl get secret zabbix-service-account -n monitoring -o jsonpath={.data.token} | base64 -d
Then set it to the macro {$KUBE.API.TOKEN}.
Set {$KUBE.STATE.ENDPOINT.NAME} with Kube state metrics endpoint name. See kubectl -n monitoring get ep. Default: zabbix-kube-state-metrics.
NOTE. If you wish to monitor Controller Manager and Scheduler components, you might need to set the --binding-address option for them to the address where Zabbix proxy can reach them.
For example, for clusters created with kubeadm it can be set in the following manifest files (changes will be applied immediately):
Depending on your Kubernetes distribution, you might need to adjust {$KUBE.CONTROL_PLANE.TAINT} macro (for example, set it to node-role.kubernetes.io/master for OpenShift).
NOTE. Some metrics may not be collected depending on your Kubernetes version and configuration.
Also, see the Macros section for a list of macros used to set trigger values.
Set up the macros to filter the metrics of discovered Kubelets by node names:
Set up macros to filter metrics by namespace:
Set up macros to filter node metrics by nodename:
Note: If you have a large cluster, it is highly recommended to set a filter for discoverable namespaces.
You can use the {$KUBE.KUBELET.FILTER.LABELS} and {$KUBE.KUBELET.FILTER.ANNOTATIONS} macros for advanced filtering of kubelets by node labels and annotations.
Notes about labels and annotations filters:
key1: value, key2: regexp).!) to invert the filter (!key: value).For example: kubernetes.io/hostname: kubernetes-node[5-25], !node-role.kubernetes.io/ingress: .*. As a result, the kubelets on nodes 5-25 without the "ingress" role will be discovered.
See the Kubernetes documentation for details about labels and annotations:
You can also set up evaluation periods for replica mismatch triggers (Deployments, ReplicaSets, StatefulSets) with the macro {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD}, which supports context and regular expressions. For example, you can create the following macros:
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:default:nginx-deployment"} = #3
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"deployment:.*:.*"} = #10 or {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:"^deployment.*"} = #10
{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:regex:".*:default:.*"} = 15m
Note that different context macros with regular expressions matching the same string can be applied in an undefined order, and simple context macros (without regular expressions) have higher priority. Read the Important notes section in Zabbix documentation for details.
| Name | Description | Default |
|---|---|---|
| {$KUBE.API.URL} | Kubernetes API endpoint URL in the format |
https://kubernetes.default.svc.cluster.local:443 |
| {$KUBE.API.READYZ.ENDPOINT} | Kubernetes API readyz endpoint /readyz |
/readyz |
| {$KUBE.API.LIVEZ.ENDPOINT} | Kubernetes API livez endpoint /livez |
/livez |
| {$KUBE.API.COMPONENTSTATUSES.ENDPOINT} | Kubernetes API componentstatuses endpoint /api/v1/componentstatuses |
/api/v1/componentstatuses |
| {$KUBE.API.TOKEN} | Service account bearer token. |
|
| {$KUBE.HTTP.PROXY} | Sets the HTTP proxy to |
|
| {$KUBE.STATE.ENDPOINT.NAME} | Kubernetes state endpoint name. |
zabbix-kube-state-metrics |
| {$OPENSHIFT.STATE.ENDPOINT.NAME} | OpenShift state endpoint name. |
openshift-state-metrics |
| {$KUBE.API_SERVER.SCHEME} | Kubernetes API servers metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.API_SERVER.PORT} | Kubernetes API servers metrics endpoint port. Used in ControlPlane LLD. |
6443 |
| {$KUBE.CONTROL_PLANE.TAINT} | Taint that applies to control plane nodes. Change if needed. Used in ControlPlane LLD. |
node-role.kubernetes.io/control-plane |
| {$KUBE.CONTROLLER_MANAGER.SCHEME} | Kubernetes Controller manager metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.CONTROLLER_MANAGER.PORT} | Kubernetes Controller manager metrics endpoint port. Used in ControlPlane LLD. |
10257 |
| {$KUBE.SCHEDULER.SCHEME} | Kubernetes Scheduler metrics endpoint scheme. Used in ControlPlane LLD. |
https |
| {$KUBE.SCHEDULER.PORT} | Kubernetes Scheduler metrics endpoint port. Used in ControlPlane LLD. |
10259 |
| {$KUBE.KUBELET.SCHEME} | Kubernetes Kubelet metrics endpoint scheme. Used in Kubelet LLD. |
https |
| {$KUBE.KUBELET.PORT} | Kubernetes Kubelet metrics endpoint port. Used in Kubelet LLD. |
10250 |
| {$KUBE.LLD.FILTER.NAMESPACE.MATCHES} | Filter of discoverable metrics by namespace. |
.* |
| {$KUBE.LLD.FILTER.NAMESPACE.NOT_MATCHES} | Filter to exclude discovered metrics by namespace. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.NODE.MATCHES} | Filter of discoverable nodes by nodename. |
.* |
| {$KUBE.LLD.FILTER.NODE.NOT_MATCHES} | Filter to exclude discovered nodes by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.LLD.FILTER.KUBELET_NODE.MATCHES} | Filter of discoverable Kubelets by nodename. |
.* |
| {$KUBE.LLD.FILTER.KUBELET_NODE.NOT_MATCHES} | Filter to exclude discovered Kubelets by nodename. |
CHANGE_IF_NEEDED |
| {$KUBE.KUBELET.FILTER.ANNOTATIONS} | Node annotations to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.KUBELET.FILTER.LABELS} | Node labels to filter Kubelets (regex in values are supported). See the template's README.md for details. |
|
| {$KUBE.LLD.FILTER.PV.MATCHES} | Filter of discoverable persistent volumes by name. |
.* |
| {$KUBE.LLD.FILTER.PV.NOT_MATCHES} | Filter to exclude discovered persistent volumes by name. |
CHANGE_IF_NEEDED |
| {$KUBE.REPLICA.MISMATCH.EVAL_PERIOD} | The evaluation period range which is used for calculation of expressions in trigger prototypes (time period or value range). Can be used with context. |
#5 |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Get state metrics | Collecting Kubernetes metrics from kube-state-metrics. |
Script | kube.state.metrics |
| Kubernetes: Control plane LLD | Generation of data for Control plane discovery rules. |
Script | kube.control_plane.lld Preprocessing
|
| Kubernetes: Node LLD | Generation of data for Kubelet discovery rules. |
Script | kube.node.lld Preprocessing
|
| Kubernetes: Get component statuses | HTTP agent | kube.componentstatuses Preprocessing
|
|
| Kubernetes: Get readyz | HTTP agent | kube.readyz Preprocessing
|
|
| Kubernetes: Get livez | HTTP agent | kube.livez Preprocessing
|
|
| Kubernetes: Namespace count | The number of namespaces. |
Dependent item | kube.namespace.count Preprocessing
|
| Kubernetes: CronJob count | Number of cronjobs. |
Dependent item | kube.cronjob.count Preprocessing
|
| Kubernetes: Job count | Number of jobs (generated by cronjob + job). |
Dependent item | kube.job.count Preprocessing
|
| Kubernetes: Endpoint count | Number of endpoints. |
Dependent item | kube.endpoint.count Preprocessing
|
| Kubernetes: Deployment count | The number of deployments. |
Dependent item | kube.deployment.count Preprocessing
|
| Kubernetes: Service count | The number of services. |
Dependent item | kube.service.count Preprocessing
|
| Kubernetes: StatefulSet count | The number of statefulsets. |
Dependent item | kube.statefulset.count Preprocessing
|
| Kubernetes: Node count | The number of nodes. |
Dependent item | kube.node.count Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| API servers discovery | Dependent item | kube.api_servers.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Controller manager nodes discovery | Dependent item | kube.controller_manager.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Scheduler servers nodes discovery | Dependent item | kube.scheduler.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubelet discovery | Dependent item | kube.kubelet.discovery |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Daemonset discovery | Dependent item | kube.daemonset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Ready | The number of nodes that should be running the daemon pod and have one or more running and ready. |
Dependent item | kube.daemonset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Scheduled | The number of nodes that run at least one daemon pod and are supposed to. |
Dependent item | kube.daemonset.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Desired | The number of nodes that should be running the daemon pod. |
Dependent item | kube.daemonset.desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Misscheduled | The number of nodes that run a daemon pod but are not supposed to. |
Dependent item | kube.daemonset.misscheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Daemonset [{#NAME}]: Updated number scheduled | The total number of nodes that are running updated daemon pod. |
Dependent item | kube.daemonset.updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PVC discovery | Dependent item | kube.pvc.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Status phase | The current status phase of the persistent volume claim. |
Dependent item | kube.pvc.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC [{#NAME}] Requested storage | The capacity of storage requested by the persistent volume claim. |
Dependent item | kube.pvc.requested.storage[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC status phase: Bound, sum | The total amount of persistent volume claims in the Bound phase. |
Dependent item | kube.pvc.status_phase.bound.sum[{#NAMESPACE}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC status phase: Lost, sum | The total amount of persistent volume claims in the Lost phase. |
Dependent item | kube.pvc.status_phase.lost.sum[{#NAMESPACE}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PVC status phase: Pending, sum | The total amount of persistent volume claims in the Pending phase. |
Dependent item | kube.pvc.status_phase.pending.sum[{#NAMESPACE}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: NS [{#NAMESPACE}] PVC [{#NAME}]: PVC is pending | count(/Kubernetes cluster state by HTTP/kube.pvc.status_phase[{#NAMESPACE}/{#NAME}],2m,,5)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PV discovery | Dependent item | kube.pv.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: PV [{#NAME}] Status phase | The current status phase of the persistent volume. |
Dependent item | kube.pv.status_phase[{#NAME}] Preprocessing
|
| Kubernetes: PV [{#NAME}] Capacity bytes | A capacity of the persistent volume in bytes. |
Dependent item | kube.pv.capacity.bytes[{#NAME}] Preprocessing
|
| Kubernetes: PV status phase: Pending, sum | The total amount of persistent volumes in the Pending phase. |
Dependent item | kube.pv.status_phase.pending.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Available, sum | The total amount of persistent volumes in the Available phase. |
Dependent item | kube.pv.status_phase.available.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Bound, sum | The total amount of persistent volumes in the Bound phase. |
Dependent item | kube.pv.status_phase.bound.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Released, sum | The total amount of persistent volumes in the Released phase. |
Dependent item | kube.pv.status_phase.released.sum[{#SINGLETON}] Preprocessing
|
| Kubernetes: PV status phase: Failed, sum | The total amount of persistent volumes in the Failed phase. |
Dependent item | kube.pv.status_phase.failed.sum[{#SINGLETON}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: PV [{#NAME}]: PV has failed | count(/Kubernetes cluster state by HTTP/kube.pv.status_phase[{#NAME}],2m,,3)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Deployment discovery | Dependent item | kube.deployment.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Paused | Whether the deployment is paused and will not be processed by the deployment controller. |
Dependent item | kube.deployment.spec_paused[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas desired | Number of desired pods for a deployment. |
Dependent item | kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Rollingupdate max unavailable | Maximum number of unavailable replicas during a rolling update of a deployment. |
Dependent item | kube.deployment.rollingupdate.max_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas | The number of replicas per deployment. |
Dependent item | kube.deployment.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas available | The number of available replicas per deployment. |
Dependent item | kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas unavailable | The number of unavailable replicas per deployment. |
Dependent item | kube.deployment.replicas_unavailable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas updated | The number of updated replicas per deployment. |
Dependent item | kube.deployment.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Replicas mismatched | The number of available replicas not matching the desired number of replicas. |
Dependent item | kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Deployment [{#NAME}]: Deployment replicas mismatch | Deployment has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.deployment.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"deployment:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.deployment.replicas_available[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Endpoint discovery | Dependent item | kube.endpoint.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address available | Number of addresses available in endpoint. |
Dependent item | kube.endpoint.address_available[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Address not ready | Number of addresses not ready in endpoint. |
Dependent item | kube.endpoint.address_not_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Endpoint [{#NAME}]: Age | Endpoint age (number of seconds since creation). |
Dependent item | kube.endpoint.age[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Node discovery | Dependent item | kube.node.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Node [{#NAME}]: CPU allocatable | The CPU resources of a node that are available for scheduling. |
Dependent item | kube.node.cpu_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Memory allocatable | The memory resources of a node that are available for scheduling. |
Dependent item | kube.node.memory_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Pods allocatable | The pods resources of a node that are available for scheduling. |
Dependent item | kube.node.pods_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Ephemeral storage allocatable | The allocatable ephemeral storage of a node that is available for scheduling. |
Dependent item | kube.node.ephemeral_storage_allocatable[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: CPU capacity | The capacity for CPU resources of a node. |
Dependent item | kube.node.cpu_capacity[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Memory capacity | The capacity for memory resources of a node. |
Dependent item | kube.node.memory_capacity[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Ephemeral storage capacity | The ephemeral storage capacity of a node. |
Dependent item | kube.node.ephemeral_storage_capacity[{#NAME}] Preprocessing
|
| Kubernetes: Node [{#NAME}]: Pods capacity | The capacity for pods resources of a node. |
Dependent item | kube.node.pods_capacity[{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Pod discovery | Dependent item | kube.pod.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Pending | Pod is in pending state. |
Dependent item | kube.pod.phase.pending[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Succeeded | Pod is in succeeded state. |
Dependent item | kube.pod.phase.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Failed | Pod is in failed state. |
Dependent item | kube.pod.phase.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Unknown | Pod is in unknown state. |
Dependent item | kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}] Phase: Running | Pod is in unknown state. |
Dependent item | kube.pod.phase.running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers terminated | Describes whether the container is currently in terminated state. |
Dependent item | kube.pod.containers_terminated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers waiting | Describes whether the container is currently in waiting state. |
Dependent item | kube.pod.containers_waiting[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers ready | Describes whether the containers readiness check succeeded. |
Dependent item | kube.pod.containers_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers restarts | The number of container restarts. |
Dependent item | kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers running | Describes whether the container is currently in running state. |
Dependent item | kube.pod.containers_running[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Ready | Describes whether the pod is ready to serve requests. |
Dependent item | kube.pod.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Scheduled | Describes the status of the scheduling process for the pod. |
Dependent item | kube.pod.scheduled[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Unschedulable | Describes the unschedulable status for the pod. |
Dependent item | kube.pod.unschedulable[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU limits | The limit on CPU cores to be used by a container. |
Dependent item | kube.pod.containers.limits.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory limits | The limit on memory to be used by a container. |
Dependent item | kube.pod.containers.limits.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers CPU requests | The number of requested CPU cores by a container. |
Dependent item | kube.pod.containers.requests.cpu[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Containers memory requests | The number of requested memory bytes by a container. |
Dependent item | kube.pod.containers.requests.memory[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is not healthy | min(/Kubernetes cluster state by HTTP/kube.pod.phase.failed[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.pending[{#NAMESPACE}/{#NAME}],10m)>0 or min(/Kubernetes cluster state by HTTP/kube.pod.phase.unknown[{#NAMESPACE}/{#NAME}],10m)>0 |
High | ||
| Kubernetes: Namespace [{#NAMESPACE}] Pod [{#NAME}]: Pod is crash looping | Containers of the pod keep restarting. This most likely indicates that the pod is in the CrashLoopBackOff state. |
(last(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}])-min(/Kubernetes cluster state by HTTP/kube.pod.containers_restarts[{#NAMESPACE}/{#NAME}],15m))>1 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| ReplicaSet discovery | Dependent item | kube.replicaset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas | The number of replicas per ReplicaSet. |
Dependent item | kube.replicaset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Desired replicas | Number of desired pods for a ReplicaSet. |
Dependent item | kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Fully labeled replicas | The number of fully labeled replicas per ReplicaSet. |
Dependent item | kube.replicaset.fully_labeled_replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Ready | The number of ready replicas per ReplicaSet. |
Dependent item | kube.replicaset.ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] ReplicaSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the desired number of replicas. |
Dependent item | kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] RS [{#NAME}]: ReplicaSet mismatch | ReplicaSet has not matched the expected number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"replicaset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.replicas_desired[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.replicaset.ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| StatefulSet discovery | Dependent item | kube.statefulset.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas | The number of replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Desired replicas | Number of desired pods for a StatefulSet. |
Dependent item | kube.statefulset.replicas_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Current replicas | The number of current replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Ready replicas | The number of ready replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Updated replicas | The number of updated replicas per StatefulSet. |
Dependent item | kube.statefulset.replicas_updated[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: Replicas mismatched | The number of ready replicas not matching the number of replicas. |
Dependent item | kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet is down | (last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}]) / last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_current[{#NAMESPACE}/{#NAME}]))<>1 |
High | ||
| Kubernetes: Namespace [{#NAMESPACE}] StatefulSet [{#NAME}]: StatefulSet replicas mismatch | StatefulSet has not matched the number of replicas during the specified trigger evaluation period. |
min(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_mismatched[{#NAMESPACE}/{#NAME}],{$KUBE.REPLICA.MISMATCH.EVAL_PERIOD:"statefulset:{#NAMESPACE}:{#NAME}"})>0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas[{#NAMESPACE}/{#NAME}])>=0 and last(/Kubernetes cluster state by HTTP/kube.statefulset.replicas_ready[{#NAMESPACE}/{#NAME}])>=0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| PodDisruptionBudget discovery | Dependent item | kube.pdb.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods healthy | Current number of healthy pods. |
Dependent item | kube.pdb.pods_healthy[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods desired | Minimum desired number of healthy pods. |
Dependent item | kube.pdb.pods_desired[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Disruptions allowed | Number of pod disruptions that are allowed. |
Dependent item | kube.pdb.disruptions_allowed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] PodDisruptionBudget [{#NAME}]: Pods total | Total number of pods counted by this disruption budget. |
Dependent item | kube.pdb.pods_total[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| CronJob discovery | Dependent item | kube.cronjob.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Suspend | Suspend flag tells the controller to suspend subsequent executions. |
Dependent item | kube.cronjob.spec_suspend[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Active | Active holds pointers to currently running jobs. |
Dependent item | kube.cronjob.status_active[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Last schedule | LastScheduleTime keeps information of when was the last time the job was successfully scheduled. |
Dependent item | kube.cronjob.last_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Next schedule | Next time the cronjob should be scheduled. The time after lastScheduleTime or after the cron job's creation time if it's never been scheduled. Use this to determine if the job is delayed. |
Dependent item | kube.cronjob.next_schedule_time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.cronjob.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.cronjob.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.cronjob.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] CronJob [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.cronjob.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Job discovery | Dependent item | kube.job.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Failed | The number of pods which reached the Failed phase and the reason for failure. |
Dependent item | kube.job.status_failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Succeeded | The number of pods which reached the Succeeded phase. |
Dependent item | kube.job.status_succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion succeeded | Number of jobs the execution of which has been completed. |
Dependent item | kube.job.completion.succeeded[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Kubernetes: Namespace [{#NAMESPACE}] Job [{#NAME}]: Completion failed | Number of jobs the execution of which has failed. |
Dependent item | kube.job.completion.failed[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Component statuses discovery | Dependent item | kube.componentstatuses.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Component [{#NAME}]: Healthy | Cluster component healthy. |
Dependent item | kube.componentstatuses.healthy[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Component [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}],#2,"ne","True")=2 and length(last(/Kubernetes cluster state by HTTP/kube.componentstatuses.healthy[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Readyz discovery | Dependent item | kube.readyz.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Readyz [{#NAME}]: Healthcheck | Result of readyz healthcheck for component. |
Dependent item | kube.readyz.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Readyz [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.readyz.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Livez discovery | Dependent item | kube.livez.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| Kubernetes: Livez [{#NAME}]: Healthcheck | Result of livez healthcheck for component. |
Dependent item | kube.livez.healthcheck[{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| Kubernetes: Livez [{#NAME}] is unhealthy | count(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}],#2,"ne","ok")=2 and length(last(/Kubernetes cluster state by HTTP/kube.livez.healthcheck[{#NAME}]))>0 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift BuildConfig discovery | Dependent item | openshift.buildconfig.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Created | OpenShift BuildConfig Unix creation timestamp. |
Dependent item | openshift.buildconfig.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.buildconfig.generation[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] BuildConfig [{#NAME}]: Latest version | The latest version of BuildConfig. |
Dependent item | openshift.buildconfig.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Build discovery | Dependent item | openshift.build.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Namespace [{#NAMESPACE}] Build [{#NAME}]: Created | OpenShift Build Unix creation timestamp. |
Dependent item | openshift.build.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] Build [{#NAME}]: Generation | Sequence number representing a specific generation of the desired state. |
Dependent item | openshift.build.sequence.number[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] Build [{#NAME}]: Status phase | The Build phase. |
Dependent item | openshift.build.status_phase[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| OpenShift: Build [{#NAME}]: Build has failed | count(/Kubernetes cluster state by HTTP/openshift.build.status_phase[{#NAMESPACE}/{#NAME}],2m,"ge",6)>=2 |
Warning |
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift ClusterResourceQuota discovery | Dependent item | openshift.cluster.resource.quota.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Quota [{#NAME}] Resource [{#RESOURCE}]: Type [{#TYPE}]] | Usage about resource quota. |
Dependent item | openshift.cluster.resource.quota[{#RESOURCE}/{#NAME}/{#TYPE}] Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift Route discovery | Dependent item | openshift.route.discovery Preprocessing
|
| Name | Description | Type | Key and additional info |
|---|---|---|---|
| OpenShift: Namespace [{#NAMESPACE}] Route [{#NAME}]: Created | OpenShift Route Unix creation timestamp. |
Dependent item | openshift.route.created.time[{#NAMESPACE}/{#NAME}] Preprocessing
|
| OpenShift: Namespace [{#NAMESPACE}] Route [{#NAME}]: Status | Information about route status. |
Dependent item | openshift.route.status[{#NAMESPACE}/{#NAME}] Preprocessing
|
| Name | Description | Expression | Severity | Dependencies and additional info |
|---|---|---|---|---|
| OpenShift: Route [{#NAME}] with issue: Status is false | count(/Kubernetes cluster state by HTTP/openshift.route.status[{#NAMESPACE}/{#NAME}],2m,,0)>=2 |
Warning |
Please report any issues with the template at https://support.zabbix.com
You can also provide feedback, discuss the template, or ask for help at ZABBIX forums
| Link | Source | Compatibility | Type, Technology | Created Updated | Rating |
|---|---|---|---|---|---|
| Template Docker Monitoring of Docker container by using Zabbix. Available CPU, mem, blkio, net container metrics and some containers config details, e.g. IP, name, ... Zabbix Docker module has native support for Docker containers (Systemd included) and should also support a few other container types (e.g. LXC) out of ... github.com/monitoringartist/zabbix-docker-monitoring |
GitHub |
Docker, Templates,Userparameters LLD, Docker |
2014-08-03 2 y |
Popular
|
|
| Kube by Prom API Descriptionzabbix-kube-prom is a batch of Zabbix LLD templates for Zabbix server.It is used for external Kubernetes monitoring by Zabbix via Prometheus API.InstallationInstall kube-prometheus-stack into the Kubernetes cluster.Import global Zabbix Template (zabbix-kube-prom.xml) into your Zabbix server.Create ... template_zabbix-kube-prom |
GitHub Community Templates |
5.0+ |
| ||
| CHECK_KUBERNETES Nagios-style checks against Kubernetes API. Designed for usage with Nagios, Icinga, Zabbix... Whatever github.com/agapoff/check_kubernetes |
GitHub |
Template, Userparameters LLD, Bash |
2018-06-29 2 y |
Popular
|
|
| Kubernetes operator Working with a postgres database container, though there is an option to provide a custom database container gitlab.com/frenchtoasters/zabbix-operator |
| 2020-05-05 |