Hello there,
We recently made many of our core services highly available, zabbix server being one of them.
We're using a Patroni cluster of postgres databases behind a load balancer running on our Firewall.
This entails only one 'master' postgres server ever being read from and/or written to at any given time.
These database servers are running with SSD based storage behind them.
Zabbix was the first service we spun up to use this new Patroni cluster - previously we had a single Zabbix server v7.0.23 with a locally running MySQL backend (as opposed to postgres).
We had the same amount of hosts and general configuration from the old server, this recent exercise was really to just change to being highly available - not to change the overall host monitoring scheme or count.
This new postgres backend isn't being loaded/used-by any other services other than Zabbix at this time.
From my current research there are some suggestions this could be database slowness - though I've yet to confirm that is truly the cause and the spiking nature makes me dubious of that diagnosis.
For whatever reason we're getting random spikes of Utilization of configuration syncer worker internal processes hitting 100% for 17mins every occurrence.
The time between occurrences appears random and does not correlate with any other backend operations in our environment.
I'm attaching some graphs to show that on both HA Zabbix servers the spikes occur at the same time.
In addition is a last 7 days graph showing the inconsistency between occurrences.
I'm really stumped as to what's happening here.
Most of our hosts are Active Zabbix Agents (Linux, macOS, Windows), however the Windows Zabbix Agent Active template is the most used (by 317 hosts at this time).
I attempted to modify the template to bring the Windows services discovery to 6hr intervals rather than the default 1hr interval - but this really had no change on these spikes we're seeing.
Ultimately we're not seeing any performance or data gathering issues, but I'd like to address these spikes rather than just ignore them.
Here are some stats on our old and new server estates, do let me know if I can clarify anything further.
Old Zabbix server with MySQL Backend stats:
Zabbix server is running Yes localhost:10051
Zabbix server version 7.0.23 New update available
Zabbix frontend version 7.0.23 New update available
Number of hosts (enabled/disabled) 493 480 / 13
Number of templates 330
Number of items (enabled/disabled/not supported) 80840 77154 / 3096 / 590
Number of triggers (enabled/disabled [problem/ok]) 54706 24899 / 29807 [1307 / 23592]
Number of users (online) 21 1
Required server performance, new values per second 1138.21
High availability cluster Disabled
Current Server stats:
Zabbix server is running Yes ns-zabbix:10051
Zabbix server version 7.0.30 Up to date
Zabbix frontend version 7.0.30 Up to date
Number of hosts (enabled/disabled) 455 455 / 0
Number of templates 375
Number of items (enabled/disabled/not supported) 77048 76150 / 385 / 513
Number of triggers (enabled/disabled [problem/ok]) 51736 26458 / 25278 [448 / 26010]
Number of users (online) 17 1
Required server performance, new values per second 1119.11
High availability cluster Enabled Fail-over delay: 1 minute
Thank you for reading.
We recently made many of our core services highly available, zabbix server being one of them.
We're using a Patroni cluster of postgres databases behind a load balancer running on our Firewall.
This entails only one 'master' postgres server ever being read from and/or written to at any given time.
These database servers are running with SSD based storage behind them.
Zabbix was the first service we spun up to use this new Patroni cluster - previously we had a single Zabbix server v7.0.23 with a locally running MySQL backend (as opposed to postgres).
We had the same amount of hosts and general configuration from the old server, this recent exercise was really to just change to being highly available - not to change the overall host monitoring scheme or count.
This new postgres backend isn't being loaded/used-by any other services other than Zabbix at this time.
From my current research there are some suggestions this could be database slowness - though I've yet to confirm that is truly the cause and the spiking nature makes me dubious of that diagnosis.
For whatever reason we're getting random spikes of Utilization of configuration syncer worker internal processes hitting 100% for 17mins every occurrence.
The time between occurrences appears random and does not correlate with any other backend operations in our environment.
I'm attaching some graphs to show that on both HA Zabbix servers the spikes occur at the same time.
In addition is a last 7 days graph showing the inconsistency between occurrences.
I'm really stumped as to what's happening here.
Most of our hosts are Active Zabbix Agents (Linux, macOS, Windows), however the Windows Zabbix Agent Active template is the most used (by 317 hosts at this time).
I attempted to modify the template to bring the Windows services discovery to 6hr intervals rather than the default 1hr interval - but this really had no change on these spikes we're seeing.
Ultimately we're not seeing any performance or data gathering issues, but I'd like to address these spikes rather than just ignore them.
Here are some stats on our old and new server estates, do let me know if I can clarify anything further.
Old Zabbix server with MySQL Backend stats:
Zabbix server is running Yes localhost:10051
Zabbix server version 7.0.23 New update available
Zabbix frontend version 7.0.23 New update available
Number of hosts (enabled/disabled) 493 480 / 13
Number of templates 330
Number of items (enabled/disabled/not supported) 80840 77154 / 3096 / 590
Number of triggers (enabled/disabled [problem/ok]) 54706 24899 / 29807 [1307 / 23592]
Number of users (online) 21 1
Required server performance, new values per second 1138.21
High availability cluster Disabled
Current Server stats:
Zabbix server is running Yes ns-zabbix:10051
Zabbix server version 7.0.30 Up to date
Zabbix frontend version 7.0.30 Up to date
Number of hosts (enabled/disabled) 455 455 / 0
Number of templates 375
Number of items (enabled/disabled/not supported) 77048 76150 / 385 / 513
Number of triggers (enabled/disabled [problem/ok]) 51736 26458 / 25278 [448 / 26010]
Number of users (online) 17 1
Required server performance, new values per second 1119.11
High availability cluster Enabled Fail-over delay: 1 minute
Thank you for reading.