Has anyone built a proper ML-based predictive monitoring layer on top of Zabbix?
Asking because I genuinely couldn't find one.
And to be clear . It's about any metric that's heading toward trouble: CPU, memory, disk, queue depth, whatever. The real question is always the same: how long until this becomes a problem? Answer that early enough, and you avoid downtime and the cost that comes with it — both the financial hit and the scramble to fix things at 3am.
Zabbix's built-in forecast()/timeleft() functions are plain linear regression. Fine for something that grows in a straight line — useless for anything with a weekly pattern (backups, log rotation) or an accelerating trend. It just doesn't see it.
So I've been building a pipeline that:
→ Pulls history data via the Zabbix API
→ Trains one model pooled across ALL monitored items together — not per-item, since most metrics never actually breach a threshold within a given training window, so per-item training on crossing events alone starves on data
→ Predicts values at multiple horizons and interpolates to estimate days-to-threshold
XGBoost is just the current model — a solid, explainable baseline to prove the approach works. The pipeline itself is model-agnostic; the plan is to eventually plug in stronger ML approaches or even LLM-based reasoning on top, once the base is solid.
What I couldn't find anywhere: someone doing this for Zabbix specifically, at this level. Grafana Cloud ML offers something similar for Prometheus — nothing comparable for Zabbix shops.
Genuine question: has anyone here built or seen a similar approach for Zabbix? Would love to compare notes.
Asking because I genuinely couldn't find one.
And to be clear . It's about any metric that's heading toward trouble: CPU, memory, disk, queue depth, whatever. The real question is always the same: how long until this becomes a problem? Answer that early enough, and you avoid downtime and the cost that comes with it — both the financial hit and the scramble to fix things at 3am.
Zabbix's built-in forecast()/timeleft() functions are plain linear regression. Fine for something that grows in a straight line — useless for anything with a weekly pattern (backups, log rotation) or an accelerating trend. It just doesn't see it.
So I've been building a pipeline that:
→ Pulls history data via the Zabbix API
→ Trains one model pooled across ALL monitored items together — not per-item, since most metrics never actually breach a threshold within a given training window, so per-item training on crossing events alone starves on data
→ Predicts values at multiple horizons and interpolates to estimate days-to-threshold
XGBoost is just the current model — a solid, explainable baseline to prove the approach works. The pipeline itself is model-agnostic; the plan is to eventually plug in stronger ML approaches or even LLM-based reasoning on top, once the base is solid.
What I couldn't find anywhere: someone doing this for Zabbix specifically, at this level. Grafana Cloud ML offers something similar for Prometheus — nothing comparable for Zabbix shops.
Genuine question: has anyone here built or seen a similar approach for Zabbix? Would love to compare notes.
Comment