Monitoring & High Availability

A system is only truly in operation when someone sees the outage coming before it happens. We build monitoring that measures the right things and spots problems. And we build architectures that absorb a failure: failover clusters since the Solaris and Veritas days, today Patroni on Kubernetes, MariaDB Galera, HAProxy/nginx/Apache/Squid with Keepalived and GlusterFS. Plus backup and recovery that we actually test on a regular basis.

For clients with high availability requirements we build systems that survive the loss of an entire data center (disaster recovery). We test the real thing with our clients on a regular basis.

Typical tasks

  • Build system and service monitoring with checkmk or Nagios, alerting with sense and proportion
  • Metrics and logs: Grafana, InfluxDB and Telegraf, ELK stack with Logstash and Filebeat
  • Database monitoring with Oracle Enterprise Manager and Grid Control, introduced from 10g to 12c
  • Backup and recovery with RMAN and Veritas NetBackup, recovery tests instead of hope
  • Failover clusters and high availability: Veritas Cluster Server, Patroni, MariaDB Galera, HAProxy and Keepalived, GlusterFS, Samba with CTDB
  • Patch management and security patching, secrets with HashiCorp Vault

From practice

For RCI Banque we built a highly available Nextcloud system with MariaDB Galera and GlusterFS and an HA Samba system with CTDB; the OpenShift clusters are monitored with checkmk. An older example: for around 120 test databases at Vodafone we designed and implemented a failover cluster concept with Veritas Cluster Server and EMC SAN between 2002 and 2007 and introduced Oracle Enterprise Manager Grid Control as the central database monitoring.

Frequently asked questions

Which monitoring tools do you use?
checkmk and Nagios for system and service monitoring, Grafana with InfluxDB and Telegraf for metrics, ELK for logs, Oracle Enterprise Manager for databases.
Do you really test backups?
Yes. A backup without a recovery test is a hope. We plan recovery tests with RMAN and Veritas NetBackup as a fixed part of operations.
How do you achieve high availability for databases?
Depending on the system: Oracle with RAC and Data Guard, PostgreSQL with Patroni on Kubernetes, MariaDB with Galera, HAProxy and Keepalived in front. We have built failover clusters since the Veritas Cluster days.

Not sure which field fits?

Describe the problem to us, not the technology.

Contact