HA setup for clusters with regular maintenance windows
Overview
In some clusters, nodes are rotated regularly, for example during node image upgrades or maintenance windows, in which nodes are drained and restarted one after another. In such clusters, the HDFS storage of the SUSE® Observability topology database is restarted multiple times within a short period.
The 150-ha, 250-ha and 500-ha sizing profiles run HDFS with 3 data nodes and store every block twice. When data nodes restart in quick succession while data is being written, HDFS becomes more vulnerable to further node failures. When a data node restarts, a maximum of two nodes are active. Consequently, any subsequent additional node issue risks downtime, or worse, data corruption and data loss. For clusters with regular maintenance windows, you are recommend to use the HDFS settings of the 4000-ha profile, which stores every block 3 times on 5 data nodes:
150-ha, 250-ha, 500-ha |
Recommended (as in 4000-ha) |
|
|---|---|---|
HDFS data nodes |
3 |
5 |
Copies of every block ( |
2 |
3 |
Minimum copies of a block ( |
2 |
2 |
These settings add resilience, but they do not replace a careful maintenance procedure. For general guidance on maintaining and rotating nodes, including the checks to perform before disrupting the next node, see Maintain or rotate nodes. That section is part of the Longhorn page, but it applies regardless of Kubernetes distribution, storage driver or cluster management platform.
Prerequisites
-
SUSE® Observability Helm chart version 2.10.0 or later, installed with the
150-ha,250-haor500-hasizing profile. -
At least 5 Kubernetes nodes. Every HDFS data node runs on a separate node.
-
Capacity for 2 additional HDFS data nodes. Each data node requests 600m CPU and 4Gi memory and uses its own persistent volume (250Gi by default). The volume of data per data node does not increase as every block is stored 3 times across 5 data nodes.
Configure the HDFS settings
Add the following to the values.yaml file to install or upgrade SUSE® Observability:
hbase:
hdfs:
replication: 3
minReplication: 2
datanode:
replicaCount: 5
-
Apply the configuration by installing or upgrading SUSE® Observability:
helm upgrade \
--install \
--namespace suse-observability \
--values values.yaml \
suse-observability \
suse-observability/suse-observability
-
If you deploy with
--setflags instead of a values file, add these flags to yourhelm upgradecommand:
--set hbase.hdfs.replication=3 \
--set hbase.hdfs.minReplication=2 \
--set hbase.hdfs.datanode.replicaCount=5 \
|
On an existing installation, the new settings apply to data written after the upgrade. Data that already exists keeps 2 copies until it is rewritten. |