Replacing a healthy etcd member¶
You might need to replace a healthy etcd member for planned hardware maintenance, hardware upgrades, or migration to new infrastructure.
About replacing a healthy etcd member¶
To replace a control plane node without disrupting etcd, remove a healthy member and add a replacement while the cluster remains operational. The procedure you follow depends on how your cluster was installed and whether it uses the Machine API and a control plane machine set.
Note
If the etcd member is unhealthy because the machine is not running, the node is not ready, or the etcd pod is crashlooping, see "Replacing an unhealthy etcd member".
If you have lost the majority of your control plane hosts, see "Restoring to an earlier cluster state".
Replacing a healthy etcd member¶
To replace a healthy etcd member without disrupting cluster operations, choose the procedure that matches your control plane configuration. You can use a control plane machine set, the Machine API, or scale up and scale down control plane nodes.
Warning
Take an etcd backup before you replace a healthy etcd member so that you can restore your cluster if any issues occur. For more information, see "Backing up etcd data".
Depending on your cluster configuration, use one of the following procedures:
- Replacing a healthy etcd member with a control plane machine set
- Replacing a healthy etcd member with the Machine API
- Replacing a healthy etcd member by scaling up and scaling down
For clusters that were installed by using the Assisted Installer, see "Replacing a control plane node in a healthy cluster" in the Assisted Installer documentation.
Determining how to replace a healthy etcd member¶
To choose the correct procedure for replacing a healthy etcd member, check whether your cluster uses the Assisted Installer, a control plane machine set, or the Machine API. Use the OpenShift CLI (oc) to identify your cluster configuration and follow the matching replacement procedure.
Prerequisites
- You installed the OpenShift CLI (
oc). - You logged in to
ocas a user with thecluster-adminrole.
Procedure
-
Check whether the cluster was installed by using the Assisted Installer by running the following command:
- If the command returns one or more
AgentClusterInstallresources, follow the procedure in "Replacing a control plane node in a healthy cluster" in the Assisted Installer documentation. - If the command returns no resources, continue with the following steps.
- If the command returns one or more
-
Check whether the cluster has a control plane machine set by running the following command:
- If the command returns a
ControlPlaneMachineSetresource, follow the procedure in "Replacing a healthy etcd member with a control plane machine set". - If the command returns no resources, continue to the next step.
- If the command returns a
-
Check whether the cluster has control plane
Machineobjects by running the following command:- If
Machineobjects exist, follow the procedure in "Replacing a healthy etcd member with the Machine API". - If there are no
Machineobjects, follow the procedure in "Replacing a healthy etcd member by scaling up and scaling down".
- If
Replacing a healthy etcd member with a control plane machine set¶
On clusters that use a control plane machine set, you can replace a healthy control plane machine by deleting the corresponding Machine object.
The control plane machine set creates a replacement machine, and the etcd Operator uses machine lifecycle hooks to protect etcd quorum during the replacement.
For more information about how quorum protection works during control plane machine deletion, see "Quorum protection with machine lifecycle hooks".
Prerequisites
-
The cluster has a
ControlPlaneMachineSetresource. -
You have access to the cluster as a user with the
cluster-adminrole. -
You have taken an etcd backup. For more information, see "Backing up etcd data".
Warning
Take an etcd backup before you replace a healthy etcd member so that you can restore your cluster if any issues occur.
Procedure
-
List the control plane machines in your cluster by running the following command:
-
Identify the control plane machine that corresponds to the node that you want to replace.
-
Optional. If you are performing planned maintenance, cordon the node by running the following command:
Replace
<node_name>with the name of the node that you are replacing.Warning
Delete only one control plane machine at a time. Deleting multiple control plane machines at the same time can cause etcd quorum loss.
-
Delete the control plane machine by running the following command:
Replace
<control_plane_machine_name>with the name of the control plane machine to delete.Note
If you delete multiple control plane machines, the control plane machine set replaces them according to the configured update strategy:
- For clusters that use the default
RollingUpdateupdate strategy, the Operator replaces one machine at a time until each machine is replaced. - For clusters that are configured to use the
OnDeleteupdate strategy, the Operator creates all of the required replacement machines simultaneously.
Both strategies maintain etcd health during control plane machine replacement.
- For clusters that use the default
-
Monitor the replacement by running the following commands:
-
Verify that a new control plane machine is created:
-
Verify that the etcd cluster Operator reports
Available=TrueandDegraded=False:Note
During the replacement
Progressing=Trueis expected and transitions toFalseonce the new member is fully reconciled. -
Verify that all control plane nodes are in the
Readystate:
-
Verification
-
Verify etcd health by running the following commands:
-
Open a remote shell session to a control plane etcd pod:
Replace
<etcd_pod_name>with the name of a running etcd pod. -
Check endpoint health:
Expected output shows
is healthyfor each endpoint. -
List etcd members and verify that the cluster has three members:
-
-
Verify that all cluster Operators are available by running the following command:
Replacing a healthy etcd member with the Machine API¶
On clusters that access the Machine API but do not use a control plane machine set, you can replace a healthy control plane machine by deleting the corresponding Machine object. The Machine API provisions a replacement machine, and the etcd cluster Operator adds the new node as an etcd member.
Prerequisites
-
The cluster has access to the Machine API.
-
The cluster does not have a
ControlPlaneMachineSetresource. -
You have access to the cluster as a user with the
cluster-adminrole. -
You have taken an etcd backup. For more information, see "Backing up etcd data".
Warning
Take an etcd backup before you replace a healthy etcd member so that you can restore your cluster if any issues occur.
Procedure
-
List the control plane machines in your cluster by running the following command:
-
Identify the control plane machine that corresponds to the node that you want to replace.
-
Optional. If you are performing planned maintenance, cordon the node by running the following command:
Replace
<node_name>with the name of the node that you are replacing. -
Delete the control plane machine by running the following command:
Replace
<control_plane_machine_name>with the name of the control plane machine to delete.Warning
Delete only one control plane machine at a time. Deleting multiple control plane machines at the same time can cause etcd quorum loss.
A new machine is automatically provisioned after you delete the control plane machine.
-
Monitor the replacement by running the following commands until the new machine reaches the
Runningphase:
Verification
-
Verify etcd health by running the following commands:
-
Open a remote shell session to a control plane etcd pod:
Replace
<etcd_pod_name>with the name of a running etcd pod. -
Check endpoint health:
Expected output shows
is healthyfor each endpoint. -
List etcd members and verify that the cluster has three members:
-
-
Verify that all cluster Operators are available by running the following command:
Replacing a healthy etcd member by scaling up and scaling down¶
On bare-metal clusters that do not use a control plane machine set, replace a healthy control plane node by temporarily scaling the control plane to four nodes, and then removing the node that you want to replace.
Warning
Red Hat supports a cluster that has 4 or 5 control plane nodes only on bare-metal infrastructure.
Prerequisites
-
The cluster does not have a
ControlPlaneMachineSetresource. -
The cluster is installed on bare-metal infrastructure.
-
You have access to the cluster as a user with the
cluster-adminrole. -
You have taken an etcd backup. For more information, see "Backing up etcd data".
-
You have created a single control plane node that you intend to add to your cluster as a postinstallation task.
Warning
Take an etcd backup before you replace a healthy etcd member so that you can restore your cluster if any issues occur.
Procedure
-
Add the new control plane node to your cluster by following the steps in "Adding a control plane node to your cluster".
-
Verify that the new control plane node is in the
Readystate and that etcd has four members by running the following commands:Replace
<etcd_pod_name>with the name of a running etcd pod.Expected output shows four etcd members and
is healthyfor each endpoint. -
Remove the control plane node that you want to replace.
-
Optional. If you are performing planned maintenance, cordon the node by running the following command:
Replace
<node_name>with the name of the node that you are replacing. -
Delete the
BareMetalHostobject for the control plane node that you want to replace by running the following command:Replace
<node_name>with the name of the node that you are replacing. -
Delete the
Machineobject for the control plane node that you want to replace by running the following command:Replace
<machine_name>with the name of the machine that is associated with the node that you are replacing.Note
After you remove the
BareMetalHostandMachineobjects, the machine controller automatically deletes theNodeobject.
-
-
Monitor the cluster until the control plane returns to three nodes and etcd is healthy by running the following commands:
Verification
-
Verify etcd health by running the following commands:
-
Open a remote shell session to a control plane etcd pod:
-
Check endpoint health:
Expected output shows
is healthyfor each endpoint. -
List etcd members and verify that the cluster has three members:
-
-
Verify that all cluster Operators are available by running the following command:
Additional resources
- Backing up etcd data
- Replacing an unhealthy etcd member
- Restoring to an earlier cluster state
- Replacing a control plane node in a healthy cluster
- Quorum protection with machine lifecycle hooks
- Replacing a control plane machine
- Adding a control plane node to your cluster
- How to replace all master nodes in OpenShift Container Platform 4 (Red Hat Knowledgebase article)