Backing up and restoring etcd on a hosted cluster¶
By backing up and restoring etcd on a hosted cluster, you can fix failures, such as corrupted or missing data in an etcd member of a three-node cluster. If members of the etcd cluster lose data or have a CrashLoopBackOff status, this approach helps prevent an etcd quorum loss.
Backing up etcd on a hosted cluster¶
Fix failures by taking a snapshot of etcd on a hosted cluster.
Prerequisites
-
The
ocandjqbinaries have been installed. -
For hosted control planes on AWS, the OIDC provider configuration must be accessible so that any necessary fixes can be completed after the restore process. See the following procedure for more information about applying any necessary fixes.
-
For hosted control planes on bare metal, the
InfraEnvresource must reside in a different namespace from the hosted control plane namespace. Do not delete theInfraEnvresource during the backup or restore process. -
Management cluster prerequisites:
- A valid
StorageClassresource is configured in the management cluster. - You have
cluster-adminaccess to the management cluster. - You have access to online storage that is compatible with OpenShift API for Data Protection (OADP) cloud storage providers, such as Amazon Web Services (AWS) S3, Microsoft Azure, Google Cloud, or MinIO. If you use S3 for backup storage, ensure that IAM roles and policies are configured. For more information, see "Configuring Amazon Web Services".
- Hosted control plane pods are accessible and functioning properly.
- You have access to the
openshift-adpsubscription through aCatalogSourceobject.
- A valid
-
Service publishing strategy prerequisites for hosted clusters:
- The
APIServerservice must have a fixed hostname. Otherwise, the restore process fails and nodes cannot rejoin the cluster. For hosted control planes on AWS, theAPIServerservice can also use aRouteservice publishing strategy with a fixed hostname. - For production environments, it is strongly recommended to configure all services with fixed hostnames. By having fixed hostnames, you can ensure full service continuity and DNS consistency during the restore process on a different management cluster.
- When you restore a hosted cluster to a different management cluster, all services in the hosted cluster must be configured with a fixed hostname in its
servicePublishingStrategyproperty. This requirement applies to all platforms. Restoring a hosted cluster to a different management cluster is a Technology Preview feature. Restoring a hosted cluster to its original management cluster is supported.
Warning
After you back up the hosted cluster, you must back up workloads in the data cluster and then delete the original hosted cluster so that the restore process can begin.
- The
Procedure
-
Set up environment variables for your hosted cluster by entering the following commands, replacing values as necessary:
-
Pause reconciliation of the hosted cluster by entering the following command, replacing values as necessary:
-
Take a snapshot of etcd by using one of the following methods:
-
Use a previously backed-up snapshot of etcd.
-
If you have an available etcd pod, take a snapshot from the active etcd pod by completing the following steps:
-
List etcd pods by entering the following command:
-
Take a snapshot of the pod database and save it locally to your machine by entering the following commands:
$ oc exec -n ${CONTROL_PLANE_NAMESPACE} -c etcd -t ${ETCD_POD} -- \ env ETCDCTL_API=3 /usr/bin/etcdctl \ --cacert /etc/etcd/tls/etcd-ca/ca.crt \ --cert /etc/etcd/tls/client/etcd-client.crt \ --key /etc/etcd/tls/client/etcd-client.key \ --endpoints=https://localhost:2379 \ snapshot save /var/lib/snapshot.db -
Verify that the snapshot is successful by entering the following command:
-
-
Make a local copy of the snapshot:
-
Copy the snapshot by entering the following command:
-
Copy the snapshot database from etcd persistent storage:
-
List etcd pods by entering the following command:
-
Find a pod that is running and set its name as the value of
ETCD_POD: ETCD_POD=etcd-0, and then copy its snapshot database by entering the following command:
-
-
-
Additional resources
Restoring etcd on a hosted cluster¶
Fix failures by restoring a snapshot of etcd on a hosted cluster.
Prerequisites
- You completed the steps in "Backing up etcd on a hosted cluster". Ensure that you meet the prerequisites listed in that procedure.
Warning
After you back up the hosted cluster, you must back up workloads in the data cluster and then delete the original hosted cluster so that the restore process can begin.
Procedure
-
If you are working in a new terminal session from the session you used to complete the steps in "Backing up etcd on a hosted cluster", set the environment variables again as described in the backup procedure.
-
Scale down the etcd statefulset by entering the following command:
-
Delete volumes for second and third members by entering the following command:
-
Create a pod to access the first etcd member’s data:
-
Get the etcd image by entering the following command:
-
Create a pod that allows access to etcd data:
$ cat << EOF | oc apply -n ${CONTROL_PLANE_NAMESPACE} -f - apiVersion: apps/v1 kind: Deployment metadata: name: etcd-data spec: replicas: 1 selector: matchLabels: app: etcd-data template: metadata: labels: app: etcd-data spec: containers: - name: access image: $ETCD_IMAGE volumeMounts: - name: data mountPath: /var/lib command: - /usr/bin/bash args: - -c - |- while true; do sleep 1000 done volumes: - name: data persistentVolumeClaim: claimName: data-etcd-0 EOF -
Check the status of the
etcd-datapod and wait for it to be running by entering the following command: -
Get the name of the
etcd-datapod by entering the following command:
-
-
Copy an etcd snapshot into the pod by entering the following command:
-
Remove old data from the
etcd-datapod by entering the following commands: -
Restore the etcd snapshot by entering the following command:
$ oc exec -n ${CONTROL_PLANE_NAMESPACE} ${DATA_POD} -- \ etcdutl snapshot restore /var/lib/restored.snap.db \ --data-dir=/var/lib/data --skip-hash-check \ --name etcd-0 \ --initial-cluster-token=etcd-cluster \ --initial-cluster etcd-0=https://etcd-0.etcd-discovery.${CONTROL_PLANE_NAMESPACE}.svc:2380,etcd-1=https://etcd-1.etcd-discovery.${CONTROL_PLANE_NAMESPACE}.svc:2380,etcd-2=https://etcd-2.etcd-discovery.${CONTROL_PLANE_NAMESPACE}.svc:2380 \ --initial-advertise-peer-urls https://etcd-0.etcd-discovery.${CONTROL_PLANE_NAMESPACE}.svc:2380 -
Remove the temporary etcd snapshot from the pod by entering the following command:
-
Delete data access deployment by entering the following command:
-
Scale up the etcd cluster by entering the following command:
-
Wait for the etcd member pods to return and report as available by entering the following command:
-
Restore reconciliation of the hosted cluster by entering the following command:
$ oc patch -n ${HOSTED_CLUSTER_NAMESPACE} hostedclusters/${CLUSTER_NAME} \ -p '{"spec":{"pausedUntil":"null"}}' --type=mergeThis command uses the
"null"string. When you use that string, the controller treats unrecognized strings as not paused, but it logs an error. Instead of"null", you can also use"false", which is valid per Common Expression Language (CEL) validation, or JSONnull, which removes the field. -
Manually roll out the hosted cluster by entering the following command:
$ oc annotate hostedcluster -n \ <hosted_cluster_namespace> <hosted_cluster_name> \ hypershift.openshift.io/restart-date=$(date --iso-8601=seconds)The Multus admission controller and network node identity pods do not start yet.
-
Delete the pods for the second and third members of etcd and their PVCs by entering the following commands:
-
Manually roll out the hosted cluster again by entering the following command:
$ oc annotate hostedcluster -n \ <hosted_cluster_namespace> <hosted_cluster_name> \ hypershift.openshift.io/restart-date=$(date --iso-8601=seconds) \ --overwriteAfter a few minutes, the control plane pods start running.
-
If your hosted cluster is on AWS and you need to apply OIDC fixes after the restore process, enter the following command:
$ hcp fix dr-oidc-iam --hc-name <hosted_cluster_name> --hc-namespace <hosted_cluster_namespace> --aws-creds ~/.aws/credentialsThis command regenerates the OIDC in S3 in case OIDC is deleted.