Automated disaster recovery for a hosted cluster by using OADP¶
In hosted clusters on bare-metal or Amazon Web Services (AWS) platforms, you can automate some backup and restore steps by using the OpenShift API for Data Protection (OADP) Operator.
The process involves the following steps:
- Configuring OADP
- Defining a Data Protection Application (DPA)
- Backing up the data plane workload
- Backing up the control plane workload
- Restoring a hosted cluster by using OADP
Prerequisites to automate disaster recovery by using OADP¶
Ensure that you meet the prerequisites to automate disaster recovery for hosted control planes by using OADP.
The following prerequisites apply to the management cluster:
- You installed the OADP Operator. For more information, see "About installing OADP".
- You created a storage class.
- You have access to the cluster with
cluster-adminprivileges. - You have access to the OADP subscription through a catalog source.
- You have access to a cloud storage provider that is compatible with OADP, such as S3, Microsoft Azure, Google Cloud, or MinIO.
- In a disconnected environment, you have access to a self-hosted storage provider that is compatible with OADP, for example Red Hat OpenShift Data Foundation or MinIO.
- Your hosted control planes pods are up and running.
- You are using a supported version of OADP for your management cluster. For example, if your management cluster is on OpenShift Container Platform 4.20, you must use OADP version 1.5. For more information, see "Support for OpenShift API for Data Protection (OADP)".
Additional resources
- About installing OADP
- Red Hat OpenShift Data Foundation
- MinIO
- Support for OpenShift API for Data Protection (OADP)
Configuring OADP to automate disaster recovery for hosted control planes¶
Before you can automate disaster recovery by using OpenShift API for Data Protection (OADP), you need to configure it for your hosted control planes platform.
Procedure
- If your hosted cluster is on AWS, follow the steps in "Configuring the OpenShift API for Data Protection with AWS S3 compatible storage" to configure OADP.
- If your hosted cluster is on a bare metal, follow the steps in "Configuring the OpenShift API for Data Protection with Multicloud Object Gateway" to configure OADP.
Additional resources
- Configuring the OpenShift API for Data Protection with AWS S3 compatible storage
- Configuring the OpenShift API for Data Protection with Multicloud Object Gateway
Automation of the backup and restore process with a DPA¶
You can automate parts of the backup and restore process by using a Data Protection Application (DPA). When you use a DPA, the steps to pause and restart the reconciliation of resources are automated. The DPA defines information including backup locations and Velero pod configurations.
Creating a Data Protection Application for bare metal¶
Automate parts of the backup and restore process on bare metal by creating a Data Protection Application (DPA). A DPA defines information including backup locations and Velero pod configurations.
You can create a DPA by defining a DataProtectionApplication object.
Procedure
-
Create a manifest file similar to the following example:
apiVersion: oadp.openshift.io/v1alpha1 kind: DataProtectionApplication metadata: name: dpa-sample namespace: openshift-adp spec: backupLocations: - name: default velero: provider: aws default: true objectStorage: bucket: <bucket_name> prefix: <bucket_prefix> config: region: minio profile: "default" s3ForcePathStyle: "true" s3Url: "<bucket_url>" insecureSkipTLSVerify: "true" credential: key: cloud name: cloud-credentials default: true snapshotLocations: - velero: provider: aws config: region: minio profile: "default" credential: key: cloud name: cloud-credentials configuration: nodeAgent: enable: true uploaderType: kopia velero: defaultPlugins: - openshift - aws - csi - hypershift resourceTimeout: 2hspec.backupLocations.velero.providerspecifies the provider for Velero. If you are using bare metal and MinIO, you can useawsas the provider.spec.backupLocations.velero.objectStorage.bucketspecifies the bucket name; for example,oadp-backup.spec.backupLocations.velero.objectStorage.prefixspecifies the bucket prefix; for example,hcp.spec.backupLocations.velero.config.regionspecifies the bucket region. In this example, the region isminio, which is a storage provider that is compatible with the S3 API.spec.backupLocations.velero.config.s3Urlspecifies the URL of the S3 endpoint.spec.snapshotLocations.velero.providerspecifies the provider for Velero. If you are using bare metal and MinIO, you can useawsas the provider.spec.snapshotLocations.velero.config.regionspecifies the region. In this example, the region isminio, which is a storage provider that is compatible with the S3 API.spec.configuration.nodeAgent.uploaderTypespecifieskopiaas the uploader type. Theresticuploader type is deprecated for OADP 1.5 and later.
-
Create the DPA object by running the following command:
After you create the
DataProtectionApplicationobject, newvelerodeployment andnode-agentpods are created in theopenshift-adpnamespace.
Next steps
- Back up the data plane workload.
Creating a Data Protection Application for AWS¶
Automate parts of the backup and restore process on AWS by creating a Data Protection Application (DPA). A DPA defines information including backup locations and Velero pod configurations.
You can create a DPA by defining a DataProtectionApplication object.
Procedure
-
Create a manifest file similar to the following example:
apiVersion: oadp.openshift.io/v1alpha1 kind: DataProtectionApplication metadata: name: dpa-sample namespace: openshift-adp spec: backupLocations: - name: default velero: provider: aws default: true objectStorage: bucket: <bucket_name> prefix: <bucket_prefix> config: region: minio profile: "backupStorage" credential: key: cloud name: cloud-credentials snapshotLocations: - velero: provider: aws config: region: minio profile: "volumeSnapshot" credential: key: cloud name: cloud-credentials configuration: nodeAgent: enable: true uploaderType: kopia velero: defaultPlugins: - openshift - aws - csi - hypershift resourceTimeout: 2hspec.backupLocations.velero.objectStorage.bucketspecifies the bucket name; for example,oadp-backup.spec.backupLocations.velero.objectStorage.prefixspecifies the bucket prefix; for example,hcp.spec.backupLocations.velero.config.regionspecifies the bucket region. The bucket region in this example isminio, which is a storage provider that is compatible with the S3 API.spec.snapshotLocations.velero.config.regionspecifies the region. The region in this example isminio, which is a storage provider that is compatible with the S3 API.spec.configuration.nodeAgent.uploaderTypespecifieskopiaas the uploader type. Theresticuploader type is deprecated for OADP 1.5 and later.
-
Create the DPA resource by running the following command:
After you create the
DataProtectionApplicationobject, newvelerodeployment andnode-agentpods are created in theopenshift-adpnamespace.
Next steps
- Back up the data plane workload.
Backing up the data plane workload by using the OADP Operator¶
You can back up the data plane workload by using the OADP Operator.
However, if the data plane workload is not important, you can skip this procedure.
Procedure
- To back up the data plane workload, follow the steps in "Backing up applications".
Additional resources
Backing up the control plane workload¶
You can back up the control plane workload by creating the Backup custom resource (CR).
To monitor and observe the backup process, see "Observing the backup and restore process".
Procedure
-
Create a YAML file that defines the
BackupCR:apiVersion: velero.io/v1 kind: Backup metadata: name: <backup_resource_name> namespace: openshift-adp labels: velero.io/storage-location: default spec: hooks: {} includedNamespaces: - <hosted_cluster_namespace> - <hosted_control_plane_namespace> includedResources: - sa - role - rolebinding - pod - pvc - pv - bmh - configmap - infraenv - priorityclasses - pdb - agents - hostedcluster - nodepool - secrets - services - deployments - hostedcontrolplane - cluster - agentcluster - agentmachinetemplate - agentmachine - machinedeployment - machineset - machine - route - clusterdeployment excludedResources: [] storageLocation: default ttl: 2h0m0s snapshotMoveData: true datamover: "velero" defaultVolumesToFsBackup: false-
metadata.namespecifies the name for yourBackupresource. -
spec.includedNamespacesspecifies namespaces to back up objects from. You must replace<hosted_cluster_namespace>with the name of the hosted cluster namespace and replace<hosted_control_plane_namespace>with the name of the hosted control plane namespace. -
spec.includedResourcesincludes theinfraenvvalue. You must create theinfraenvresource in a separate namespace. Do not delete theinfraenvresource during the backup process. -
spec.snapshotMoveData: trueandspec.datamover: veleroenable the CSI volume snapshots and upload the control plane workload automatically to cloud storage. -
spec.defaultVolumesToFsBackupspecifies that thefs-backupbacking up method for persistent volumes (PVs) is not used.Note
If you want to use CSI volume snapshots, you must add the
backup.velero.io/backup-volumes-excludes=<pv_name>annotation to your PVs.
-
-
Apply the
BackupCR by running the following command:
Verification
-
Verify that the value of the
status.phaseisCompletedby running the following command:
Next steps
- Restore the hosted cluster by using OADP.
Restoring a hosted cluster by using OADP¶
You can restore the hosted cluster by creating the Restore custom resource (CR).
- If you are using an in-place update, the
InfraEnvresource does not need spare nodes. You need to re-provision the worker nodes from the new management cluster. - If you are using a replace update, you need some spare nodes for the
InfraEnvresource to deploy the worker nodes.
Warning
After you back up your hosted cluster, you must delete it to start the restoring process. To start node provisioning, you must back up workloads in the data plane before deleting the hosted cluster.
Prerequisites
- You completed the steps in Removing a cluster by using the console (RHACM documentation) to delete your hosted cluster.
- You completed the steps in Removing remaining resources after removing a cluster (RHACM documentation).
To monitor and observe the backup process, see "Observing the backup and restore process".
Procedure
-
Verify that no pods and persistent volume claims (PVCs) are present in the hosted control plane namespace by running the following command:
-
Create a YAML file that defines the
RestoreCR:apiVersion: velero.io/v1 kind: Restore metadata: name: <restore_resource_name> namespace: openshift-adp spec: backupName: <backup_resource_name> restorePVs: true existingResourcePolicy: update excludedResources: - nodes - events - events.events.k8s.io - backups.velero.io - restores.velero.io - resticrepositories.velero.io-
metadata.namespecifies the name for yourRestoreresource. -
spec.backupNamespecifies the name of yourBackupresource. -
spec.restorePVs: trueindicates the recovery of persistent volumes (PVs) and their pods. -
spec.existingResourcePolicy: updateensures that the existing objects are overwritten with the backed up content.Warning
You must create the
InfraEnvresource in a separate namespace. Do not delete theInfraEnvresource during the restore process. TheInfraEnvresource is mandatory for the new nodes to be reprovisioned.
-
-
Apply the
RestoreCR by running the following command: -
Verify if the value of the
status.phaseisCompletedby running the following command:
Observing the backup and restore process¶
When you use OpenShift API for Data Protection (OADP) to back up and restore a hosted cluster, you can monitor and observe the process.
Procedure
-
Observe the backup process by running the following command:
-
Observe the restore process by running the following command:
-
Observe the Velero logs by running the following command:
-
Observe the progress of all of the OADP objects by running the following command:
$ watch "echo BackupRepositories:;echo;oc get backuprepositories.velero.io -A;echo; echo BackupStorageLocations: ;echo; oc get backupstoragelocations.velero.io -A;echo;echo DataUploads: ;echo;oc get datauploads.velero.io -A;echo;echo DataDownloads: ;echo;oc get datadownloads.velero.io -n openshift-adp; echo;echo VolumeSnapshotLocations: ;echo;oc get volumesnapshotlocations.velero.io -A;echo;echo Backups:;echo;oc get backup -A; echo;echo Restores:;echo;oc get restore -A"
Using the OADP CLI to describe the Backup and Restore resources¶
When you use OpenShift API for Data Protection, you can get more details of the Backup and Restore resources by using the OADP command-line interface (CLI).
Procedure
-
Get details of your
Restorecustom resource (CR) by running the following command:Replace
<restore_resource_name>with the name of yourRestoreresource. -
Get details of your
BackupCR by running the following command:Replace
<backup_resource_name>with the name of yourBackupresource.