Performing an image-based upgrade for single-node OpenShift clusters using GitOps ZTP
You can use a single resource on the hub cluster, the ImageBasedGroupUpgrade custom resource (CR), to manage an imaged-based upgrade on a selected group of managed clusters through all stages. Topology Aware Lifecycle Manager (TALM) reconciles the ImageBasedGroupUpgrade CR and creates the underlying resources to complete the defined stage transitions, either in a manually controlled or a fully automated upgrade flow.
For more information about the image-based upgrade, see "Understanding the image-based upgrade for single-node OpenShift clusters".
Additional resources
Managing the image-based upgrade at scale using the ImageBasedGroupUpgrade CR on the hub
The ImageBasedGroupUpgrade CR combines the ImageBasedUpgrade and ClusterGroupUpgrade APIs. For example, you can define the cluster selection and rollout strategy with the ImageBasedGroupUpgrade API in the same way as the ClusterGroupUpgrade API. The stage transitions are different from the ImageBasedUpgrade API. You can use the ImageBasedGroupUpgrade API to combine several stage transitions, also called actions, into one step that share one rollout strategy.
**Example **ImageBasedGroupUpgrade.yaml
apiVersion: lcm.openshift.io/v1alpha1
kind: ImageBasedGroupUpgrade
metadata:
name: <filename>
namespace: default
spec:
clusterLabelSelectors:
- matchExpressions:
- key: name
operator: In
values:
- spoke1
- spoke4
- spoke6
ibuSpec:
seedImageRef:
image: quay.io/seed/image:4.22.0
version: 4.22.0
pullSecretRef:
name: "<seed_pull_secret>"
extraManifests:
- name: example-extra-manifests
namespace: openshift-lifecycle-agent
oadpContent:
- name: oadp-cm
namespace: openshift-adp
plan:
- actions: ["Prep", "Upgrade", "FinalizeUpgrade"]
rolloutStrategy:
maxConcurrency: 200
timeout: 2400
Where:
clusterLabelSelectors: Clusters to upgrade.seedImageRef: Target platform version, the seed image, and the secret required to access the image.
If you add the seed image pull secret in the hub cluster, in the same namespace as the ImageBasedGroupUpgrade resource, the secret is added to the manifest list for the Prep stage. The secret is recreated in each spoke cluster in the openshift-lifecycle-agent namespace.
extraManifests: Optional: Applies additional manifests, which are not in the seed image, to the target cluster. Also appliesConfigMapobjects for custom catalog sources.oadpContent:ConfigMapresources that contain the OADPBackupandRestoreCRs.plan: Upgrade plan details.maxConcurrency: Number of clusters to update in a batch.timeout: Timeout limit to complete the action in minutes.
Supported action combinations
Actions are the list of stage transitions that TALM completes in the steps of an upgrade plan for the selected group of clusters. Each action entry in the ImageBasedGroupUpgrade CR is a separate step and a step has one or several actions that share the same rollout strategy. You can achieve more control over the rollout strategy for each action by separating actions into steps.
You can combine these actions differently in your upgrade plan and you can add the next steps later. Wait until the earlier steps either complete or fail before adding a step to your plan. The first action of an added step for clusters that failed a earlier steps must be either Abort or Rollback.
You cannot remove actions or steps from an ongoing plan.
The following table shows example plans for different levels of control over the rollout strategy:
Example upgrade plans
| Example plan | Description |
|---|---|
|
All actions share the same strategy |
|
Some actions share the same strategy |
|
All actions have different strategies |
Clusters that fail one of the actions will skip the remaining actions in the same step.
The ImageBasedGroupUpgrade API accepts the following actions:
Prep- Start preparing the upgrade resources by moving to the
Prepstage. Upgrade- Start the upgrade by moving to the
Upgradestage. FinalizeUpgrade- Complete the upgrade on selected clusters that completed the
Upgradeaction by moving to theIdlestage. Rollback- Start a rollback only on successfully upgraded clusters by moving to the
Rollbackstage. FinalizeRollback- Complete the rollback by moving to the
Idlestage. AbortOnFailure- Cancel the upgrade on selected clusters that failed the
PreporUpgradeactions by moving to theIdlestage. Abort- Cancel an ongoing upgrade only on clusters that are not yet upgraded by moving to the
Idlestage.
The following action combinations are supported. A pair of brackets signifies one step in the plan section:
["Prep"],["Abort"]["Prep", "Upgrade", "FinalizeUpgrade"]["Prep"],["AbortOnFailure"],["Upgrade"],["AbortOnFailure"],["FinalizeUpgrade"]["Rollback", "FinalizeRollback"]
Use one of the following combinations when you need to resume or cancel an ongoing upgrade from a completely new ImageBasedGroupUpgrade CR:
["Upgrade","FinalizeUpgrade"]["FinalizeUpgrade"]["FinalizeRollback"]["Abort"]["AbortOnFailure"]
Labeling for cluster selection
Use the spec.clusterLabelSelectors field for initial cluster selection. In addition, TALM labels the managed clusters according to the results of their last stage transition.
When a stage completes or fails, TALM marks the relevant clusters with the following labels:
lcm.openshift.io/ibgu-<stage>-completedlcm.openshift.io/ibgu-<stage>-failed
Use these cluster labels to cancel or roll back an upgrade on a group of clusters after troubleshooting the issues.
If you are using the ImageBasedGroupUpgrade CR to upgrade your clusters, ensure that you update the lcm.openshift.io/ibgu-<stage>-completed or lcm.openshift.io/ibgu-<stage>-failed cluster labels properly after performing troubleshooting or recovery steps on the managed clusters. This ensures that the TALM continues to manage the image-based upgrade for the cluster.
For example, if you want to cancel the upgrade for all managed clusters except for clusters that successfully completed the upgrade, you can add an Abort action to your plan. The Abort action moves back the ImageBasedUpgrade CR to the Idle stage, which cancels the upgrade on clusters that are not yet upgraded. Adding a separate Abort action ensures that the TALM does not perform the Abort action on clusters that have the lcm.openshift.io/ibgu-upgrade-completed label.
The TALM removes the cluster labels after successfully canceling or finalizing the upgrade.
Status monitoring
The ImageBasedGroupUpgrade CR ensures a better monitoring experience by aggregating status reporting for all clusters in one place. You can monitor the following actions:
status.clusters.completedActions- Shows all completed actions defined in the
plansection. status.clusters.currentAction- Shows all actions that are currently in progress.
status.clusters.failedActions- Shows all failed actions along with a detailed error message.
Performing an image-based upgrade on managed clusters at scale in several steps
For use cases when you need better control of when the upgrade interrupts your service, you can upgrade a set of your managed clusters by using the ImageBasedGroupUpgrade CR. You can use the ImageBasedGroupUpgrade CR to add actions after the earlier step is complete. After evaluating the results of the earlier steps, you can move to the next upgrade stage or troubleshoot any failed steps throughout the procedure.
Only certain action combinations are supported and listed in Supported action combinations.
Prerequisites
- You have logged in to the hub cluster as a user with
cluster-adminprivileges. - You have created policies and
ConfigMapobjects for resources used in the image-based upgrade. - You have installed the Lifecycle Agent and OADP Operators on all managed clusters through the hub cluster.
Procedure
-
Create a YAML file on the hub cluster that has the
ImageBasedGroupUpgradeCR:apiVersion: lcm.openshift.io/v1alpha1kind: ImageBasedGroupUpgrademetadata:name: <filename>namespace: defaultspec:clusterLabelSelectors:- matchExpressions:- key: nameoperator: Invalues:- spoke1- spoke4- spoke6ibuSpec:seedImageRef:image: quay.io/seed/image:4.16.0-rc.1version: 4.16.0-rc.1pullSecretRef:name: "<seed_pull_secret>"extraManifests:- name: example-extra-manifestsnamespace: openshift-lifecycle-agentoadpContent:- name: oadp-cmnamespace: openshift-adpplan:- actions: ["Prep"]rolloutStrategy:maxConcurrency: 2timeout: 2400Where:
clusterLabelSelectors: Clusters to upgrade.seedImageRef: Target platform version, the seed image, and the secret required to access the image.
noteIf you add the seed image pull secret in the hub cluster, in the same namespace as the
ImageBasedGroupUpgraderesource, the {lco} adds the secret to the manifest list for thePrepstage. The {lco} recreates the secret in each spoke cluster in theopenshift-lifecycle-agentnamespace.extraManifests: Optional: Applies additional manifests, which are not in the seed image, to the target cluster. Also appliesConfigMapobjects for custom catalog sources.oadpContent: List ofConfigMapresources that contain the OADPBackupandRestoreCRs.plan: Upgrade plan details.
-
Apply the created file by running the following command on the hub cluster:
$ oc apply -f <filename>.yaml -
Monitor the status updates by running the following command on the hub cluster:
$ oc get ibgu -o yamlExample output# ...status:clusters:- completedActions:- action: Prepname: spoke1- completedActions:- action: Prepname: spoke4- failedActions:- action: Prepname: spoke6# ...The earlier output of an example plan starts with the
Prepstage only and you add actions to the plan based on the results of the earlier step. The TALM adds a label to the clusters to mark if the upgrade succeeded or failed. For example, the TALM applies thelcm.openshift.io/ibgu-prep-failedlabel to clusters that failed thePrepstage.After investigating the failure, you can add the
AbortOnFailurestep to your upgrade plan. It moves the clusters labeled withlcm.openshift.io/ibgu-<action>-failedback to theIdlestage. The TALM deletes the resources that are related to the upgrade on the selected clusters. -
Optional: Add the
AbortOnFailureaction to your existingImageBasedGroupUpgradeCR by running the following command:$ oc patch ibgu <filename> --type=json -p \'[{"op": "add", "path": "/spec/plan/-", "value": {"actions": ["AbortOnFailure"], "rolloutStrategy": {"maxConcurrency": 5, "timeout": 10}}}]'- Continue monitoring the status updates by running the following command:
$ oc get ibgu -o yaml
- Continue monitoring the status updates by running the following command:
-
Add the action to your existing
ImageBasedGroupUpgradeCR by running the following command:$ oc patch ibgu <filename> --type=json -p \'[{"op": "add", "path": "/spec/plan/-", "value": {"actions": ["Upgrade"], "rolloutStrategy": {"maxConcurrency": 2, "timeout": 30}}}]' -
Optional: Add the
AbortOnFailureaction to your existingImageBasedGroupUpgradeCR by running the following command:$ oc patch ibgu <filename> --type=json -p \'[{"op": "add", "path": "/spec/plan/-", "value": {"actions": ["AbortOnFailure"], "rolloutStrategy": {"maxConcurrency": 5, "timeout": 10}}}]'- Continue monitoring the status updates by running the following command:
$ oc get ibgu -o yaml
- Continue monitoring the status updates by running the following command:
-
Add the action to your existing
ImageBasedGroupUpgradeCR by running the following command:$ oc patch ibgu <filename> --type=json -p \'[{"op": "add", "path": "/spec/plan/-", "value": {"actions": ["FinalizeUpgrade"], "rolloutStrategy": {"maxConcurrency": 10, "timeout": 3}}}]'
Verification
-
Monitor the status updates by running the following command:
$ oc get ibgu -o yamlExample output# ...status:clusters:- completedActions:- action: Prep- action: AbortOnFailurefailedActions:- action: Upgradename: spoke1- completedActions:- action: Prep- action: Upgrade- action: FinalizeUpgradename: spoke4- completedActions:- action: AbortOnFailurefailedActions:- action: Prepname: spoke6# ...
Additional resources
- Configuring a shared container partition between ostree stateroots when using GitOps ZTP
- Creating ConfigMap objects for the image-based upgrade with Lifecycle Agent using GitOps ZTP
- About backup and snapshot locations and their secrets
- Creating a Backup CR
- Creating a Restore CR
- Supported action combinations
Performing an image-based upgrade on managed clusters at scale in one step
For use cases when service interruption is not a concern, you can upgrade a set of your managed clusters by using the ImageBasedGroupUpgrade custom resource (CR). You can use the ImageBasedGroupUpgrade CR to combine several actions in one step with one rollout strategy. With one rollout strategy, you can reduce the upgrade time but you can only troubleshoot failed clusters after the upgrade plan is complete.
Prerequisites
- You have logged in to the hub cluster as a user with
cluster-adminprivileges. - You have created policies and
ConfigMapobjects for resources used in the image-based upgrade. - You have installed the Lifecycle Agent and OADP Operators on all managed clusters through the hub cluster.
Procedure
-
Create a YAML file on the hub cluster that has the
ImageBasedGroupUpgradeCR:apiVersion: lcm.openshift.io/v1alpha1kind: ImageBasedGroupUpgrademetadata:name: <filename>namespace: defaultspec:clusterLabelSelectors:- matchExpressions:- key: nameoperator: Invalues:- spoke1- spoke4- spoke6ibuSpec:seedImageRef:image: quay.io/seed/image:4.22.0version: 4.22.0pullSecretRef:name: "<seed_pull_secret>"extraManifests:- name: example-extra-manifestsnamespace: openshift-lifecycle-agentoadpContent:- name: oadp-cmnamespace: openshift-adpplan:- actions: ["Prep", "Upgrade", "FinalizeUpgrade"]rolloutStrategy:maxConcurrency: 200timeout: 2400Where:
clusterLabelSelectors: Clusters to upgrade.seedImageRef: Target platform version, the seed image, and the secret required to access the image.
noteIf you add the seed image pull secret in the hub cluster, in the same namespace as the
ImageBasedGroupUpgraderesource, the secret is added to the manifest list for thePrepstage. The secret is recreated in each spoke cluster in theopenshift-lifecycle-agentnamespace.extraManifests: Optional: Applies additional manifests, which are not in the seed image, to the target cluster. Also appliesConfigMapobjects for custom catalog sources.oadpContent:ConfigMapresources that contain the OADPBackupandRestoreCRs.plan: Upgrade plan details.maxConcurrency: Number of clusters to update in a batch.timeout: Timeout limit to complete the action in minutes.
-
Apply the created file by running the following command on the hub cluster:
$ oc apply -f <filename>.yaml
Verification
-
Monitor the status updates by running the following command:
$ oc get ibgu -o yamlExample output# ...status:clusters:- completedActions:- action: PrepfailedActions:- action: Upgradename: spoke1- completedActions:- action: Prep- action: Upgrade- action: FinalizeUpgradename: spoke4- failedActions:- action: Prepname: spoke6# ...