Updating managed clusters in a disconnected environment with PolicyGenTemplate resources and TALM
You can use the Topology Aware Lifecycle Manager (TALM) to manage the software lifecycle of managed clusters that you have deployed by using GitOps Zero Touch Provisioning (ZTP) and Topology Aware Lifecycle Manager (TALM). TALM uses Red Hat Advanced Cluster Management (RHACM) PolicyGenTemplate policies to manage and control changes applied to target clusters.
Using PolicyGenTemplate CRs to manage and deploy policies to managed clusters will be deprecated in an upcoming OpenShift Container Platform release. Equivalent and improved functionality is available using Red Hat Advanced Cluster Management (RHACM) and PolicyGenerator CRs.
For more information about PolicyGenerator resources, see the RHACM Integrating Policy Generator documentation.
Additional resources
- Configuring managed cluster policies by using PolicyGenerator resources
- Comparing RHACM PolicyGenerator and PolicyGenTemplate resource patching
- About the Topology Aware Lifecycle Manager
Setting up the disconnected environment
TALM can perform both platform and Operator updates.
You must mirror both the platform image and Operator images that you want to update to in your mirror registry before you can use TALM to update your disconnected clusters.
Procedure
- For platform updates, you must perform the following steps:
-
Mirror the required OpenShift Container Platform image repository. Ensure that the required platform image is mirrored by following the "Mirroring the OpenShift Container Platform image repository" procedure linked in the Additional resources. Save the contents of the
imageContentSourcessection in theimageContentSources.yamlfile: The following is example output:imageContentSources:- mirrors:- mirror-ocp-registry.ibmcloud.io.cpak:5000/openshift-release-dev/openshift4source: quay.io/openshift-release-dev/ocp-release- mirrors:- mirror-ocp-registry.ibmcloud.io.cpak:5000/openshift-release-dev/openshift4source: quay.io/openshift-release-dev/ocp-v4.0-art-dev -
Save the image signature of the required platform image that was mirrored. You must add the image signature to the
PolicyGenTemplateCR for platform updates. To get the image signature, perform the following steps:-
Specify the required OpenShift Container Platform tag by running the following command:
$ OCP_RELEASE_NUMBER=<release_version> -
Specify the architecture of the cluster by running the following command:
$ ARCHITECTURE=<cluster_architecture><cluster_architecture>specifies the architecture of the cluster, such asx86_64,aarch64,s390x, orppc64le.
-
Get the release image digest from Quay by running the following command
$ DIGEST="$(oc adm release info quay.io/openshift-release-dev/ocp-release:${OCP_RELEASE_NUMBER}-${ARCHITECTURE} | sed -n 's/Pull From: .*@//p')" -
Set the digest algorithm by running the following command:
$ DIGEST_ALGO="${DIGEST%%:*}" -
Set the digest signature by running the following command:
$ DIGEST_ENCODED="${DIGEST#*:}" -
Get the image signature from the mirror.openshift.com website by running the following command:
$ SIGNATURE_BASE64=$(curl -s "https://mirror.openshift.com/pub/openshift-v4/signatures/openshift/release/${DIGEST_ALGO}=${DIGEST_ENCODED}/signature-1" | base64 -w0 && echo) -
Save the image signature to the
checksum-<OCP_RELEASE_NUMBER>.yamlfile by running the following commands:$ cat >checksum-${OCP_RELEASE_NUMBER}.yaml <<EOF${DIGEST_ALGO}-${DIGEST_ENCODED}: ${SIGNATURE_BASE64}EOF
-
-
Prepare the update graph. You have two options to prepare the update graph:
- Use the OpenShift Update Service. For more information about how to set up the graph on the hub cluster, see Deploy the operator for OpenShift Update Service and Build the graph data init container.
- Make a local copy of the upstream graph. Host the update graph on an
httporhttpsserver in the disconnected environment that has access to the managed cluster. To download the update graph, use the following command:$ curl -s https://api.openshift.com/api/upgrades_info/v1/graph?channel=stable-4.22 -o ~/upgrade-graph_stable-4.22
-
- For Operator updates, you must perform the following task:
- Mirror the Operator catalogs. Ensure that the required Operator images are mirrored by following the procedure in the "Mirroring Operator catalogs for use with disconnected clusters" section.
Additional resources
- Upgrading GitOps ZTP
- Mirroring the OpenShift Container Platform image repository
- Mirroring Operator catalogs for use with disconnected clusters
- Preparing the disconnected environment
- Understanding update channels and releases
Performing a platform update with PolicyGenTemplate CRs
You can perform a platform update with the TALM.
Prerequisites
- Install the Topology Aware Lifecycle Manager (TALM).
- Update GitOps Zero Touch Provisioning (ZTP) to the latest version.
- Provision one or more managed clusters with GitOps ZTP.
- Mirror the required image repository.
- Log in as a user with
cluster-adminprivileges. - Create RHACM policies in the hub cluster.
Procedure
- Create a
PolicyGenTemplateCR for the platform update:-
Save the following
PolicyGenTemplateCR in thedu-upgrade.yamlfile: The following example shows thePolicyGenTemplateCR for platform update:apiVersion: ran.openshift.io/v1kind: PolicyGenTemplatemetadata:name: "du-upgrade"namespace: "ztp-group-du-sno"spec:bindingRules:group-du-sno: ""mcp: "master"remediationAction: informsourceFiles:- fileName: ImageSignature.yamlpolicyName: "platform-upgrade-prep"binaryData:${DIGEST_ALGO}-${DIGEST_ENCODED}: ${SIGNATURE_BASE64}- fileName: DisconnectedICSP.yamlpolicyName: "platform-upgrade-prep"metadata:name: disconnected-internal-icsp-for-ocpspec:repositoryDigestMirrors:- mirrors:- quay-intern.example.com/ocp4/openshift-release-devsource: quay.io/openshift-release-dev/ocp-release- mirrors:- quay-intern.example.com/ocp4/openshift-release-devsource: quay.io/openshift-release-dev/ocp-v4.0-art-dev- fileName: ClusterVersion.yamlpolicyName: "platform-upgrade"metadata:name: versionspec:channel: "stable-4.22"upstream: http://upgrade.example.com/images/upgrade-graph_stable-4.22desiredUpdate:version: 4.22.4status:history:- version: 4.22.4state: "Completed"ImageSignature.yaml- TheConfigMapCR contains the signature of the required release image to update to.${DIGEST_ALGO}-${DIGEST_ENCODED}: ${SIGNATURE_BASE64}- Shows the image signature of the required OpenShift Container Platform release. Get the signature from thechecksum-${OCP_RELEASE_NUMBER}.yamlfile you saved when following the procedures in the "Setting up the environment" section.repositoryDigestMirrors- Shows the mirror repository that contains the required OpenShift Container Platform image. Get the mirrors from theimageContentSources.yamlfile that you saved when following the procedures in the "Setting up the environment" section.ClusterVersion.yaml- Shows theClusterVersionCR to trigger the update. Thechannel,upstream, anddesiredVersionfields are all required for image precaching.
The
PolicyGenTemplateCR generates two policies:- The
du-upgrade-platform-upgrade-preppolicy does the preparation work for the platform update. It creates theConfigMapCR for the required release image signature, creates the image content source of the mirrored release image repository, and updates the cluster version with the required update channel and the update graph reachable by the managed cluster in the disconnected environment. - The
du-upgrade-platform-upgradepolicy is used to perform platform upgrade.
-
Add the
du-upgrade.yamlfile contents to thekustomization.yamlfile located in the GitOps ZTP Git repository for thePolicyGenTemplateCRs and push the changes to the Git repository. ArgoCD pulls the changes from the Git repository and generates the policies on the hub cluster. -
Check the created policies by running the following command:
$ oc get policies -A | grep platform-upgrade
-
- Create the
ClusterGroupUpdateCR for the platform update with thespec.enablefield set tofalse.- Save the content of the platform update
ClusterGroupUpdateCR with thedu-upgrade-platform-upgrade-prepand thedu-upgrade-platform-upgradepolicies and the target clusters to thecgu-platform-upgrade.ymlfile, as shown in the following example:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:name: cgu-platform-upgradenamespace: defaultspec:managedPolicies:- du-upgrade-platform-upgrade-prep- du-upgrade-platform-upgradepreCaching: falseclusters:- spoke1remediationStrategy:maxConcurrency: 1enable: false - Apply the
ClusterGroupUpdateCR to the hub cluster by running the following command:$ oc apply -f cgu-platform-upgrade.yml
- Save the content of the platform update
- Optional: Precache the images for the platform update.
- Enable precaching in the
ClusterGroupUpdateCR by running the following command:$ oc --namespace=default patch clustergroupupgrade.ran.openshift.io/cgu-platform-upgrade \--patch '{"spec":{"preCaching": true}}' --type=merge - Monitor the update process and wait for the pre-caching to complete. Check the status of pre-caching by running the following command on the hub cluster:
$ oc get cgu cgu-platform-upgrade -o jsonpath='{.status.precaching.status}'
- Enable precaching in the
- Start the platform update:
- Enable the
cgu-platform-upgradepolicy and disable pre-caching by running the following command:$ oc --namespace=default patch clustergroupupgrade.ran.openshift.io/cgu-platform-upgrade \--patch '{"spec":{"enable":true, "preCaching": false}}' --type=merge - Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
$ oc get policies --all-namespaces
- Enable the
Additional resources
Performing an Operator update with PolicyGenTemplate CRs
You can perform an Operator update with the TALM.
Prerequisites
- Install the Topology Aware Lifecycle Manager (TALM).
- Update GitOps Zero Touch Provisioning (ZTP) to the latest version.
- Provision one or more managed clusters with GitOps ZTP.
- Mirror the required index image, bundle images, and all Operator images referenced in the bundle images.
- Log in as a user with
cluster-adminprivileges. - Create RHACM policies in the hub cluster.
Procedure
- Update the
PolicyGenTemplateCR for the Operator update.-
Update the
du-upgradePolicyGenTemplateCR with the following additional contents in thedu-upgrade.yamlfile:apiVersion: ran.openshift.io/v1kind: PolicyGenTemplatemetadata:name: "du-upgrade"namespace: "ztp-group-du-sno"spec:bindingRules:group-du-sno: ""mcp: "master"remediationAction: informsourceFiles:- fileName: DefaultCatsrc.yamlremediationAction: informpolicyName: "operator-catsrc-policy"metadata:name: redhat-operators-disconnectedspec:displayName: Red Hat Operators Catalogimage: registry.example.com:5000/olm/redhat-operators-disconnected:v4.22updateStrategy:registryPoll:interval: 1hstatus:connectionState:lastObservedState: READYimage- The index image URL contains the required Operator images. If the index images are always pushed to the same image name and tag, this change is not needed.updateStrategy- Set how frequently the Operator Lifecycle Manager (OLM) polls the index image for new Operator versions with theregistryPoll.intervalfield. This change is not needed if a new index image tag is always pushed for y-stream and z-stream Operator updates. TheregistryPoll.intervalfield can be set to a shorter interval to expedite the update, however shorter intervals increase computational load. To counteract this behavior, you can restoreregistryPoll.intervalto the default value once the update is complete.lastObservedState- Last observed state of the catalog connection. TheREADYvalue ensures that theCatalogSourcepolicy is ready, indicating that the index pod is pulled and is running. This way, TALM upgrades the Operators based on up-to-date policy compliance states.
-
This update generates one policy,
du-upgrade-operator-catsrc-policy, to update theredhat-operators-disconnectedcatalog source with the new index images that contain the required Operators images.noteIf you want to use the image precaching for Operators and there are Operators from a different catalog source other than
redhat-operators-disconnected, you must perform the following tasks:- Prepare a separate catalog source policy with the new index image or registry poll interval update for the different catalog source.
- Prepare a separate subscription policy for the required Operators that are from the different catalog source.
For example, the required SRIOV-FEC Operator is available in the
certified-operatorscatalog source. To update the catalog source and the Operator subscription, add the following contents to generate two policies,du-upgrade-fec-catsrc-policyanddu-upgrade-subscriptions-fec-policy:apiVersion: ran.openshift.io/v1kind: PolicyGenTemplatemetadata:name: "du-upgrade"namespace: "ztp-group-du-sno"spec:bindingRules:group-du-sno: ""mcp: "master"remediationAction: informsourceFiles:# ...- fileName: DefaultCatsrc.yamlremediationAction: informpolicyName: "fec-catsrc-policy"metadata:name: certified-operatorsspec:displayName: Intel SRIOV-FEC Operatorimage: registry.example.com:5000/olm/far-edge-sriov-fec:v4.10updateStrategy:registryPoll:interval: 10m- fileName: AcceleratorsSubscription.yamlpolicyName: "subscriptions-fec-policy"spec:channel: "stable"source: certified-operators -
Remove the specified subscriptions channels in the common
PolicyGenTemplateCR, if they exist. The default subscriptions channels from the GitOps ZTP image are used for the update.noteThe default channel for the Operators applied through GitOps ZTP 4.22 is
stable, except for theperformance-addon-operator. As of OpenShift Container Platform 4.11, theperformance-addon-operatorfunctionality was moved to thenode-tuning-operator. For the 4.10 release, the default channel for PAO isv4.10. You can also specify the default channels in the commonPolicyGenTemplateCR. -
Push the
PolicyGenTemplateCRs updates to the GitOps ZTP Git repository. ArgoCD pulls the changes from the Git repository and generates the policies on the hub cluster. -
Check the created policies by running the following command:
$ oc get policies -A | grep -E "catsrc-policy|subscription"
-
- Apply the required catalog source updates before starting the Operator update.
- Save the content of the
ClusterGroupUpgradeCR namedoperator-upgrade-prepwith the catalog source policies and the target managed clusters to thecgu-operator-upgrade-prep.ymlfile:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:name: cgu-operator-upgrade-prepnamespace: defaultspec:clusters:- spoke1enable: truemanagedPolicies:- du-upgrade-operator-catsrc-policyremediationStrategy:maxConcurrency: 1 - Apply the policy to the hub cluster by running the following command:
$ oc apply -f cgu-operator-upgrade-prep.yml
- Monitor the update process. Upon completion, ensure that the policy is compliant by running the following command:
$ oc get policies -A | grep -E "catsrc-policy"
- Save the content of the
- Create the
ClusterGroupUpgradeCR for the Operator update with thespec.enablefield set tofalse.-
Save the content of the Operator update
ClusterGroupUpgradeCR with thedu-upgrade-operator-catsrc-policypolicy and the subscription policies created from the commonPolicyGenTemplateand the target clusters to thecgu-operator-upgrade.ymlfile, as shown in the following example:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:name: cgu-operator-upgradenamespace: defaultspec:managedPolicies:- du-upgrade-operator-catsrc-policy- common-subscriptions-policypreCaching: falseclusters:- spoke1remediationStrategy:maxConcurrency: 1enable: falsedu-upgrade-operator-catsrc-policyis needed by the image precaching feature to retrieve the Operator images from the catalog source.common-subscriptions-policycontains Operator subscriptions. If you have followed the structure and content of the referencePolicyGenTemplates, all Operator subscriptions are grouped into thecommon-subscriptions-policypolicy.
noteOne
ClusterGroupUpgradeCR can only precache the images of the required Operators defined in the subscription policy from one catalog source included in theClusterGroupUpgradeCR. If the required Operators are from different catalog sources, such as in the example of the SRIOV-FEC Operator, anotherClusterGroupUpgradeCR must be created withdu-upgrade-fec-catsrc-policyanddu-upgrade-subscriptions-fec-policypolicies for the SRIOV-FEC Operator images precaching and update. -
Apply the
ClusterGroupUpgradeCR to the hub cluster by running the following command:$ oc apply -f cgu-operator-upgrade.yml
-
- Optional: Precache the images for the Operator update.
-
Before starting image precaching, verify the subscription policy is
NonCompliantat this point by running the following command:$ oc get policy common-subscriptions-policy -n <policy_namespace>The following is example output:
NAME REMEDIATION ACTION COMPLIANCE STATE AGEcommon-subscriptions-policy inform NonCompliant 27d -
Enable precaching in the
ClusterGroupUpgradeCR by running the following command:$ oc --namespace=default patch clustergroupupgrade.ran.openshift.io/cgu-operator-upgrade \--patch '{"spec":{"preCaching": true}}' --type=merge -
Monitor the process and wait for the precaching to complete. Check the status of precaching by running the following command on the managed cluster:
$ oc get cgu cgu-operator-upgrade -o jsonpath='{.status.precaching.status}' -
Check if the precaching is completed before starting the update by running the following command:
$ oc get cgu -n default cgu-operator-upgrade -ojsonpath='{.status.conditions}' | jqThe following is example output:
[{"lastTransitionTime": "2022-03-08T20:49:08.000Z","message": "The ClusterGroupUpgrade CR is not enabled","reason": "UpgradeNotStarted","status": "False","type": "Ready"},{"lastTransitionTime": "2022-03-08T20:55:30.000Z","message": "Precaching is completed","reason": "PrecachingCompleted","status": "True","type": "PrecachingDone"}]
-
- Start the Operator update.
- Enable the
cgu-operator-upgradeClusterGroupUpgradeCR and disable precaching to start the Operator update by running the following command:$ oc --namespace=default patch clustergroupupgrade.ran.openshift.io/cgu-operator-upgrade \--patch '{"spec":{"enable":true, "preCaching": false}}' --type=merge - Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
$ oc get policies --all-namespaces
- Enable the
Additional resources
Troubleshooting missed Operator updates with PolicyGenTemplate CRs
In some scenarios, Topology Aware Lifecycle Manager (TALM) might miss Operator updates due to an out-of-date policy compliance state.
After a catalog source update, it takes time for the Operator Lifecycle Manager (OLM) to update the subscription status. The status of the subscription policy might continue to show as compliant while TALM decides whether remediation is needed. As a result, the Operator specified in the subscription policy does not get upgraded.
To avoid this scenario, add another catalog source configuration to the PolicyGenTemplate and specify this configuration in the subscription for any Operators that require an update.
Procedure
-
Add a catalog source configuration in the
PolicyGenTemplateresource:- fileName: DefaultCatsrc.yamlremediationAction: informpolicyName: "operator-catsrc-policy"metadata:name: redhat-operators-disconnectedspec:displayName: Red Hat Operators Catalogimage: registry.example.com:5000/olm/redhat-operators-disconnected:v{product-version}updateStrategy:registryPoll:interval: 1hstatus:connectionState:lastObservedState: READY- fileName: DefaultCatsrc.yamlremediationAction: informpolicyName: "operator-catsrc-policy"metadata:name: redhat-operators-disconnected-v2spec:displayName: Red Hat Operators Catalog v2image: registry.example.com:5000/olm/redhat-operators-disconnected:<version>updateStrategy:registryPoll:interval: 1hstatus:connectionState:lastObservedState: READYname- Update the name for the new configuration.displayName- Update the display name for the new configuration.image- Update the index image URL. ThisfileName.spec.imagefield overrides any configuration in theDefaultCatsrc.yamlfile.
-
Update the
Subscriptionresource to point to the new configuration for Operators that require an update:apiVersion: operators.coreos.com/v1alpha1kind: Subscriptionmetadata:name: operator-subscriptionnamespace: operator-namspace# ...spec:source: redhat-operators-disconnected-v2# ...redhat-operators-disconnected-v2specifies the name of the additional catalog source configuration that you defined in thePolicyGenTemplateresource.
Performing a platform and an Operator update together
You can perform a platform and an Operator update at the same time.
Prerequisites
- Install the Topology Aware Lifecycle Manager (TALM).
- Update GitOps Zero Touch Provisioning (ZTP) to the latest version.
- Provision one or more managed clusters with GitOps ZTP.
- Log in as a user with
cluster-adminprivileges. - Create RHACM policies in the hub cluster.
Procedure
- Create the
PolicyGenTemplateCR for the updates by following the steps described in the "Performing a platform update" and "Performing an Operator update" sections. - Apply the prep work for the platform and the Operator update.
- Save the content of the
ClusterGroupUpgradeCR with the policies for platform update preparation work, catalog source updates, and target clusters to thecgu-platform-operator-upgrade-prep.ymlfile, for example:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:name: cgu-platform-operator-upgrade-prepnamespace: defaultspec:managedPolicies:- du-upgrade-platform-upgrade-prep- du-upgrade-operator-catsrc-policyclusterSelector:- group-du-snoremediationStrategy:maxConcurrency: 10enable: true - Apply the
cgu-platform-operator-upgrade-prep.ymlfile to the hub cluster by running the following command:$ oc apply -f cgu-platform-operator-upgrade-prep.yml - Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
$ oc get policies --all-namespaces
- Save the content of the
- Create the
ClusterGroupUpdateCR for the platform and the Operator update with thespec.enablefield set tofalse.-
Save the contents of the platform and Operator update
ClusterGroupUpdateCR with the policies and the target clusters to thecgu-platform-operator-upgrade.ymlfile, as shown in the following example:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:name: cgu-du-upgradenamespace: defaultspec:managedPolicies:- du-upgrade-platform-upgrade- du-upgrade-operator-catsrc-policy- common-subscriptions-policypreCaching: trueclusterSelector:- group-du-snoremediationStrategy:maxConcurrency: 1enable: falsedu-upgrade-platform-upgradeis the platform update policy.du-upgrade-operator-catsrc-policyis the policy containing the catalog source information for the Operators to be updated. It is needed for the precaching feature to determine which Operator images to download to the managed cluster.common-subscriptions-policyis the policy to update the Operators.
-
Apply the
cgu-platform-operator-upgrade.ymlfile to the hub cluster by running the following command:$ oc apply -f cgu-platform-operator-upgrade.yml
-
- Optional: Precache the images for the platform and the Operator update.
- Enable precaching in the
ClusterGroupUpgradeCR by running the following command:$ oc --namespace=default patch clustergroupupgrade.ran.openshift.io/cgu-du-upgrade \--patch '{"spec":{"preCaching": true}}' --type=merge - Monitor the update process and wait for the precaching to complete. Check the status of precaching by running the following command on the managed cluster:
$ oc get jobs,pods -n openshift-talm-pre-cache
- Check if the precaching is completed before starting the update by running the following command:
$ oc get cgu cgu-du-upgrade -ojsonpath='{.status.conditions}'
- Enable precaching in the
- Start the platform and Operator update.
-
Enable the
cgu-du-upgradeClusterGroupUpgradeCR to start the platform and the Operator update by running the following command:$ oc --namespace=default patch clustergroupupgrade.ran.openshift.io/cgu-du-upgrade \--patch '{"spec":{"enable":true, "preCaching": false}}' --type=merge -
Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
$ oc get policies --all-namespacesnoteThe CRs for the platform and Operator updates can be created from the beginning by configuring the setting to
spec.enable: true. In this case, the update starts immediately after precaching completes and there is no need to manually enable the CR.Both precaching and the update create extra resources, such as policies, placement bindings, placement rules, managed cluster actions, and managed cluster view, to help complete the procedures. Setting the
afterCompletion.deleteObjectsfield totruedeletes all these resources after the updates complete.
-
Removing Performance Addon Operator subscriptions from deployed clusters with PolicyGenTemplate CRs
In earlier versions of OpenShift Container Platform, the Performance Addon Operator provided automatic, low latency performance tuning for applications. In OpenShift Container Platform 4.11 or later, these functions are part of the Node Tuning Operator.
Do not install the Performance Addon Operator on clusters running OpenShift Container Platform 4.11 or later. If you upgrade to OpenShift Container Platform 4.11 or later, the Node Tuning Operator automatically removes the Performance Addon Operator.
You need to remove any policies that create Performance Addon Operator subscriptions to prevent a re-installation of the Operator.
The reference DU profile includes the Performance Addon Operator in the PolicyGenTemplate CR truecommon-ranGen.yaml. To remove the subscription from deployed managed clusters, you must update truecommon-ranGen.yaml.
If you install Performance Addon Operator 4.10.3-5 or later on OpenShift Container Platform 4.11 or later, the Performance Addon Operator detects the cluster version and automatically hibernates to avoid interfering with the Node Tuning Operator functions. However, to ensure best performance, remove the Performance Addon Operator from your OpenShift Container Platform 4.11 clusters.
Prerequisites
- Create a Git repository where you manage your custom site configuration data. The repository must be accessible from the hub cluster and be defined as a source repository for ArgoCD.
- Update to OpenShift Container Platform 4.11 or later.
- Log in as a user with
cluster-adminprivileges.
Procedure
- Change the
complianceTypetomustnothavefor the Performance Addon Operator namespace, Operator group, and subscription in thetruecommon-ranGen.yamlfile.- fileName: PaoSubscriptionNS.yamlpolicyName: "subscriptions-policy"complianceType: mustnothave- fileName: PaoSubscriptionOperGroup.yamlpolicyName: "subscriptions-policy"complianceType: mustnothave- fileName: PaoSubscription.yamlpolicyName: "subscriptions-policy"complianceType: mustnothave - Merge the changes with your custom site repository and wait for the ArgoCD application to synchronize the change to the hub cluster. The status of the
common-subscriptions-policypolicy changes toNon-Compliant. - Apply the change to your target clusters by using the Topology Aware Lifecycle Manager. For more information about rolling out configuration changes, see the "Additional resources" section.
- Monitor the process. When the status of the
common-subscriptions-policypolicy for a target cluster isCompliant, the Performance Addon Operator has been removed from the cluster. Get the status of thecommon-subscriptions-policyby running the following command:$ oc get policy -n ztp-common common-subscriptions-policy - Delete the Performance Addon Operator namespace, Operator group and subscription CRs from
spec.sourceFilesin thetruecommon-ranGen.yamlfile. - Merge the changes with your custom site repository and wait for the ArgoCD application to synchronize the change to the hub cluster. The policy remains compliant.
Precaching user-specified images with TALM on single-node OpenShift clusters
You can precache application-specific workload images on single-node OpenShift clusters before updating your applications.
You can specify the configuration options for the precaching jobs by using the following custom resources (CR):
PreCachingConfigCRClusterGroupUpgradeCR
TALM derives the platform image from the ClusterVersion object in the managed policies. TALM derives Operator index images from CatalogSource objects that the managed policies reference.
All fields in the PreCachingConfig CR are optional.
The following example shows a PreCachingConfig CR:
apiVersion: ran.openshift.io/v1alpha1
kind: PreCachingConfig
metadata:
name: exampleconfig
namespace: exampleconfig-ns
spec:
overrides:
operatorsPackagesAndChannels:
- local-storage-operator: stable
- ptp-operator: stable
- sriov-network-operator: stable
spaceRequired: 30 Gi
excludePrecachePatterns:
- aws
- vsphere
additionalImages:
- quay.io/exampleconfig/application1@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef
- quay.io/exampleconfig/application2@sha256:3d5800123dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47adfaef
- quay.io/exampleconfig/applicationN@sha256:4fe1334adfafadsf987123adfffdaf1243340adfafdedga0991234afdadfsa09
overrides- Specifies Operator packages and channels to precache instead of the values that TALM derives from the managed policies. The only supported override isoperatorsPackagesAndChannels. TALM ignores the deprecatedplatformImageandoperatorsIndexesoverride fields if they are present.spaceRequired- Specifies the minimum required disk space on the cluster. If unspecified, TALM defines a default value for OpenShift Container Platform images. The disk space field must include an integer value and the storage unit. For example:40 GiB,200 MB,1 TiB.excludePrecachePatterns- Specifies the images to exclude from precaching based on image name matching.additionalImages- Specifies the list of additional images to precache.
The following example shows a ClusterGroupUpgrade CR with a PreCachingConfig CR reference:
apiVersion: ran.openshift.io/v1alpha1
kind: ClusterGroupUpgrade
metadata:
name: cgu
spec:
preCaching: true
preCachingConfigRef:
name: exampleconfig
namespace: exampleconfig-ns
preCachingset totrueenables the precaching job.preCachingConfigRef.namespecifies thePreCachingConfigCR that you want to use.preCachingConfigRef.namespacespecifies the namespace of thePreCachingConfigCR that you want to use.
Creating the custom resources for precaching
You must create the PreCachingConfig CR before or concurrently with the ClusterGroupUpgrade CR.
Procedure
-
Create the
PreCachingConfigCR with the list of additional images you want to precache.apiVersion: ran.openshift.io/v1alpha1kind: PreCachingConfigmetadata:name: exampleconfignamespace: defaultspec:# ...spaceRequired: 30GiadditionalImages:- quay.io/exampleconfig/application1@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef- quay.io/exampleconfig/application2@sha256:3d5800123dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47adfaef- quay.io/exampleconfig/applicationN@sha256:4fe1334adfafadsf987123adfffdaf1243340adfafdedga0991234afdadfsa09namespacemust be accessible to the hub cluster.spaceRequired- It is recommended to set the minimum disk space required field to ensure that there is sufficient storage space for the precached images.
-
Create a
ClusterGroupUpgradeCR with thepreCachingfield set totrueand specify thePreCachingConfigCR created in the previous step:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:name: cgunamespace: defaultspec:clusters:- sno1- sno2preCaching: truepreCachingConfigRef:- name: exampleconfignamespace: defaultmanagedPolicies:- du-upgrade-platform-upgrade- du-upgrade-operator-catsrc-policy- common-subscriptions-policyremediationStrategy:timeout: 240warningOnce you install the images on the cluster, you cannot change or delete them.
-
When you want to start precaching the images, apply the
ClusterGroupUpgradeCR by running the following command:$ oc apply -f cgu.yamlTALM verifies the
ClusterGroupUpgradeCR. From this point, you can continue with the TALM precaching workflow.noteAll sites are precached concurrently.
Verification
-
Check the precaching status on the hub cluster where the
ClusterGroupUpgradeCR is applied by running the following command:$ oc get cgu <cgu_name> -n <cgu_namespace> -oyamlThe following example shows the derived precaching specification. The
platformImageandoperatorsIndexesvalues come from the managed policies, not fromPreCachingConfigoverrides.precaching:spec:platformImage: quay.io/openshift-release-dev/ocp-release@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1efoperatorsIndexes:- registry.example.com:5000/custom-redhat-operators:1.0.0operatorsPackagesAndChannels:- local-storage-operator: stable- ptp-operator: stable- sriov-network-operator: stableexcludePrecachePatterns:- aws- vsphereadditionalImages:- quay.io/exampleconfig/application1@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef- quay.io/exampleconfig/application2@sha256:3d5800123dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47adfaef- quay.io/exampleconfig/applicationN@sha256:4fe1334adfafadsf987123adfffdaf1243340adfafdedga0991234afdadfsa09spaceRequired: "30"status:sno1: Startingsno2: StartingThe precaching configurations are validated by checking if the managed policies exist. Valid configurations of the
ClusterGroupUpgradeand thePreCachingConfigCRs result in the following statuses:The following example shows the output of valid CRs:
- lastTransitionTime: "2023-01-01T00:00:01Z"message: All selected clusters are validreason: ClusterSelectionCompletedstatus: "True"type: ClusterSelected- lastTransitionTime: "2023-01-01T00:00:02Z"message: Completed validationreason: ValidationCompletedstatus: "True"type: Validated- lastTransitionTime: "2023-01-01T00:00:03Z"message: Precaching spec is valid and consistentreason: PrecacheSpecIsWellFormedstatus: "True"type: PrecacheSpecValid- lastTransitionTime: "2023-01-01T00:00:04Z"message: Precaching in progress for 1 clustersreason: InProgressstatus: "False"type: PrecachingSucceededThe following example shows an invalid
PreCachingConfigCR:Type: "PrecacheSpecValid"Status: False,Reason: "PrecacheSpecIncomplete"Message: "Precaching spec is incomplete: failed to get PreCachingConfig resource due to PreCachingConfig.ran.openshift.io "<precaching_cr_name>" not found" -
You can find the precaching job by running the following command on the managed cluster:
$ oc get jobs -n openshift-talo-pre-cacheThe following example shows a precaching job in progress:
NAME COMPLETIONS DURATION AGEpre-cache 0/1 1s 1s -
You can check the status of the pod created for the precaching job by running the following command:
$ oc describe pod pre-cache -n openshift-talo-pre-cacheThe following example shows a precaching job in progress:
Type Reason Age From MessageNormal SuccesfulCreate 19s job-controller Created pod: pre-cache-abcd1 -
You can get live updates on the status of the job by running the following command:
$ oc logs -f pre-cache-abcd1 -n openshift-talo-pre-cache -
To verify the precache job is successfully completed, run the following command:
$ oc describe pod pre-cache -n openshift-talo-pre-cacheThe following example shows a completed precache job:
Type Reason Age From MessageNormal SuccesfulCreate 5m19s job-controller Created pod: pre-cache-abcd1Normal Completed 19s job-controller Job completed -
To verify that the images are successfully precached on the single-node OpenShift, do the following:
- Enter into the node in debug mode:
$ oc debug node/cnfdf00.example.lab
- Change root to
host:$ chroot /host/ - Search for the required images:
$ sudo podman images | grep <operator_name>
- Enter into the node in debug mode:
Additional resources
About the auto-created ClusterGroupUpgrade CR for GitOps ZTP
TALM has a controller called ManagedClusterForCGU that monitors the Ready state of the ManagedCluster CRs on the hub cluster and creates the ClusterGroupUpgrade CRs for GitOps Zero Touch Provisioning (ZTP).
For any managed cluster in the Ready state without a ztp-done label applied, the ManagedClusterForCGU controller automatically creates a ClusterGroupUpgrade CR in the ztp-install namespace with its associated RHACM policies that are created during the GitOps ZTP process. TALM then remediates the set of configuration policies that are listed in the auto-created ClusterGroupUpgrade CR to push the configuration CRs to the managed cluster.
If there are no policies for the managed cluster at the time when the cluster becomes Ready, a ClusterGroupUpgrade CR with no policies is created. Upon completion of the ClusterGroupUpgrade the managed cluster is labeled as ztp-done. If there are policies that you want to apply for that managed cluster, manually create a ClusterGroupUpgrade as a Day 2 operation.
Procedure
-
View the auto-created
ClusterGroupUpgradeCR for GitOps ZTP: The following example shows an auto-createdClusterGroupUpgradeCR for GitOps ZTP:apiVersion: ran.openshift.io/v1alpha1kind: ClusterGroupUpgrademetadata:generation: 1name: spoke1namespace: ztp-installownerReferences:- apiVersion: cluster.open-cluster-management.io/v1blockOwnerDeletion: truecontroller: truekind: ManagedClustername: spoke1uid: 98fdb9b2-51ee-4ee7-8f57-a84f7f35b9d5resourceVersion: "46666836"uid: b8be9cd2-764f-4a62-87d6-6b767852c7daspec:actions:afterCompletion:addClusterLabels:ztp-done: ""deleteClusterLabels:ztp-running: ""deleteObjects: truebeforeEnable:addClusterLabels:ztp-running: ""clusters:- spoke1enable: truemanagedPolicies:- common-spoke1-config-policy- common-spoke1-subscriptions-policy- group-spoke1-config-policy- spoke1-config-policy- group-spoke1-validator-du-policypreCaching: falseremediationStrategy:maxConcurrency: 1timeout: 240ztp-done: ""is applied to the managed cluster when TALM completes the cluster configuration.ztp-running: ""is applied to the managed cluster when TALM starts deploying the configuration policies.