Updating managed clusters in a disconnected environment with PolicyGenerator resources and TALM¶
You can use the Topology Aware Lifecycle Manager (TALM) to manage the software lifecycle of managed clusters that you have deployed using GitOps Zero Touch Provisioning (ZTP) and Topology Aware Lifecycle Manager (TALM). TALM uses Red Hat Advanced Cluster Management (RHACM) PolicyGenerator policies to manage and control changes applied to target clusters.
Setting up the disconnected environment¶
TALM can perform both platform and Operator updates.
You must mirror both the platform image and Operator images that you want to update to in your mirror registry before you can use TALM to update your disconnected clusters.
Procedure
-
For platform updates, you must perform the following steps:
-
Mirror the required OpenShift Container Platform image repository. Ensure that the required platform image is mirrored by following the "Mirroring the OpenShift Container Platform image repository" procedure linked in the Additional resources. Save the contents of the
imageContentSourcessection in theimageContentSources.yamlfile:The following is example output:
imageContentSources: - mirrors: - mirror-ocp-registry.ibmcloud.io.cpak:5000/openshift-release-dev/openshift4 source: quay.io/openshift-release-dev/ocp-release - mirrors: - mirror-ocp-registry.ibmcloud.io.cpak:5000/openshift-release-dev/openshift4 source: quay.io/openshift-release-dev/ocp-v4.0-art-dev -
Save the image signature of the required platform image that was mirrored. You must add the image signature to the
PolicyGeneratorCR for platform updates. To get the image signature, perform the following steps:-
Specify the required OpenShift Container Platform tag by running the following command:
-
Specify the architecture of the cluster by running the following command:
<cluster_architecture>specifies the architecture of the cluster, such asx86_64,aarch64,s390x, orppc64le.
-
Get the release image digest from Quay by running the following command
-
Set the digest algorithm by running the following command:
-
Set the digest signature by running the following command:
-
Get the image signature from the mirror.openshift.com website by running the following command:
-
Save the image signature to the
checksum-<OCP_RELEASE_NUMBER>.yamlfile by running the following commands:
-
-
Prepare the update graph. You have two options to prepare the update graph:
-
Use the OpenShift Update Service.
For more information about how to set up the graph on the hub cluster, see Deploy the operator for OpenShift Update Service and Build the graph data init container.
-
Make a local copy of the upstream graph. Host the update graph on an
httporhttpsserver in the disconnected environment that has access to the managed cluster. To download the update graph, use the following command:
-
-
-
For Operator updates, you must perform the following task:
- Mirror the Operator catalogs. Ensure that the required Operator images are mirrored by following the procedure in the "Mirroring Operator catalogs for use with disconnected clusters" section.
Additional resources
- About the Topology Aware Lifecycle Manager
- Upgrading GitOps ZTP
- Mirroring the OpenShift Container Platform image repository
- Mirroring Operator catalogs for use with disconnected clusters
- Preparing the disconnected environment
- Understanding update channels and releases
Performing a platform update with PolicyGenerator CRs¶
You can perform a platform update with the TALM.
Prerequisites
- Install the Topology Aware Lifecycle Manager (TALM).
- Update GitOps Zero Touch Provisioning (ZTP) to the latest version.
- Provision one or more managed clusters with GitOps ZTP.
- Mirror the required image repository.
- Log in as a user with
cluster-adminprivileges. - Create RHACM policies in the hub cluster.
Procedure
-
Create a
PolicyGeneratorCR for the platform update:-
Save the following
PolicyGeneratorCR in thedu-upgrade.yamlfile: The following example shows thePolicyGeneratorCR for platform update:apiVersion: policy.open-cluster-management.io/v1 kind: PolicyGenerator metadata: name: du-upgrade placementBindingDefaults: name: du-upgrade-placement-binding policyDefaults: namespace: ztp-group-du-sno placement: labelSelector: matchExpressions: - key: group-du-sno operator: Exists remediationAction: inform severity: low namespaceSelector: exclude: - kube-* include: - '*' evaluationInterval: compliant: 10m noncompliant: 10s policies: - name: du-upgrade-platform-upgrade policyAnnotations: ran.openshift.io/ztp-deploy-wave: "100" manifests: - path: source-crs/ClusterVersion.yaml patches: - metadata: name: version spec: channel: stable-4.22 desiredUpdate: version: 4.22.4 upstream: http://upgrade.example.com/images/upgrade-graph_stable-4.22 status: history: - state: Completed version: 4.22.4 - name: du-upgrade-platform-upgrade-prep policyAnnotations: ran.openshift.io/ztp-deploy-wave: "1" manifests: - path: source-crs/ImageSignature.yaml - path: source-crs/DisconnectedICSP.yaml patches: - metadata: name: disconnected-internal-icsp-for-ocp spec: repositoryDigestMirrors: - mirrors: - quay-intern.example.com/ocp4/openshift-release-dev source: quay.io/openshift-release-dev/ocp-release - mirrors: - quay-intern.example.com/ocp4/openshift-release-dev source: quay.io/openshift-release-dev/ocp-v4.0-art-devsource-crs/ClusterVersion.yaml- Shows theClusterVersionCR to trigger the update. Thechannel,upstream, anddesiredVersionfields are all required for image precaching.source-crs/ImageSignature.yaml- Contains the image signature of the required release image. The image signature is used to verify the image before applying the platform update.repositoryDigestMirrors- Shows the mirror repository that contains the required OpenShift Container Platform image. Get the mirrors from theimageContentSources.yamlfile that you saved when following the procedures in the "Setting up the environment" section.
The
PolicyGeneratorCR generates two policies:- The
du-upgrade-platform-upgrade-preppolicy does the preparation work for the platform update. It creates theConfigMapCR for the required release image signature, creates the image content source of the mirrored release image repository, and updates the cluster version with the required update channel and the update graph reachable by the managed cluster in the disconnected environment. - The
du-upgrade-platform-upgradepolicy is used to perform platform upgrade.
-
Add the
du-upgrade.yamlfile contents to thekustomization.yamlfile located in the GitOps ZTP Git repository for thePolicyGeneratorCRs and push the changes to the Git repository.ArgoCD pulls the changes from the Git repository and generates the policies on the hub cluster.
-
Check the created policies by running the following command:
-
-
Create the
ClusterGroupUpdateCR for the platform update with thespec.enablefield set tofalse.-
Save the content of the platform update
ClusterGroupUpdateCR with thedu-upgrade-platform-upgrade-prepand thedu-upgrade-platform-upgradepolicies and the target clusters to thecgu-platform-upgrade.ymlfile, as shown in the following example: -
Apply the
ClusterGroupUpdateCR to the hub cluster by running the following command:
-
-
Optional: Precache the images for the platform update.
-
Enable precaching in the
ClusterGroupUpdateCR by running the following command: -
Monitor the update process and wait for the pre-caching to complete. Check the status of pre-caching by running the following command on the hub cluster:
-
-
Start the platform update:
-
Enable the
cgu-platform-upgradepolicy and disable pre-caching by running the following command: -
Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
-
Additional resources
Performing an Operator update with PolicyGenerator CRs¶
You can perform an Operator update with the TALM.
Prerequisites
- Install the Topology Aware Lifecycle Manager (TALM).
- Update GitOps Zero Touch Provisioning (ZTP) to the latest version.
- Provision one or more managed clusters with GitOps ZTP.
- Mirror the required index image, bundle images, and all Operator images referenced in the bundle images.
- Log in as a user with
cluster-adminprivileges. - Create RHACM policies in the hub cluster.
Procedure
-
Update the
PolicyGeneratorCR for the Operator update.-
Update the
du-upgradePolicyGeneratorCR with the following additional contents in thedu-upgrade.yamlfile:apiVersion: policy.open-cluster-management.io/v1 kind: PolicyGenerator metadata: name: du-upgrade placementBindingDefaults: name: du-upgrade-placement-binding policyDefaults: namespace: ztp-group-du-sno placement: labelSelector: matchExpressions: - key: group-du-sno operator: Exists remediationAction: inform severity: low namespaceSelector: exclude: - kube-* include: - '*' evaluationInterval: compliant: 10m noncompliant: 10s policies: - name: du-upgrade-operator-catsrc-policy policyAnnotations: ran.openshift.io/ztp-deploy-wave: "1" manifests: - path: source-crs/DefaultCatsrc.yaml patches: - metadata: name: redhat-operators-disconnected spec: displayName: Red Hat Operators Catalog image: registry.example.com:5000/olm/redhat-operators-disconnected:v4.22 updateStrategy: registryPoll: interval: 1h status: connectionState: lastObservedState: READYimage- Contains the required Operator images. If the index images are always pushed to the same image name and tag, this change is not needed.updateStrategy- Sets how frequently the Operator Lifecycle Manager (OLM) polls the index image for new Operator versions with theregistryPoll.intervalfield. This change is not needed if a new index image tag is always pushed for y-stream and z-stream Operator updates. TheregistryPoll.intervalfield can be set to a shorter interval to expedite the update, however shorter intervals increase computational load. To counteract this, you can restoreregistryPoll.intervalto the default value once the update is complete.lastObservedState- Displays the observed state of the catalog connection. TheREADYvalue ensures that theCatalogSourcepolicy is ready, indicating that the index pod is pulled and is running. This way, TALM upgrades the Operators based on up-to-date policy compliance states.
-
This update generates one policy,
du-upgrade-operator-catsrc-policy, to update theredhat-operators-disconnectedcatalog source with the new index images that contain the required Operators images.Note
If you want to use the image precaching for Operators and there are Operators from a different catalog source other than
redhat-operators-disconnected, you must perform the following tasks:- Prepare a separate catalog source policy with the new index image or registry poll interval update for the different catalog source.
- Prepare a separate subscription policy for the required Operators that are from the different catalog source.
For example, the required SRIOV-FEC Operator is available in the
certified-operatorscatalog source. To update the catalog source and the Operator subscription, add the following contents to generate two policies,du-upgrade-fec-catsrc-policyanddu-upgrade-subscriptions-fec-policy:apiVersion: policy.open-cluster-management.io/v1 kind: PolicyGenerator metadata: name: du-upgrade placementBindingDefaults: name: du-upgrade-placement-binding policyDefaults: namespace: ztp-group-du-sno placement: labelSelector: matchExpressions: - key: group-du-sno operator: Exists remediationAction: inform severity: low namespaceSelector: exclude: - kube-* include: - '*' evaluationInterval: compliant: 10m noncompliant: 10s policies: - name: du-upgrade-fec-catsrc-policy policyAnnotations: ran.openshift.io/ztp-deploy-wave: "1" manifests: - path: source-crs/DefaultCatsrc.yaml patches: - metadata: name: certified-operators spec: displayName: Intel SRIOV-FEC Operator image: registry.example.com:5000/olm/far-edge-sriov-fec:v4.10 updateStrategy: registryPoll: interval: 10m - name: du-upgrade-subscriptions-fec-policy policyAnnotations: ran.openshift.io/ztp-deploy-wave: "2" manifests: - path: source-crs/AcceleratorsSubscription.yaml patches: - spec: channel: stable source: certified-operators -
Remove the specified subscriptions channels in the common
PolicyGeneratorCR, if they exist. The default subscriptions channels from the GitOps ZTP image are used for the update.Note
The default channel for the Operators applied through GitOps ZTP 4.22 is
stable, except for theperformance-addon-operator. As of OpenShift Container Platform 4.11, theperformance-addon-operatorfunctionality was moved to thenode-tuning-operator. For the 4.10 release, the default channel for PAO isv4.10. You can also specify the default channels in the commonPolicyGeneratorCR. -
Push the
PolicyGeneratorCRs updates to the GitOps ZTP Git repository.ArgoCD pulls the changes from the Git repository and generates the policies on the hub cluster.
-
Check the created policies by running the following command:
-
-
Apply the required catalog source updates before starting the Operator update.
-
Save the content of the
ClusterGroupUpgradeCR namedoperator-upgrade-prepwith the catalog source policies and the target managed clusters to thecgu-operator-upgrade-prep.ymlfile: -
Apply the policy to the hub cluster by running the following command:
-
Monitor the update process. Upon completion, ensure that the policy is compliant by running the following command:
-
-
Create the
ClusterGroupUpgradeCR for the Operator update with thespec.enablefield set tofalse.-
Save the content of the Operator update
ClusterGroupUpgradeCR with thedu-upgrade-operator-catsrc-policypolicy and the subscription policies created from the commonPolicyGeneratorand the target clusters to thecgu-operator-upgrade.ymlfile, as shown in the following example:apiVersion: ran.openshift.io/v1alpha1 kind: ClusterGroupUpgrade metadata: name: cgu-operator-upgrade namespace: default spec: managedPolicies: - du-upgrade-operator-catsrc-policy - common-subscriptions-policy preCaching: false clusters: - spoke1 remediationStrategy: maxConcurrency: 1 enable: falsedu-upgrade-operator-catsrc-policyis needed by the image precaching feature to retrieve the Operator images from the catalog source.common-subscriptions-policycontains Operator subscriptions. If you have followed the structure and content of the referencePolicyGenTemplates, all Operator subscriptions are grouped into thecommon-subscriptions-policypolicy.
Note
One
ClusterGroupUpgradeCR can only precache the images of the required Operators defined in the subscription policy from one catalog source included in theClusterGroupUpgradeCR. If the required Operators are from different catalog sources, such as in the example of the SRIOV-FEC Operator, anotherClusterGroupUpgradeCR must be created withdu-upgrade-fec-catsrc-policyanddu-upgrade-subscriptions-fec-policypolicies for the SRIOV-FEC Operator images precaching and update. -
Apply the
ClusterGroupUpgradeCR to the hub cluster by running the following command:
-
-
Optional: Precache the images for the Operator update.
-
Before starting image precaching, verify the subscription policy is
NonCompliantat this point by running the following command:The following is example output:
-
Enable precaching in the
ClusterGroupUpgradeCR by running the following command: -
Monitor the process and wait for the precaching to complete. Check the status of precaching by running the following command on the managed cluster:
-
Check if the precaching is completed before starting the update by running the following command:
The following is example output:
[ { "lastTransitionTime": "2022-03-08T20:49:08.000Z", "message": "The ClusterGroupUpgrade CR is not enabled", "reason": "UpgradeNotStarted", "status": "False", "type": "Ready" }, { "lastTransitionTime": "2022-03-08T20:55:30.000Z", "message": "Precaching is completed", "reason": "PrecachingCompleted", "status": "True", "type": "PrecachingDone" } ]
-
-
Start the Operator update.
-
Enable the
cgu-operator-upgradeClusterGroupUpgradeCR and disable precaching to start the Operator update by running the following command: -
Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
-
Additional resources
Troubleshooting missed Operator updates with PolicyGenerator CRs¶
In some scenarios, Topology Aware Lifecycle Manager (TALM) might miss Operator updates due to an out-of-date policy compliance state.
After a catalog source update, it takes time for the Operator Lifecycle Manager (OLM) to update the subscription status. The status of the subscription policy might continue to show as compliant while TALM decides whether remediation is needed. As a result, the Operator specified in the subscription policy does not get upgraded.
To avoid this scenario, add another catalog source configuration to the PolicyGenerator and specify this configuration in the subscription for any Operators that require an update.
Procedure
-
Add a catalog source configuration in the
PolicyGeneratorresource:manifests: - path: source-crs/DefaultCatsrc.yaml patches: - metadata: name: redhat-operators-disconnected spec: displayName: Red Hat Operators Catalog image: registry.example.com:5000/olm/redhat-operators-disconnected:v{product-version} updateStrategy: registryPoll: interval: 1h status: connectionState: lastObservedState: READY - path: source-crs/DefaultCatsrc.yaml patches: - metadata: name: redhat-operators-disconnected-v2 spec: displayName: Red Hat Operators Catalog v2 image: registry.example.com:5000/olm/redhat-operators-disconnected:<version> updateStrategy: registryPoll: interval: 1h status: connectionState: lastObservedState: READYname- Update the name for the new configuration.displayName- Update the display name for the new configuration.image- Update the index image URL. Thispolicies.manifests.patches.spec.imagefield overrides any configuration in theDefaultCatsrc.yamlfile.
-
Update the
Subscriptionresource to point to the new configuration for Operators that require an update:apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: operator-subscription namespace: operator-namspace # ... spec: source: redhat-operators-disconnected-v2 # ...redhat-operators-disconnected-v2specifies the name of the additional catalog source configuration that you defined in thePolicyGeneratorresource.
Performing a platform and an Operator update together¶
You can perform a platform and an Operator update at the same time.
Prerequisites
- Install the Topology Aware Lifecycle Manager (TALM).
- Update GitOps Zero Touch Provisioning (ZTP) to the latest version.
- Provision one or more managed clusters with GitOps ZTP.
- Log in as a user with
cluster-adminprivileges. - Create RHACM policies in the hub cluster.
Procedure
-
Create the
PolicyGeneratorCR for the updates by following the steps described in the "Performing a platform update" and "Performing an Operator update" sections. -
Apply the prep work for the platform and the Operator update.
-
Save the content of the
ClusterGroupUpgradeCR with the policies for platform update preparation work, catalog source updates, and target clusters to thecgu-platform-operator-upgrade-prep.ymlfile, for example:apiVersion: ran.openshift.io/v1alpha1 kind: ClusterGroupUpgrade metadata: name: cgu-platform-operator-upgrade-prep namespace: default spec: managedPolicies: - du-upgrade-platform-upgrade-prep - du-upgrade-operator-catsrc-policy clusterSelector: - group-du-sno remediationStrategy: maxConcurrency: 10 enable: true -
Apply the
cgu-platform-operator-upgrade-prep.ymlfile to the hub cluster by running the following command: -
Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
-
-
Create the
ClusterGroupUpdateCR for the platform and the Operator update with thespec.enablefield set tofalse.-
Save the contents of the platform and Operator update
ClusterGroupUpdateCR with the policies and the target clusters to thecgu-platform-operator-upgrade.ymlfile, as shown in the following example:apiVersion: ran.openshift.io/v1alpha1 kind: ClusterGroupUpgrade metadata: name: cgu-du-upgrade namespace: default spec: managedPolicies: - du-upgrade-platform-upgrade - du-upgrade-operator-catsrc-policy - common-subscriptions-policy preCaching: true clusterSelector: - group-du-sno remediationStrategy: maxConcurrency: 1 enable: falsedu-upgrade-platform-upgradeis the platform update policy.du-upgrade-operator-catsrc-policyis the policy containing the catalog source information for the Operators to be updated. It is needed for the precaching feature to determine which Operator images to download to the managed cluster.common-subscriptions-policyis the policy to update the Operators.
-
Apply the
cgu-platform-operator-upgrade.ymlfile to the hub cluster by running the following command:
-
-
Optional: Precache the images for the platform and the Operator update.
-
Enable precaching in the
ClusterGroupUpgradeCR by running the following command: -
Monitor the update process and wait for the precaching to complete. Check the status of precaching by running the following command on the managed cluster:
-
Check if the precaching is completed before starting the update by running the following command:
-
-
Start the platform and Operator update.
-
Enable the
cgu-du-upgradeClusterGroupUpgradeCR to start the platform and the Operator update by running the following command: -
Monitor the process. Upon completion, ensure that the policy is compliant by running the following command:
Note
The CRs for the platform and Operator updates can be created from the beginning by configuring the setting to
spec.enable: true. In this case, the update starts immediately after precaching completes and there is no need to manually enable the CR.Both precaching and the update create extra resources, such as policies, placement bindings, placement rules, managed cluster actions, and managed cluster view, to help complete the procedures. Setting the
afterCompletion.deleteObjectsfield totruedeletes all these resources after the updates complete.
-
Removing Performance Addon Operator subscriptions from deployed clusters with PolicyGenerator CRs¶
In earlier versions of OpenShift Container Platform, the Performance Addon Operator provided automatic, low latency performance tuning for applications. In OpenShift Container Platform 4.11 or later, these functions are part of the Node Tuning Operator.
Do not install the Performance Addon Operator on clusters running OpenShift Container Platform 4.11 or later. If you upgrade to OpenShift Container Platform 4.11 or later, the Node Tuning Operator automatically removes the Performance Addon Operator.
Note
You need to remove any policies that create Performance Addon Operator subscriptions to prevent a re-installation of the Operator.
The reference DU profile includes the Performance Addon Operator in the PolicyGenerator CR acm-common-ranGen.yaml. To remove the subscription from deployed managed clusters, you must update acm-common-ranGen.yaml.
Note
If you install Performance Addon Operator 4.10.3-5 or later on OpenShift Container Platform 4.11 or later, the Performance Addon Operator detects the cluster version and automatically hibernates to avoid interfering with the Node Tuning Operator functions. However, to ensure best performance, remove the Performance Addon Operator from your OpenShift Container Platform 4.11 clusters.
Prerequisites
- Create a Git repository where you manage your custom site configuration data. The repository must be accessible from the hub cluster and be defined as a source repository for ArgoCD.
- Update to OpenShift Container Platform 4.11 or later.
- Log in as a user with
cluster-adminprivileges.
Procedure
-
Change the
complianceTypetomustnothavefor the Performance Addon Operator namespace, Operator group, and subscription in theacm-common-ranGen.yamlfile. -
Merge the changes with your custom site repository and wait for the ArgoCD application to synchronize the change to the hub cluster. The status of the
common-subscriptions-policypolicy changes toNon-Compliant. -
Apply the change to your target clusters by using the Topology Aware Lifecycle Manager. For more information about rolling out configuration changes, see the "Additional resources" section.
-
Monitor the process. When the status of the
common-subscriptions-policypolicy for a target cluster isCompliant, the Performance Addon Operator has been removed from the cluster. Get the status of thecommon-subscriptions-policyby running the following command: -
Delete the Performance Addon Operator namespace, Operator group and subscription CRs from
policies.manifestsin theacm-common-ranGen.yamlfile. -
Merge the changes with your custom site repository and wait for the ArgoCD application to synchronize the change to the hub cluster. The policy remains compliant.
Precaching user-specified images with TALM on single-node OpenShift clusters¶
You can precache application-specific workload images on single-node OpenShift clusters before updating your applications.
You can specify the configuration options for the precaching jobs by using the following custom resources (CR):
PreCachingConfigCRClusterGroupUpgradeCR
TALM derives the platform image from the ClusterVersion object in the managed policies. TALM derives Operator index images from CatalogSource objects that the managed policies reference.
Note
All fields in the PreCachingConfig CR are optional.
The following example shows a PreCachingConfig CR:
apiVersion: ran.openshift.io/v1alpha1
kind: PreCachingConfig
metadata:
name: exampleconfig
namespace: exampleconfig-ns
spec:
overrides:
operatorsPackagesAndChannels:
- local-storage-operator: stable
- ptp-operator: stable
- sriov-network-operator: stable
spaceRequired: 30 Gi
excludePrecachePatterns:
- aws
- vsphere
additionalImages:
- quay.io/exampleconfig/application1@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef
- quay.io/exampleconfig/application2@sha256:3d5800123dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47adfaef
- quay.io/exampleconfig/applicationN@sha256:4fe1334adfafadsf987123adfffdaf1243340adfafdedga0991234afdadfsa09
overrides- Specifies Operator packages and channels to precache instead of the values that TALM derives from the managed policies. The only supported override isoperatorsPackagesAndChannels. TALM ignores the deprecatedplatformImageandoperatorsIndexesoverride fields if they are present.spaceRequired- Specifies the minimum required disk space on the cluster. If unspecified, TALM defines a default value for OpenShift Container Platform images. The disk space field must include an integer value and the storage unit. For example:40 GiB,200 MB,1 TiB.excludePrecachePatterns- Specifies the images to exclude from precaching based on image name matching.additionalImages- Specifies the list of additional images to precache.
The following example shows a ClusterGroupUpgrade CR with a PreCachingConfig CR reference:
apiVersion: ran.openshift.io/v1alpha1
kind: ClusterGroupUpgrade
metadata:
name: cgu
spec:
preCaching: true
preCachingConfigRef:
name: exampleconfig
namespace: exampleconfig-ns
preCachingset totrueenables the precaching job.preCachingConfigRef.namespecifies thePreCachingConfigCR that you want to use.preCachingConfigRef.namespacespecifies the namespace of thePreCachingConfigCR that you want to use.
Creating the custom resources for precaching¶
You must create the PreCachingConfig CR before or concurrently with the ClusterGroupUpgrade CR.
Procedure
-
Create the
PreCachingConfigCR with the list of additional images you want to precache.apiVersion: ran.openshift.io/v1alpha1 kind: PreCachingConfig metadata: name: exampleconfig namespace: default spec: # ... spaceRequired: 30Gi additionalImages: - quay.io/exampleconfig/application1@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef - quay.io/exampleconfig/application2@sha256:3d5800123dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47adfaef - quay.io/exampleconfig/applicationN@sha256:4fe1334adfafadsf987123adfffdaf1243340adfafdedga0991234afdadfsa09namespacemust be accessible to the hub cluster.spaceRequired- It is recommended to set the minimum disk space required field to ensure that there is sufficient storage space for the precached images.
-
Create a
ClusterGroupUpgradeCR with thepreCachingfield set totrueand specify thePreCachingConfigCR created in the previous step:apiVersion: ran.openshift.io/v1alpha1 kind: ClusterGroupUpgrade metadata: name: cgu namespace: default spec: clusters: - sno1 - sno2 preCaching: true preCachingConfigRef: - name: exampleconfig namespace: default managedPolicies: - du-upgrade-platform-upgrade - du-upgrade-operator-catsrc-policy - common-subscriptions-policy remediationStrategy: timeout: 240Warning
Once you install the images on the cluster, you cannot change or delete them.
-
When you want to start precaching the images, apply the
ClusterGroupUpgradeCR by running the following command:TALM verifies the
ClusterGroupUpgradeCR. From this point, you can continue with the TALM precaching workflow.Note
All sites are precached concurrently.
Verification
-
Check the precaching status on the hub cluster where the
ClusterGroupUpgradeCR is applied by running the following command:The following example shows the derived precaching specification. The
platformImageandoperatorsIndexesvalues come from the managed policies, not fromPreCachingConfigoverrides.precaching: spec: platformImage: quay.io/openshift-release-dev/ocp-release@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef operatorsIndexes: - registry.example.com:5000/custom-redhat-operators:1.0.0 operatorsPackagesAndChannels: - local-storage-operator: stable - ptp-operator: stable - sriov-network-operator: stable excludePrecachePatterns: - aws - vsphere additionalImages: - quay.io/exampleconfig/application1@sha256:3d5800990dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47e2e1ef - quay.io/exampleconfig/application2@sha256:3d5800123dee7cd4727d3fe238a97e2d2976d3808fc925ada29c559a47adfaef - quay.io/exampleconfig/applicationN@sha256:4fe1334adfafadsf987123adfffdaf1243340adfafdedga0991234afdadfsa09 spaceRequired: "30" status: sno1: Starting sno2: StartingThe precaching configurations are validated by checking if the managed policies exist. Valid configurations of the
ClusterGroupUpgradeand thePreCachingConfigCRs result in the following statuses:The following example shows the output of valid CRs:
- lastTransitionTime: "2023-01-01T00:00:01Z" message: All selected clusters are valid reason: ClusterSelectionCompleted status: "True" type: ClusterSelected - lastTransitionTime: "2023-01-01T00:00:02Z" message: Completed validation reason: ValidationCompleted status: "True" type: Validated - lastTransitionTime: "2023-01-01T00:00:03Z" message: Precaching spec is valid and consistent reason: PrecacheSpecIsWellFormed status: "True" type: PrecacheSpecValid - lastTransitionTime: "2023-01-01T00:00:04Z" message: Precaching in progress for 1 clusters reason: InProgress status: "False" type: PrecachingSucceededThe following example shows an invalid
PreCachingConfigCR: -
You can find the precaching job by running the following command on the managed cluster:
The following example shows a precaching job in progress:
-
You can check the status of the pod created for the precaching job by running the following command:
The following example shows a precaching job in progress:
-
You can get live updates on the status of the job by running the following command:
-
To verify the precache job is successfully completed, run the following command:
The following example shows a completed precache job:
-
To verify that the images are successfully precached on the single-node OpenShift, do the following:
-
Enter into the node in debug mode:
-
Change root to
host: -
Search for the required images:
-
Additional resources
About the auto-created ClusterGroupUpgrade CR for GitOps ZTP¶
TALM has a controller called ManagedClusterForCGU that monitors the Ready state of the ManagedCluster CRs on the hub cluster and creates the ClusterGroupUpgrade CRs for GitOps Zero Touch Provisioning (ZTP).
For any managed cluster in the Ready state without a ztp-done label applied, the ManagedClusterForCGU controller automatically creates a ClusterGroupUpgrade CR in the ztp-install namespace with its associated RHACM policies that are created during the GitOps ZTP process. TALM then remediates the set of configuration policies that are listed in the auto-created ClusterGroupUpgrade CR to push the configuration CRs to the managed cluster.
If there are no policies for the managed cluster at the time when the cluster becomes Ready, a ClusterGroupUpgrade CR with no policies is created. Upon completion of the ClusterGroupUpgrade the managed cluster is labeled as ztp-done. If there are policies that you want to apply for that managed cluster, manually create a ClusterGroupUpgrade as a Day 2 operation.
Procedure
-
View the auto-created
ClusterGroupUpgradeCR for GitOps ZTP:The following example shows an auto-created
ClusterGroupUpgradeCR for GitOps ZTP:apiVersion: ran.openshift.io/v1alpha1 kind: ClusterGroupUpgrade metadata: generation: 1 name: spoke1 namespace: ztp-install ownerReferences: - apiVersion: cluster.open-cluster-management.io/v1 blockOwnerDeletion: true controller: true kind: ManagedCluster name: spoke1 uid: 98fdb9b2-51ee-4ee7-8f57-a84f7f35b9d5 resourceVersion: "46666836" uid: b8be9cd2-764f-4a62-87d6-6b767852c7da spec: actions: afterCompletion: addClusterLabels: ztp-done: "" deleteClusterLabels: ztp-running: "" deleteObjects: true beforeEnable: addClusterLabels: ztp-running: "" clusters: - spoke1 enable: true managedPolicies: - common-spoke1-config-policy - common-spoke1-subscriptions-policy - group-spoke1-config-policy - spoke1-config-policy - group-spoke1-validator-du-policy preCaching: false remediationStrategy: maxConcurrency: 1 timeout: 240ztp-done: ""is applied to the managed cluster when TALM completes the cluster configuration.ztp-running: ""is applied to the managed cluster when TALM starts deploying the configuration policies.