Deploying hosted control planes on IBM Power
You can deploy hosted control planes on IBM Power by configuring a cluster to function as a hosting cluster. This configuration provides an efficient and scalable solution for managing many clusters. The hosting cluster is an OpenShift Container Platform cluster that hosts control planes. The hosting cluster is also known as the management cluster.
The management cluster is not the managed cluster. A managed cluster is a cluster that the hub cluster manages.
The multicluster engine Operator supports only the default local-cluster, which is a managed hub cluster, and the hub cluster as the hosting cluster.
To provision hosted control planes on bare-metal infrastructure, you can use the Agent platform. The Agent platform uses the central infrastructure management service to add compute nodes to a hosted cluster. For more information, see "Enabling the central infrastructure management service".
You must start each IBM Power host with a Discovery image that the central infrastructure management provides. After each host starts, it runs an Agent process to discover the details of the host and completes the installation. An Agent custom resource represents each host.
When you create a hosted cluster with the Agent platform, HyperShift installs the Agent Cluster API provider in the hosted control plane namespace.
Prerequisites to configure hosted control planes on IBM Power
Ensure you meet the prerequisites to configure hosted control planes on IBM Power.
- The multicluster engine for Kubernetes Operator version 2.7 and later installed on an OpenShift Container Platform cluster. The multicluster engine Operator is automatically installed when you install Red Hat Advanced Cluster Management (RHACM). You can also install the multicluster engine Operator without RHACM as an Operator from the OpenShift Container Platform software catalog.
- The multicluster engine Operator must have at least one managed OpenShift Container Platform cluster. The
local-clustermanaged hub cluster is automatically imported in the multicluster engine Operator version 2.7 and later. For more information aboutlocal-cluster, see Advanced configuration in the RHACM documentation. You can check the status of your hub cluster by running the following command:$ oc get managedclusters local-cluster - You need a hosting cluster with at least 3 compute nodes to run the HyperShift Operator.
- You need to enable the central infrastructure management service. For more information, see "Enabling the central infrastructure management service".
- You need to install the hosted control planes command-line interface. For more information, see "Installing the hosted control plane command-line interface".
The hosted control planes feature is enabled by default. If you disabled the feature and want to manually enable the feature, see "Manually enabling the hosted control planes feature". If you need to disable the feature, see "Disabling the hosted control planes feature".
Additional resources
- Advanced configuration
- Enabling the central infrastructure management service
- Installing the hosted control plane command-line interface
- Manually enabling the hosted control planes feature
- Disabling the hosted control planes feature
IBM Power infrastructure requirements
The Agent platform does not create any infrastructure, but requires resources for infrastructure.
- Agents: An Agent represents a host that boots with a Discovery image and that you can provision as an OpenShift Container Platform node.
- DNS: The API and Ingress endpoints must be routable.
DNS configuration for hosted control planes on IBM Power
Clients outside the network can access the API server for the hosted cluster. A DNS entry must exist for the api.<hosted_cluster_name>.<basedomain> entry that points to the destination where the API server is reachable.
The DNS entry can be as simple as a record that points to one of the nodes in the managed cluster that runs the hosted control plane.
The entry can also point to a deployed load balancer to redirect incoming traffic to the ingress pods.
See the following example of a DNS configuration:
$ cat /var/named/<example.krnl.es.zone>
$ TTL 900
@ IN SOA bastion.example.krnl.es.com. hostmaster.example.krnl.es.com. (
2019062002
1D 1H 1W 3H )
IN NS bastion.example.krnl.es.com.
;
;
api IN A 1xx.2x.2xx.1xx
api-int IN A 1xx.2x.2xx.1xx
;
;
*.apps.<hosted_cluster_name>.<basedomain> IN A 1xx.2x.2xx.1xx
;
;EOF
The api record refers to the IP address of the API load balancer that handles ingress and egress traffic for hosted control planes.
For IBM Power, add IP addresses that correspond to the IP address of the agent.
compute-0 IN A 1xx.2x.2xx.1yy
compute-1 IN A 1xx.2x.2xx.1yy
Defining a custom DNS name
As a cluster administrator, you can create a hosted cluster with an external API DNS name that differs from the internal endpoint that gets used for node bootstraps and control plane communication.
You might want to define a different DNS name for the following reasons:
- To replace the user-facing TLS certificate with one from a public CA without breaking the control plane functions that bind to the internal root CA
- To support split-horizon DNS and NAT scenarios
- To ensure a similar experience to standalone control planes, where you can use functions, such as the
Show Login Commandfunction, with the correctkubeconfigand DNS configuration
You can define a DNS name either during your initial setup or during postinstallation operations, by entering a domain name in the kubeAPIServerDNSName parameter of a HostedCluster object.
Prerequisites
- You have a valid TLS certificate that covers the DNS name that you set in the
kubeAPIServerDNSNameparameter. - You have a resolvable DNS name URI that can reach and point to the correct address.
Procedure
-
In the specification for the
HostedClusterobject, add thekubeAPIServerDNSNameparameter and the address for the domain and specify which certificate to use, as shown in the following example:#...spec:configuration:apiServer:servingCerts:namedCertificates:- names:- xxx.example.com- yyy.example.comservingCertificate:name: <my_serving_certificate>kubeAPIServerDNSName: <custom_address>The value for the
kubeAPIServerDNSNameparameter must be a valid and addressable domain.After you define the
kubeAPIServerDNSNameparameter and specify the certificate, the Control Plane Operator controllers create akubeconfigfile namedcustom-admin-kubeconfig, where the file gets stored in theHostedControlPlanenamespace. The generation of certificates happen from the root CA, and theHostedControlPlanenamespace manages their expiration and renewal.The Control Plane Operator reports a new
kubeconfigfile namedCustomKubeconfigin theHostedControlPlanenamespace. That file uses the defined new server in thekubeAPIServerDNSNameparameter.A reference for the custom
kubeconfigfile exists in thestatusparameter asCustomKubeconfigof theHostedClusterobject. TheCustomKubeConfigparameter is optional, and you can add the parameter only if thekubeAPIServerDNSNameparameter is not empty. After you set theCustomKubeConfigparameter, the parameter triggers the generation of a secret named<hosted_cluster_name>-custom-admin-kubeconfigin theHostedClusternamespace. You can use the secret to access theHostedClusterAPI server. If you remove theCustomKubeConfigparameter during postinstallation operations, deletion of all related secrets and status references occur.noteDefining a custom DNS name does not directly impact the data plane, so no expected rollouts occur. The
HostedControlPlanenamespace receives the changes from the HyperShift Operator and deletes the corresponding parameters.If you remove the
kubeAPIServerDNSNameparameter from the specification for theHostedClusterobject, all newly generated secrets and theCustomKubeconfigreference are removed from the cluster and from thestatusparameter.
Creating a hosted cluster by using the CLI
On bare-metal infrastructure, you can import a hosted cluster or create one by using the command-line interface (CLI).
After you enable the Assisted Installer as an add-on to multicluster engine Operator and you create a hosted cluster with the Agent platform, the HyperShift Operator installs the Agent Cluster API provider in the hosted control plane namespace. The Agent Cluster API provider connects a management cluster that hosts the control plane and a hosted cluster that consists of only the compute nodes.
Prerequisites
-
Each hosted cluster must have a cluster-wide unique name. A hosted cluster name cannot be the same as any existing managed cluster. Otherwise, the multicluster engine Operator cannot manage the hosted cluster.
-
Do not use the word
clustersas a hosted cluster name. -
You cannot create a hosted cluster in the namespace of a multicluster engine Operator managed cluster.
-
For best security and management practices, create a hosted cluster separate from other hosted clusters.
-
Verify that you have a default storage class configured for your cluster. Otherwise, you might see pending persistent volume claims (PVCs).
-
By default when you use the
hcp create cluster agentcommand, the command creates a hosted cluster with configured node ports. The preferred publishing strategy for hosted clusters on bare metal exposes services through a load balancer. If you create a hosted cluster by using the web console or by using Red Hat Advanced Cluster Management, to set a publishing strategy for a service besides the Kubernetes API server, you must manually specify theservicePublishingStrategyinformation in theHostedClustercustom resource. -
Ensure that you meet the requirements described in "Requirements for hosted control planes on bare metal", which includes requirements related to infrastructure, firewalls, ports, and services. For example, those requirements describe how to add the appropriate zone labels to the bare-metal hosts in your management cluster, as shown in the following example commands:
$ oc label node [compute-node-1] topology.kubernetes.io/zone=zone1$ oc label node [compute-node-2] topology.kubernetes.io/zone=zone2$ oc label node [compute-node-3] topology.kubernetes.io/zone=zone3 -
Ensure that you have added bare-metal nodes to a hardware inventory.
Procedure
-
Create a namespace by entering the following command:
$ oc create ns <hosted_cluster_namespace>Replace
<hosted_cluster_namespace>with an identifier for your hosted cluster namespace. The HyperShift Operator creates the namespace. During the hosted cluster creation process on bare-metal infrastructure, a generated Cluster API provider role requires that the namespace already exists. -
Create the configuration file for your hosted cluster by entering the following command:
$ hcp create cluster agent \--name=<hosted_cluster_name> \--pull-secret=<path_to_pull_secret> \--agent-namespace=<hosted_control_plane_namespace> \--base-domain=<base_domain> \--api-server-address=api.<hosted_cluster_name>.<base_domain> \--etcd-storage-class=<etcd_storage_class> \--ssh-key=<path_to_ssh_key> \--namespace=<hosted_cluster_namespace> \--control-plane-availability-policy=HighlyAvailable \--release-image=quay.io/openshift-release-dev/ocp-release:<ocp_release_image>-multi \--node-pool-replicas=<node_pool_replica_count> \--disable-cluster-capabilities=<capability> \--enable-cluster-capabilities=<capability> \--render \--render-sensitive > hosted-cluster-config.yamlwhere:
--namespecifies the name of your hosted cluster, such asexample.--pull-secretspecifies the path to your pull secret, such as/user/name/pullsecret.--agent-namespacespecifies your hosted control plane namespace, such asclusters-example. Ensure that agents are available in this namespace by using theoc get agent -n <hosted_control_plane_namespace>command.--base-domainspecifies your base domain, such askrnl.es.--api-server-addressspecifies the IP address that gets used for the Kubernetes API communication in the hosted cluster. If you do not set the--api-server-addressflag, you must log in to connect to the management cluster.--etcd-storage-classspecifies the etcd storage class name, such aslvm-storageclass.--ssh-keyspecifies the path to your SSH public key. The default file path is~/.ssh/id_rsa.pub.--namespacespecifies your hosted cluster namespace.--control-plane-availability-policyspecifies the availability policy for the hosted control plane components. Supported options areSingleReplicaandHighlyAvailable. The default value isHighlyAvailable.--release-imagespecifies the supported OpenShift Container Platform version that you want to use, such as4.22.0-multi. If you are using a disconnected environment, replace<ocp_release_image>with the digest image. To extract the OpenShift Container Platform release image digest, see "Extracting the release image digest".--node-pool-replicasspecifies the node pool replica count, such as3. You must specify the replica count as0or greater to create the same number of replicas. Otherwise, you do not create node pools.--disable-cluster-capabilitiesspecifies that you want to disable optional capabilities in the hosted cluster. This flag is optional. For more information, see "Capabilities for hosted clusters".--enable-cluster-capabilitiesspecifies that you want to enable optional capabilities in the hosted cluster. This flag is optional. For more information, see "Capabilities for hosted clusters".--renderrenders the output as YAML to stdout instead of applying the resources to the cluster. By default, secrets are not included in the rendered output.--render-sensitiveincludes secrets in the rendered output when used with the--renderflag.
-
Configure the service publishing strategy. By default, hosted clusters use the
NodePortservice publishing strategy because node ports are always available without additional infrastructure. However, you can configure the service publishing strategy to use a load balancer.-
If you are using the default
NodePortstrategy, configure the DNS to point to the hosted cluster compute nodes, not the management cluster nodes. For more information, see "DNS configurations on bare metal". -
For production environments, use the
LoadBalancerstrategy because this strategy provides certificate handling and automatic DNS resolution. The following example demonstrates changing the service publishingLoadBalancerstrategy in your hosted cluster configuration file:apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:# ...spec:services:- service: APIServerservicePublishingStrategy:type: LoadBalancer- service: IgnitionservicePublishingStrategy:type: Route- service: KonnectivityservicePublishingStrategy:type: Route- service: OAuthServerservicePublishingStrategy:type: Route- service: OIDCservicePublishingStrategy:type: RoutesshKey:name: <ssh_key># ...Specify
LoadBalanceras the API Server type. For all other services, specifyRouteas the type.
-
-
If you use external load balancers, configure the ingress endpoint as shown in the following example. If you do not configure the endpoint, the default behavior is to randomize the node port that the service exposes the ingress on. To configure how the ingress controller publishes the default ingress route, set the
endpointPublishingStrategyparameter and its underlying functions by editing theHostedClusterresource:apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:#...spec:operatorConfiguration:ingressOperator:endpointPublishingStrategy:hostNetwork:httpPort: 80httpsPort: 443protocol: TCPstatsPort: 1936type: HostNetwork#...The
spec.operatorConfiguration.ingressOperator.endPointPublishingStrategy.typeparameter specifies the endpoint for the load balancer. For bare-metal installations, use theHostNetworktype. -
Apply the changes to the hosted cluster configuration file by entering the following command:
$ oc apply -f hosted_cluster_config.yaml -
Check for the creation of the hosted cluster, node pools, and pods by entering the following commands:
$ oc get hostedcluster \<hosted_cluster_namespace> -n \<hosted_cluster_namespace> -o \jsonpath='{.status.conditions[?(@.status=="False")]}' | jq .$ oc get nodepool \<hosted_cluster_namespace> -n \<hosted_cluster_namespace> -o \jsonpath='{.status.conditions[?(@.status=="False")]}' | jq .$ oc get pods -n <hosted_cluster_namespace> -
Confirm that the hosted cluster is ready. The status of
Available: Trueindicates the readiness of the cluster and the node pool status showsAllMachinesReady: True. These statuses indicate the healthiness of all cluster Operators. -
Install MetalLB in the hosted cluster:
-
Extract the
kubeconfigfile from the hosted cluster and set the environment variable for hosted cluster access by entering the following commands:$ oc get secret \<hosted_cluster_namespace>-admin-kubeconfig \-n <hosted_cluster_namespace> \-o jsonpath='{.data.kubeconfig}' \| base64 -d > \kubeconfig-<hosted_cluster_namespace>.yaml$ export KUBECONFIG="/path/to/kubeconfig-<hosted_cluster_namespace>.yaml" -
Install the MetalLB Operator by creating the
install-metallb-operator.yamlfile:apiVersion: v1kind: Namespacemetadata:name: metallb-system---apiVersion: operators.coreos.com/v1kind: OperatorGroupmetadata:name: metallb-operatornamespace: metallb-system---apiVersion: operators.coreos.com/v1alpha1kind: Subscriptionmetadata:name: metallb-operatornamespace: metallb-systemspec:channel: "stable"name: metallb-operatorsource: redhat-operatorssourceNamespace: openshift-marketplaceinstallPlanApproval: Automatic# ... -
Apply the file by entering the following command:
$ oc apply -f install-metallb-operator.yaml -
Configure the MetalLB IP address pool by creating the
deploy-metallb-ipaddresspool.yamlfile:apiVersion: metallb.io/v1beta1kind: IPAddressPoolmetadata:name: metallbnamespace: metallb-systemspec:autoAssign: trueaddresses:- 10.11.176.71-10.11.176.75---apiVersion: metallb.io/v1beta1kind: L2Advertisementmetadata:name: l2advertisementnamespace: metallb-systemspec:ipAddressPools:- metallb# ... -
Apply the configuration by entering the following command:
$ oc apply -f deploy-metallb-ipaddresspool.yaml -
Verify the installation of MetalLB by checking the Operator status, the IP address pool, and the
L2Advertisementresource by entering the following commands:$ oc get pods -n metallb-system$ oc get ipaddresspool -n metallb-system$ oc get l2advertisement -n metallb-system
-
-
Configure the load balancer for ingress:
-
Create the
ingress-loadbalancer.yamlfile:apiVersion: v1kind: Servicemetadata:annotations:metallb.universe.tf/address-pool: metallbname: metallb-ingressnamespace: openshift-ingressspec:ports:- name: httpprotocol: TCPport: 80targetPort: 80- name: httpsprotocol: TCPport: 443targetPort: 443selector:ingresscontroller.operator.openshift.io/deployment-ingresscontroller: defaulttype: LoadBalancer# ... -
Apply the configuration by entering the following command:
$ oc apply -f ingress-loadbalancer.yaml -
Verify that the load balancer service works as expected by entering the following command:
$ oc get svc metallb-ingress -n openshift-ingressExample outputNAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGEmetallb-ingress LoadBalancer 172.31.127.129 10.11.176.71 80:30961/TCP,443:32090/TCP 16h
-
-
Configure the DNS to work with the load balancer:
-
Configure the DNS for the
appsdomain by pointing the*.apps.<hosted_cluster_namespace>.<base_domain>wildcard DNS record to the load balancer IP address. -
Verify the DNS resolution by entering the following command:
$ nslookup console-openshift-console.apps.<hosted_cluster_namespace>.<base_domain> <load_balancer_ip_address>Example outputServer: 10.11.176.1Address: 10.11.176.1#53Name: console-openshift-console.apps.my-hosted-cluster.sample-base-domain.comAddress: 10.11.176.71
-
Verification
-
Check the cluster Operators by entering the following command:
$ oc get clusteroperatorsEnsure that all Operators show
AVAILABLE: True,PROGRESSING: False, andDEGRADED: False. -
Check the nodes by entering the following command:
$ oc get nodesEnsure that each node has the
READYstatus. -
Test access to the console by entering the following URL in a web browser:
https://console-openshift-console.apps.<hosted_cluster_namespace>.<base_domain>
Additional resources
- Requirements for hosted control planes
- DNS configurations on bare metal
- Manually importing a hosted cluster
- Extracting the release image digest
Capabilities for hosted clusters
To reduce resource consumption and prevent unnecessary Operators and operands from being deployed, administrators can enable or disable optional OpenShift Container Platform components when they create a hosted cluster.
When capabilities are not specified on a HostedCluster resource, the cluster uses the OpenShift Container Platform version’s DefaultCapabilitySet settings, excluding the baremetal capability. As a result, most optional components are enabled by default.
Capabilities are immutable after cluster creation. You cannot change them after you create the HostedCluster resource.
Capabilities that you can enable or disable for a hosted cluster
Familiarize yourself with the supported capabilities that you can enable or disable for a HostedCluster resource.
The capabilities are described in the following table:
| Capability | Description |
|---|---|
ImageRegistry | The OpenShift Image Registry Operator and its operands, including cloud storage infrastructure, such as S3 buckets and Identity and Access Management (IAM) users. |
openshift-samples | The OpenShift Samples Operator, which manages example ImageStreams and templates. |
Insights | The Insights Operator, which collects and uploads cluster telemetry data. |
baremetal | The Bare Metal Infrastructure Operator. This capability is excluded from the default set. If needed, you must explicitly enable it. |
Console | The OpenShift Web Console Operator and its operands. |
NodeTuning | The Node Tuning Operator, which manages node-level performance tuning by using TuneD and performance profiles. |
Ingress | The OpenShift Ingress Operator, which manages the default router of the cluster. |
The following rules apply when you combine capability settings:
- No overlap
- A capability cannot be in both the
enabledanddisabledlists simultaneously. - Console requires Ingress
- You can disable the
Ingresscapability only if theConsolecapability is also disabled because the console depends on Ingress. - Version requirement
- You must use OpenShift Container Platform 4.20 or later to disable any of the following capabilities:
openshift-samples,Insights,Console,NodeTuning, andIngress. You can disableImageRegistryandbaremetalon versions earlier than 4.20. - Bare metal default exclusion
- The
baremetalcapability is excluded from the default set. You can add it to the cluster by explicitly enabling it.
Setting capabilities for a hosted cluster
To reduce unnecessary resource consumption, you can control which optional capabilities are enabled for a hosted cluster when you create the cluster.
Capabilities are immutable after cluster creation. You cannot change them after you create the HostedCluster resource.
You can specify which capabilities are enabled by either using the hcp command-line interface (CLI) or by setting the HostedCluster manifest.
Procedure
-
To specify which capabilities are enabled in a hosted cluster by using the CLI, you can add the
--disable-cluster-capabilitiesflag, the--enable-cluster-capabilitiesflag, or both. The following example shows how to disable theImageRegistry,Console, andIngresscapabilities and enable thebaremetalcapability while you create a hosted cluster on AWS by using thehcpcommand-line interface:$ hcp create cluster aws \--name my-hosted-cluster \--disable-cluster-capabilities=ImageRegistry,Console,Ingress \--enable-cluster-capabilities=baremetalYou can specify multiple capabilities as a comma-separated list. The supported values are as follows:
ImageRegistryopenshift-samplesInsightsbaremetalConsoleNodeTuningIngress
-
To specify which capabilities are enabled in a hosted cluster by using the
HostedClustermanifest at cluster creation time, see the following examples:- To directly disable capabilities in a hosted cluster, add the
spec.capabilities.disabledsection in theHostedClusterresource:apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:name: my-hosted-clusternamespace: my-cluster-namespacespec:capabilities:disabled:- ImageRegistry- Console- Ingress# ... - To explicitly enable a capability that is not part of the default set of capabilities, such as the
baremetalcapability, see the following example:apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:name: my-hosted-clusternamespace: my-cluster-namespacespec:capabilities:enabled:- baremetal# ... - You can use both
enabledanddisabledif no capabilities are in both lists. See the following example:apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:name: my-hosted-clusternamespace: my-cluster-namespacespec:capabilities:enabled:- baremetaldisabled:- ImageRegistry- openshift-samples# ...
- To directly disable capabilities in a hosted cluster, add the
About creating heterogeneous node pools on agent hosted clusters
A node pool is a group of nodes within a cluster that share the same configuration. Heterogeneous node pools have different configurations so that you can create pools and optimize them for various workloads.
You can create heterogeneous node pools on the agent platform. The platform enables clusters to run diverse machine types, such as x86_64 or ppc64le, within a single hosted cluster.
Creating a heterogeneous node pool requires completion of the following general steps:
- Create an
AgentServiceConfigcustom resource (CR) that informs the Operator how much storage it needs for components such as the database and filesystem. The CR also defines which OpenShift Container Platform versions to support. - Create an agent cluster.
- Create the heterogeneous node pool.
- Configure DNS for hosted control planes
- Create an
InfraEnvcustom resource (CR) for each architecture. - Add agents to the heterogeneous cluster.
Creating the AgentServiceConfig custom resource
To create heterogeneous node pools on an agent hosted cluster, you need to create the AgentServiceConfig CR with two heterogeneous architecture operating system (OS) images.
Procedure
-
Run the following command:
$ envsubst <<"EOF" | oc apply -f -apiVersion: agent-install.openshift.io/v1beta1kind: AgentServiceConfigmetadata:name: agentspec:databaseStorage:accessModes:- ReadWriteOnceresources:requests:storage: <db_volume_name>filesystemStorage:accessModes:- ReadWriteOnceresources:requests:storage: <fs_volume_name>osImages:- openshiftVersion: <ocp_version>version: <ocp_release_version_x86>url: <iso_url_x86>rootFSUrl: <root_fs_url_x86>cpuArchitecture: <arch_x86>- openshiftVersion: <ocp_version>version: <ocp_release_version_ppc64le>url: <iso_url_ppc64le>rootFSUrl: <root_fs_url_ppc64le>cpuArchitecture: <arch_ppc64le>EOFwhere:
<db_volume_name>- Specifies the multicluster engine for Kubernetes Operator
agentserviceconfigconfig, database volume name. <fs_volume_name>- Specifies the multicluster engine Operator
agentserviceconfigconfig, filesystem volume name. <ocp_version>- Specifies the current version of OpenShift Container Platform.
<ocp_release_version_x86>- Specifies the current OpenShift Container Platform release version for x86.
<iso_url_x86>- Specifies the ISO URL for x86.
<root_fs_url_x86>- Specifies the root filesystem URL for x86.
<arch_x86>- Specifies the CPU architecture for x86.
<ocp_release_version_ppc64le>- Specifies the OpenShift Container Platform release version for
ppc64le. <iso_url_ppc64le>- Specifies the ISO URL for
ppc64le. <root_fs_url_ppc64le>- Specifies the root filesystem URL for
ppc64le. <arch_ppc64le>- Specifies the CPU architecture for
ppc64le.
Create an agent cluster
An agent-based approach manages and provisions an agent cluster. An agent cluster can use heterogeneous node pools, allowing the use of different types of compute nodes within the same cluster.
Prerequisites
- You used a multi-architecture release image to enable support for heterogeneous node pools when creating a hosted cluster. Find the latest multi-architecture images on the "Multi-arch release images" page.
Procedure
-
Create an environment variable for the cluster namespace by running the following command:
$ export CLUSTERS_NAMESPACE=<hosted_cluster_namespace> -
Create an environment variable for the machine classless inter-domain routing (CIDR) notation by running the following command:
$ export MACHINE_CIDR=192.168.122.0/24 -
Create the hosted control namespace by running the following command:
$ oc create ns <hosted_control_plane_namespace> -
Create the cluster by running the following command:
$ hcp create cluster agent \--name=<hosted_cluster_name> \--pull-secret=<pull_secret_file> \--agent-namespace=<hosted_control_plane_namespace> \--base-domain=<base_domain> \--api-server-address=api.<hosted_cluster_name>.<basedomain> \--release-image=quay.io/openshift-release-dev/ocp-release:<ocp_release>where:
<hosted_cluster_name>- Specifies the hosted cluster name.
<pull_secret_file>- Specifies the pull secret file path.
<hosted_control_plane_namespace>- Specifies the namespace for the hosted control plane.
<base_domain>- Specifies the base domain for the hosted cluster.
<ocp_release>- Specifies the current OpenShift Container Platform release version.
Additional resources
Creating heterogeneous node pools
You create heterogeneous node pools by using the NodePool custom resource (CR), so that you can optimize costs and performance by associating different workloads to specific hardware.
Procedure
-
To define a
NodePoolCR, create a YAML file similar to the following example:envsubst <<"EOF" | oc apply -f -apiVersion:apiVersion: hypershift.openshift.io/v1beta1kind: NodePoolmetadata:name: <hosted_cluster_name>namespace: <clusters_namespace>spec:arch: <arch_ppc64le>clusterName: <hosted_cluster_name>management:autoRepair: falseupgradeType: InPlacenodeDrainTimeout: 0snodeVolumeDetachTimeout: 0splatform:agent:agentLabelSelector:matchLabels:inventory.agent-install.openshift.io/cpu-architecture: <arch_ppc64le>type: Agentrelease:image: quay.io/openshift-release-dev/ocp-release:<ocp_release>replicas: 0EOF<arch_ppc64le>: The selector block selects the agents that match the specified label. To create a node pool of architectureppc64lewith zero replicas, specifyppc64le. This ensures that the selector block selects only agents fromppc64learchitecture during a scaling operation.
DNS configuration for hosted control planes
A Domain Name Service (DNS) configuration for hosted control planes means that external clients can reach ingress controllers, so that the clients can route traffic to internal components. Configuring this setting ensures that traffic gets routed to either a ppc64le or an x86_64 compute node.
You can point an *.apps.<cluster_name> record to either of the compute nodes that hosts the ingress application. Or, if you can set up a load balancer on top of the compute nodes, point the record to this load balancer. When you are creating a heterogeneous node pool, make sure the compute nodes can reach each other or keep them in the same network.
Creating infrastructure environment resources
For heterogeneous node pools, you must create an infraEnv custom resource (CR) for each architecture. This configuration ensures that the correct architecture-specific operating system and boot artifacts get used during the node provisioning process.
For example, for node pools with x86_64 and ppc64le architectures, create an InfraEnv CR for x86_64 and ppc64le.
Before starting the procedure, ensure that you add the operating system images for both x86_64 and ppc64le architectures to the AgentServiceConfig resource. After this, you can use the InfraEnv resources to get the minimal ISO image.
Procedure
-
Create the
InfraEnvresource withx86_64architecture for heterogeneous node pools by running the following command:$ envsubst <<"EOF" | oc apply -f -apiVersion: agent-install.openshift.io/v1beta1kind: InfraEnvmetadata:name: <hosted_cluster_name>-<arch_x86>namespace: <hosted_control_plane_namespace>spec:cpuArchitecture: <arch_x86>pullSecretRef:name: pull-secretsshAuthorizedKey: <ssh_pub_key>EOFwhere:
<hosted_cluster_name>- Specifies the hosted cluster name.
<arch_x86>- Specifies the
x86_64architecture. <hosted_control_plane_namespace>- Specifies the hosted control plane namespace.
<ssh_pub_key>- Specifies the SSH public key.
-
Create the
InfraEnvresource withppc64learchitecture for heterogeneous node pools by running the following command:envsubst <<"EOF" | oc apply -f -apiVersion: agent-install.openshift.io/v1beta1kind: InfraEnvmetadata:name: <hosted_cluster_name>-<arch_ppc64le>namespace: <hosted_control_plane_namespace>spec:cpuArchitecture: <arch_ppc64le>pullSecretRef:name: pull-secretsshAuthorizedKey: <ssh_pub_key>EOFwhere:
<hosted_cluster_name>- Specifies the hosted cluster name.
<arch_ppc64le>- Specifies the
ppc64learchitecture. <hosted_control_plane_namespace>- Specifies the hosted control plane namespace.
<ssh_pub_key>- Specifies the SSH public key.
-
Verify the successful creation of the
InfraEnvresources by running the following commands:- Verify the successful creation of the
x86_64InfraEnvresource:$ oc describe InfraEnv <hosted_cluster_name>-<arch_x86> - Verify the successful creation of the
ppc64leInfraEnvresource:$ oc describe InfraEnv <hosted_cluster_name>-<arch_ppc64le>
- Verify the successful creation of the
-
Generate a live ISO that allows either a virtual machine or a bare-metal machine to join as agents by running the following commands:
- Generate a live ISO for
x86_64:$ oc -n <hosted_control_plane_namespace> get InfraEnv <hosted_cluster_name>-<arch_x86> -ojsonpath="{.status.isoDownloadURL}" - Generate a live ISO for
ppc64le:$ oc -n <hosted_control_plane_namespace> get InfraEnv <hosted_cluster_name>-<arch_ppc64le> -ojsonpath="{.status.isoDownloadURL}"
- Generate a live ISO for
Adding agents to the heterogeneous cluster
You add agents by manually configuring the machine to boot with a live ISO. You can download the live ISO and use it to boot a bare-metal node or a virtual machine.
On boot, the node communicates with the assisted-service and registers as an agent in the same namespace as the InfraEnv resource. After the creation of each agent, you can optionally set its installation_disk_id and hostname parameters in the specifications. You can then approve the agent to indicate the agent as ready for use.
Procedure
-
Obtain a list of agents by running the following command:
$ oc -n <hosted_control_plane_namespace> get agentsExample outputNAME CLUSTER APPROVED ROLE STAGE86f7ac75-4fc4-4b36-8130-40fa12602218 auto-assigne57a637f-745b-496e-971d-1abbf03341ba auto-assign -
Patch an agent by running the following command:
$ oc -n <hosted_control_plane_namespace> patch agent 86f7ac75-4fc4-4b36-8130-40fa12602218 -p '{"spec":{"installation_disk_id":"/dev/sda","approved":true,"hostname":"worker-0.example.krnl.es"}}' --type merge -
Patch the second agent by running the following command:
$ oc -n <hosted_control_plane_namespace> patch agent 23d0c614-2caa-43f5-b7d3-0b3564688baa -p '{"spec":{"installation_disk_id":"/dev/sda","approved":true,"hostname":"worker-1.example.krnl.es"}}' --type merge -
Check the agent approval status by running the following command:
$ oc -n <hosted_control_plane_namespace> get agentsExample outputNAME CLUSTER APPROVED ROLE STAGE86f7ac75-4fc4-4b36-8130-40fa12602218 true auto-assigne57a637f-745b-496e-971d-1abbf03341ba true auto-assign
Scaling the node pool
After you approve your agents, you can scale the node pools. The agentLabelSelector value that you configured in the node pool ensures that only matching agents get added to the cluster. This also helps scale down the node pool.
To remove specific architecture nodes from the cluster, scale down the corresponding node pool.
Procedure
-
Scale the node pool by running the following command:
$ oc -n <clusters_namespace> scale nodepool <nodepool_name> --replicas 2noteThe Cluster API agent provider picks two agents randomly to assign to the hosted cluster. These agents pass through different states and then join the hosted cluster as OpenShift Container Platform nodes. The various agent states are
binding,discovering,insufficient,installing,installing-in-progress, andadded-to-existing-cluster.
Verification
-
List the agents by running the following command:
$ oc -n <hosted_control_plane_namespace> get agentExample outputNAME CLUSTER APPROVED ROLE STAGE4dac1ab2-7dd5-4894-a220-6a3473b67ee6 hypercluster1 true auto-assignd9198891-39f4-4930-a679-65fb142b108b true auto-assignda503cf1-a347-44f2-875c-4960ddb04091 hypercluster1 true auto-assign -
Check the status of a specific scaled agent by running the following command:
$ oc -n <hosted_control_plane_namespace> get agent -o jsonpath='{range .items[*]}BMH: {@.metadata.labels.agent-install\.openshift\.io/bmh} Agent: {@.metadata.name} State: {@.status.debugInfo.state}{"\n"}{end}'Example outputBMH: ocp-worker-2 Agent: 4dac1ab2-7dd5-4894-a220-6a3473b67ee6 State: bindingBMH: ocp-worker-0 Agent: d9198891-39f4-4930-a679-65fb142b108b State: known-unboundBMH: ocp-worker-1 Agent: da503cf1-a347-44f2-875c-4960ddb04091 State: insufficient -
After the agents reach the
added-to-existing-clusterstate, verify that the OpenShift Container Platform nodes are ready by running the following command:$ oc --kubeconfig <hosted_cluster_name>.kubeconfig get nodesExample outputNAME STATUS ROLES AGE VERSIONocp-worker-1 Ready worker 5m41s v1.24.0+3882f8focp-worker-2 Ready worker 6m3s v1.24.0+3882f8f -
Adding workloads to the nodes can reconcile some cluster operators. The following command displays the creation of two machines that happened after scaling up the node pool:
$ oc -n <hosted_control_plane_namespace> get machinesExample outputNAME CLUSTER NODENAME PROVIDERID PHASE AGE VERSIONhypercluster1-c96b6f675-m5vch hypercluster1-b2qhl ocp-worker-1 agent://da503cf1-a347-44f2-875c-4960ddb04091 Running 15m 4.11.5hypercluster1-c96b6f675-tl42p hypercluster1-b2qhl ocp-worker-2 agent://4dac1ab2-7dd5-4894-a220-6a3473b67ee6 Running 15m 4.11.5