Managing hosted control planes on AWS
After you deploy a hosted cluster on Amazon Web Services (AWS), you can manage the cluster.
Prerequisites to manage AWS infrastructure and IAM permissions
To configure hosted control planes for OpenShift Container Platform on Amazon Web Services (AWS), you must meet the infrastructure requirements.
- You must configure hosted control planes before you can create hosted clusters.
- You must create an AWS Identity and Access Management (IAM) role and AWS Security Token Service (STS) credentials.
Infrastructure requirements for AWS
When you use hosted control planes on Amazon Web Services (AWS), the infrastructure requirements vary based on your setup.
The infrastructure requirements fit in the following categories:
- Prerequired and unmanaged infrastructure for the HyperShift Operator in an arbitrary AWS account
- Prerequired and unmanaged infrastructure in a hosted cluster AWS account
- Hosted control planes-managed infrastructure in a management AWS account
- Hosted control planes-managed infrastructure in a hosted cluster AWS account
- Kubernetes-managed infrastructure in a hosted cluster AWS account
Prerequired means that hosted control planes requires AWS infrastructure to properly work. Unmanaged means that no Operator or controller creates the infrastructure for you.
Unmanaged infrastructure for the HyperShift Operator in an AWS account
An arbitrary Amazon Web Services (AWS) account depends on the provider of the hosted control planes service.
In self-managed hosted control planes, the cluster service provider controls the AWS account. The cluster service provider is the administrator who hosts cluster control planes and is responsible for uptime. In managed hosted control planes, the AWS account belongs to Red Hat.
In a prerequired and unmanaged infrastructure for the HyperShift Operator, the following infrastructure requirements apply for a management cluster AWS account:
- One S3 Bucket
- OpenID Connect (OIDC)
- Route 53 hosted zones
- A domain to host private and public entries for hosted clusters
Unmanaged infrastructure requirements for a management AWS account
When your infrastructure is prerequired and unmanaged in a hosted cluster Amazon Web Services (AWS) account, ensure that you are familiar with the infrastructure requirements for all access modes.
-
One VPC
-
One DHCP Option
-
Two subnets
- A private subnet that is an internal data plane subnet
- A public subnet that enables access to the internet from the data plane
-
One internet gateway
-
One elastic IP
-
One NAT gateway
-
One security group (worker nodes)
-
Two route tables (one private and one public)
-
Two Route 53 hosted zones
-
Enough quota for the following items:
- One Ingress service load balancer for public hosted clusters
- One private link endpoint for private hosted clusters
noteFor private link networking to work, the endpoint zone in the hosted cluster AWS account must match the zone of the instance that is resolved by the service endpoint in the management cluster AWS account. In AWS, the zone names are aliases, such as us-east-2b, which do not necessarily map to the same zone in different accounts. As a result, for private link to work, the management cluster must have subnets or workers in all zones of its region.
Infrastructure requirements for a management AWS account
When your infrastructure is managed by hosted control planes in a management AWS account, the infrastructure requirements differ depending on whether your clusters are public, private, or a combination.
For accounts with public clusters, the infrastructure requirements are as follows:
- Network load balancer: a load balancer Kube API server
- Kubernetes creates a security group
- Volumes
- For etcd (one or three depending on high availability)
- For OVN-Kube
For accounts with private clusters, the infrastructure requirements are as follows:
- Network load balancer: a load balancer private router
- Endpoint service (private link)
For accounts with public and private clusters, the infrastructure requirements are as follows:
- Network load balancer: a load balancer public router
- Network load balancer: a load balancer private router
- Endpoint service (private link)
- Volumes
- For etcd (one or three depending on high availability)
- For OVN-Kube
Infrastructure requirements for an AWS account in a hosted cluster
When your infrastructure is managed by hosted control planes in a hosted cluster Amazon Web Services (AWS) account, the infrastructure requirements differ depending on whether your clusters are public, private, or a combination.
For accounts with public clusters, the infrastructure requirements are as follows:
- Node pools must have EC2 instances that have
RoleandRolePolicydefined.
For accounts with private clusters, the infrastructure requirements are as follows:
- One private link endpoint for each availability zone
- EC2 instances for node pools
For accounts with public and private clusters, the infrastructure requirements are as follows:
- One private link endpoint for each availability zone
- EC2 instances for node pools
Kubernetes-managed infrastructure in a hosted cluster AWS account
When Kubernetes manages your infrastructure in a hosted cluster Amazon Web Services (AWS) account, ensure that you meet the infrastructure requirements.
The infrastructure requirements are as follows:
- A network load balancer for default Ingress
- An S3 bucket for registry
Identity and Access Management (IAM) permissions for hosted control planes on AWS
In the context of hosted control planes, the consumer is responsible to create the Amazon Resource Name (ARN) roles. The consumer is an automated process to generate the permissions files. It might be the command-line interface (CLI) or OpenShift Cluster Manager.
Hosted control planes can enable granularity to honor the principle of least-privilege components, which means that every component uses its own role to operate or create Amazon Web Services (AWS) objects, and the roles are limited to what is required for the product to function normally.
The hosted cluster receives the ARN roles as input and the consumer creates an AWS permission configuration for each component. As a result, the component can authenticate through STS and preconfigured OIDC IDP.
The following roles are consumed by some of the components from hosted control planes that run on the control plane and operate on the data plane:
controlPlaneOperatorARNimageRegistryARNingressARNkubeCloudControllerARNnodePoolManagementARNstorageARNnetworkARN
The following example shows a reference to the IAM roles from the hosted cluster:
...
endpointAccess: Public
region: us-east-2
resourceTags:
- key: kubernetes.io/cluster/example-cluster-bz4j5
value: owned
rolesRef:
controlPlaneOperatorARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-control-plane-operator
imageRegistryARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-openshift-image-registry
ingressARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-openshift-ingress
kubeCloudControllerARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-cloud-controller
networkARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-cloud-network-config-controller
nodePoolManagementARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-node-pool
storageARN: arn:aws:iam::820196288204:role/example-cluster-bz4j5-aws-ebs-csi-driver-controller
type: AWS
...
The roles that hosted control planes uses are shown in the following examples:
ingressARN{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Action": ["elasticloadbalancing:DescribeLoadBalancers","tag:GetResources","route53:ListHostedZones"],"Resource": "\*"},{"Effect": "Allow","Action": ["route53:ChangeResourceRecordSets"],"Resource": ["arn:aws:route53:::PUBLIC_ZONE_ID","arn:aws:route53:::PRIVATE_ZONE_ID"]}]}imageRegistryARN{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Action": ["s3:CreateBucket","s3:DeleteBucket","s3:PutBucketTagging","s3:GetBucketTagging","s3:PutBucketPublicAccessBlock","s3:GetBucketPublicAccessBlock","s3:PutEncryptionConfiguration","s3:GetEncryptionConfiguration","s3:PutLifecycleConfiguration","s3:GetLifecycleConfiguration","s3:GetBucketLocation","s3:ListBucket","s3:GetObject","s3:PutObject","s3:DeleteObject","s3:ListBucketMultipartUploads","s3:AbortMultipartUpload","s3:ListMultipartUploadParts"],"Resource": "\*"}]}storageARN{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Action": ["ec2:AttachVolume","ec2:CreateSnapshot","ec2:CreateTags","ec2:CreateVolume","ec2:DeleteSnapshot","ec2:DeleteTags","ec2:DeleteVolume","ec2:DescribeInstances","ec2:DescribeSnapshots","ec2:DescribeTags","ec2:DescribeVolumes","ec2:DescribeVolumesModifications","ec2:DetachVolume","ec2:ModifyVolume"],"Resource": "\*"}]}networkARN{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Action": ["ec2:DescribeInstances","ec2:DescribeInstanceStatus","ec2:DescribeInstanceTypes","ec2:UnassignPrivateIpAddresses","ec2:AssignPrivateIpAddresses","ec2:UnassignIpv6Addresses","ec2:AssignIpv6Addresses","ec2:DescribeSubnets","ec2:DescribeNetworkInterfaces"],"Resource": "\*"}]}kubeCloudControllerARN{"Version": "2012-10-17","Statement": [{"Action": ["ec2:DescribeInstances","ec2:DescribeImages","ec2:DescribeRegions","ec2:DescribeRouteTables","ec2:DescribeSecurityGroups","ec2:DescribeSubnets","ec2:DescribeVolumes","ec2:CreateSecurityGroup","ec2:CreateTags","ec2:CreateVolume","ec2:ModifyInstanceAttribute","ec2:ModifyVolume","ec2:AttachVolume","ec2:AuthorizeSecurityGroupIngress","ec2:CreateRoute","ec2:DeleteRoute","ec2:DeleteSecurityGroup","ec2:DeleteVolume","ec2:DetachVolume","ec2:RevokeSecurityGroupIngress","ec2:DescribeVpcs","elasticloadbalancing:AddTags","elasticloadbalancing:AttachLoadBalancerToSubnets","elasticloadbalancing:ApplySecurityGroupsToLoadBalancer","elasticloadbalancing:CreateLoadBalancer","elasticloadbalancing:CreateLoadBalancerPolicy","elasticloadbalancing:CreateLoadBalancerListeners","elasticloadbalancing:ConfigureHealthCheck","elasticloadbalancing:DeleteLoadBalancer","elasticloadbalancing:DeleteLoadBalancerListeners","elasticloadbalancing:DescribeLoadBalancers","elasticloadbalancing:DescribeLoadBalancerAttributes","elasticloadbalancing:DetachLoadBalancerFromSubnets","elasticloadbalancing:DeregisterInstancesFromLoadBalancer","elasticloadbalancing:ModifyLoadBalancerAttributes","elasticloadbalancing:RegisterInstancesWithLoadBalancer","elasticloadbalancing:SetLoadBalancerPoliciesForBackendServer","elasticloadbalancing:AddTags","elasticloadbalancing:CreateListener","elasticloadbalancing:CreateTargetGroup","elasticloadbalancing:DeleteListener","elasticloadbalancing:DeleteTargetGroup","elasticloadbalancing:DescribeListeners","elasticloadbalancing:DescribeLoadBalancerPolicies","elasticloadbalancing:DescribeTargetGroups","elasticloadbalancing:DescribeTargetHealth","elasticloadbalancing:ModifyListener","elasticloadbalancing:ModifyTargetGroup","elasticloadbalancing:RegisterTargets","elasticloadbalancing:SetLoadBalancerPoliciesOfListener","iam:CreateServiceLinkedRole","kms:DescribeKey"],"Resource": ["\*"],"Effect": "Allow"}]}nodePoolManagementARN{"Version": "2012-10-17","Statement": [{"Action": ["ec2:AllocateAddress","ec2:AssociateRouteTable","ec2:AttachInternetGateway","ec2:AuthorizeSecurityGroupIngress","ec2:CreateInternetGateway","ec2:CreateNatGateway","ec2:CreateRoute","ec2:CreateRouteTable","ec2:CreateSecurityGroup","ec2:CreateSubnet","ec2:CreateTags","ec2:DeleteInternetGateway","ec2:DeleteNatGateway","ec2:DeleteRouteTable","ec2:DeleteSecurityGroup","ec2:DeleteSubnet","ec2:DeleteTags","ec2:DescribeAccountAttributes","ec2:DescribeAddresses","ec2:DescribeAvailabilityZones","ec2:DescribeImages","ec2:DescribeInstances","ec2:DescribeInternetGateways","ec2:DescribeNatGateways","ec2:DescribeNetworkInterfaces","ec2:DescribeNetworkInterfaceAttribute","ec2:DescribeRouteTables","ec2:DescribeSecurityGroups","ec2:DescribeSubnets","ec2:DescribeVpcs","ec2:DescribeVpcAttribute","ec2:DescribeVolumes","ec2:DetachInternetGateway","ec2:DisassociateRouteTable","ec2:DisassociateAddress","ec2:ModifyInstanceAttribute","ec2:ModifyNetworkInterfaceAttribute","ec2:ModifySubnetAttribute","ec2:ReleaseAddress","ec2:RevokeSecurityGroupIngress","ec2:RunInstances","ec2:TerminateInstances","tag:GetResources","ec2:CreateLaunchTemplate","ec2:CreateLaunchTemplateVersion","ec2:DescribeLaunchTemplates","ec2:DescribeLaunchTemplateVersions","ec2:DeleteLaunchTemplate","ec2:DeleteLaunchTemplateVersions"],"Resource": ["\*"],"Effect": "Allow"},{"Condition": {"StringLike": {"iam:AWSServiceName": "elasticloadbalancing.amazonaws.com"}},"Action": ["iam:CreateServiceLinkedRole"],"Resource": ["arn:*:iam::*:role/aws-service-role/elasticloadbalancing.amazonaws.com/AWSServiceRoleForElasticLoadBalancing"],"Effect": "Allow"},{"Action": ["iam:PassRole"],"Resource": ["arn:*:iam::*:role/*-worker-role"],"Effect": "Allow"}]}controlPlaneOperatorARN{"Version": "2012-10-17","Statement": [{"Effect": "Allow","Action": ["ec2:CreateVpcEndpoint","ec2:DescribeVpcEndpoints","ec2:ModifyVpcEndpoint","ec2:DeleteVpcEndpoints","ec2:CreateTags","route53:ListHostedZones"],"Resource": "\*"},{"Effect": "Allow","Action": ["route53:ChangeResourceRecordSets","route53:ListResourceRecordSets"],"Resource": "arn:aws:route53:::%s"}]}
Separate creation of the hosted cluster and its resources
By default, the hcp create cluster aws command creates the cloud infrastructure with the hosted cluster and applies it. However, you can create the cloud infrastructure separately so that you can use the command to create only the cluster, or render it to modify it before you apply it.
The process to create the hosted cluster and its resources separately involves creating the cloud infrastructure, creating the AWS Identity and Access (IAM) resources, and then creating the cluster.
Creating the AWS infrastructure separately
To create the Amazon Web Services (AWS) infrastructure, you need to create a Virtual Private Cloud (VPC) and other resources for your cluster. You can use the AWS console or an infrastructure automation and provisioning tool.
For instructions to use the AWS console, see Create a VPC plus other VPC resources in the AWS Documentation.
The VPC must include private and public subnets and resources for external access, such as a network address translation (NAT) gateway and an internet gateway. In addition to the VPC, you need a private hosted zone for the ingress of your cluster. If you are creating clusters that use PrivateLink (Private or PublicAndPrivate access modes), you need an additional hosted zone for PrivateLink.
Procedure
-
Create the AWS infrastructure for your hosted cluster by using the following example configuration:
---apiVersion: v1kind: Namespacemetadata:creationTimestamp: nullname: clustersspec: {}status: {}---apiVersion: v1data:.dockerconfigjson: xxxxxxxxxxxkind: Secretmetadata:creationTimestamp: nulllabels:hypershift.openshift.io/safe-to-delete-with-cluster: "true"name: <pull_secret_name>namespace: clusters---apiVersion: v1data:key: xxxxxxxxxxxxxxxxxkind: Secretmetadata:creationTimestamp: nulllabels:hypershift.openshift.io/safe-to-delete-with-cluster: "true"name: <etcd_encryption_key_name>namespace: clusterstype: Opaque---apiVersion: v1data:id_rsa: xxxxxxxxxid_rsa.pub: xxxxxxxxxkind: Secretmetadata:creationTimestamp: nulllabels:hypershift.openshift.io/safe-to-delete-with-cluster: "true"name: <ssh_key_name>namespace: clusters---apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:creationTimestamp: nullname: <hosted_cluster_name>namespace: clustersspec:autoscaling: {}configuration: {}controllerAvailabilityPolicy: SingleReplicadns:baseDomain: <dns_domain>privateZoneID: xxxxxxxxpublicZoneID: xxxxxxxxetcd:managed:storage:persistentVolume:size: 8GistorageClassName: gp3-csitype: PersistentVolumemanagementType: Managedfips: falseinfraID: <infra_id>issuerURL: <issuer_url>networking:clusterNetwork:- cidr: 10.132.0.0/14machineNetwork:- cidr: 10.0.0.0/16networkType: OVNKubernetesserviceNetwork:- cidr: 172.31.0.0/16olmCatalogPlacement: managementplatform:aws:cloudProviderConfig:subnet:id: <subnet_xxx>vpc: <vpc_xxx>zone: us-west-1bendpointAccess: PublicmultiArch: falseregion: us-west-1rolesRef:controlPlaneOperatorARN: arn:aws:iam::820196288204:role/<infra_id>-control-plane-operatorimageRegistryARN: arn:aws:iam::820196288204:role/<infra_id>-openshift-image-registryingressARN: arn:aws:iam::820196288204:role/<infra_id>-openshift-ingresskubeCloudControllerARN: arn:aws:iam::820196288204:role/<infra_id>-cloud-controllernetworkARN: arn:aws:iam::820196288204:role/<infra_id>-cloud-network-config-controllernodePoolManagementARN: arn:aws:iam::820196288204:role/<infra_id>-node-poolstorageARN: arn:aws:iam::820196288204:role/<infra_id>-aws-ebs-csi-driver-controllertype: AWSpullSecret:name: <pull_secret_name>release:image: quay.io/openshift-release-dev/ocp-release:4.16-x86_64secretEncryption:aescbc:activeKey:name: <etcd_encryption_key_name>type: aescbcservices:- service: APIServerservicePublishingStrategy:type: LoadBalancer- service: OAuthServerservicePublishingStrategy:type: Route- service: KonnectivityservicePublishingStrategy:type: Route- service: IgnitionservicePublishingStrategy:type: Route- service: OVNSbDbservicePublishingStrategy:type: RoutesshKey:name: <ssh_key_name>status:controlPlaneEndpoint:host: ""port: 0---apiVersion: hypershift.openshift.io/v1beta1kind: NodePoolmetadata:creationTimestamp: nullname: <node_pool_name>namespace: clustersspec:arch: amd64clusterName: <hosted_cluster_name>management:autoRepair: trueupgradeType: ReplacenodeDrainTimeout: 0splatform:aws:instanceProfile: <instance_profile_name>instanceType: m6i.xlargerootVolume:size: 120type: gp3subnet:id: <subnet_xxx>type: AWSrelease:image: quay.io/openshift-release-dev/ocp-release:4.16-x86_64replicas: 2status:replicas: 0<pull_secret_name>specifies the name of your pull secret.<etcd_encryption_key_name>specifies the name of your etcd encryption key.<ssh_key_name>specifies the name of your SSH key.<hosted_cluster_name>specifies the name of your hosted cluster.<dns_domain>specifies your base DNS domain, such asexample.com.<infra_id>specifies the value that identifies the IAM resources that are associated with the hosted cluster.<issuer_url>specifies your issuer URL, which ends with yourinfra_idvalue. For example,https://example-hosted-us-west-1.s3.us-west-1.amazonaws.com/example-hosted-infra-id.<subnet_xxx>specifies your subnet ID. Both private and public subnets need to be tagged. For public subnets, usekubernetes.io/role/elb=1. For private subnets, usekubernetes.io/role/internal-elb=1.<vpc_xxx>specifies your VPC ID.<node_pool_name>specifies the name of yourNodePoolresource.<instance_profile_name>specifies the name of your AWS instance.
Creating a hosted cluster separately
In hosted control planes on AWS, you can create a hosted cluster separately from creating the infrastructure and Identity and Access Management (IAM) resources.
Prerequisites
- You created infrastructure resources separately. For more information, see "Creating the AWS infrastructure separately".
- You created the following IAM resources:
- An OpenID Connect (OIDC) identity provider in IAM, which is required to enable STS authentication. For more information, see Create an OpenID Connect (OIDC) identity provider in IAM.
- The seven roles that are listed in "Identity and Access Management (IAM) permissions." The roles are separate for every component that interacts with the provider, such as the Kubernetes controller manager, cluster API provider, and registry. For more information, see IAM role creation.
- The instance profile, which is the profile that is assigned to all worker instances of the cluster. For more information, see Use instance profiles.
Procedure
-
To create a hosted cluster separately, enter the following command:
$ hcp create cluster aws \--infra-id <infra_id> \--name <hosted_cluster_name> \--sts-creds <path_to_sts_credential_file> \--pull-secret <path_to_pull_secret> \--generate-ssh \--node-pool-replicas 3 \--role-arn <role_name> \--render-sensitive \--render > <file_name>.yaml--infra-idspecifies the same ID that you specified in thecreate infra awscommand. This value identifies the IAM resources that are associated with the hosted cluster.--namespecifies the name of your hosted cluster.--sts-credsspecifies the same name that you specified in thecreate infra awscommand.--pull-secretspecifies the name of the file that contains a valid OpenShift Container Platform pull secret.--generate-sshis an optional flag, but it is good to include in case you need to SSH to your workers. An SSH key is generated for you and is stored as a secret in the same namespace as the hosted cluster.--role-arnspecifies the Amazon Resource Name (ARN); for example,arn:aws:iam::820196288204:role/myrole. For more information about ARN roles, see "Identity and Access Management (IAM) permissions".--render-sensitivegenerates secrets that are stored in the YAML file.--renderis an optional flag. You can include this flag to redirect output to a file where you can edit the resources before you apply them to the cluster.
-
Apply the manifests by entering the following command:
$ oc apply -f <file_name>.yaml
Verification
- After you run the command, you can verify that the following resources are applied to your cluster:
- A namespace
- A secret with your pull secret
- A
HostedCluster - A
NodePool - Three AWS STS secrets for control plane components
- If you specified the
--generate-sshflag, one SSH key secret.
Transitioning a hosted cluster from single-architecture to multi-architecture
You can transition your single-architecture 64-bit AMD hosted cluster to a multi-architecture hosted cluster on Amazon Web Services (AWS) to reduce the cost of running workloads on your cluster.
For example, you can run existing workloads on 64-bit AMD while transitioning to 64-bit ARM and you can manage these workloads from a central Kubernetes cluster.
A single-architecture hosted cluster can manage node pools of only one particular CPU architecture. However, a multi-architecture hosted cluster can manage node pools with different CPU architectures. On AWS, a multi-architecture hosted cluster can manage both 64-bit AMD and 64-bit ARM node pools.
Prerequisites
- You installed an OpenShift Container Platform management cluster for AWS with the multicluster engine for Kubernetes Operator.
- You have an existing single-architecture hosted cluster that uses 64-bit AMD variant of the OpenShift Container Platform release payload.
- You have an existing node pool that uses the same 64-bit AMD variant of the OpenShift Container Platform release payload and is managed by an existing hosted cluster.
- You installed the following command-line tools:
ockubectlhcpskopeo
Procedure
-
Review an existing OpenShift Container Platform release image of the single-architecture hosted cluster by running the following command:
$ oc get hostedcluster/<hosted_cluster_name> \-o jsonpath='{.spec.release.image}'Replace
<hosted_cluster_name>with your hosted cluster name.Example outputquay.io/openshift-release-dev/ocp-release:4.20.0-x86_64 -
In your OpenShift Container Platform release image, if you use the digest instead of a tag, find the multi-architecture tag version of your release image:
-
Set the
OCP_VERSIONenvironment variable for the OpenShift Container Platform version by running the following command:$ OCP_VERSION=$(oc image info quay.io/openshift-release-dev/ocp-release@sha256:ac78ebf77f95ab8ff52847ecd22592b545415e1ff6c7ff7f66bf81f158ae4f5e \-o jsonpath='{.config.config.Labels["io.openshift.release"]}') -
Set the
MULTI_ARCH_TAGenvironment variable for the multi-architecture tag version of your release image by running the following command:$ MULTI_ARCH_TAG=$(skopeo inspect docker://quay.io/openshift-release-dev/ocp-release@sha256:ac78ebf77f95ab8ff52847ecd22592b545415e1ff6c7ff7f66bf81f158ae4f5e \| jq -r '.RepoTags' | sed 's/"//g' | sed 's/,//g' \| grep -w "$OCP_VERSION-multi$" | xargs) -
Set the
IMAGEenvironment variable for the multi-architecture release image name by running the following command:$ IMAGE=quay.io/openshift-release-dev/ocp-release:$MULTI_ARCH_TAG -
To see the list of multi-architecture image digests, run the following command:
$ oc image info $IMAGEExample outputOS DIGESTlinux/amd64 sha256:b4c7a91802c09a5a748fe19ddd99a8ffab52d8a31db3a081a956a87f22a22ff8linux/ppc64le sha256:66fda2ff6bd7704f1ba72be8bfe3e399c323de92262f594f8e482d110ec37388linux/s390x sha256:b1c1072dc639aaa2b50ec99b530012e3ceac19ddc28adcbcdc9643f2dfd14f34linux/arm64 sha256:7b046404572ac96202d82b6cb029b421dddd40e88c73bbf35f602ffc13017f21
-
-
Transition the hosted cluster from single-architecture to multi-architecture:
-
Set the multi-architecture OpenShift Container Platform release image for the hosted cluster by ensuring that you use the same OpenShift Container Platform version as the hosted cluster. Run the following command:
$ oc patch -n clusters hostedclusters/<hosted_cluster_name> -p \'{"spec":{"release":{"image":"quay.io/openshift-release-dev/ocp-release:<4.x.y>-multi"}}}' \--type=mergeReplace
<4.y.z>with the supported OpenShift Container Platform version that you use. -
Confirm that the multi-architecture image is set in your hosted cluster by running the following command:
$ oc get hostedcluster/<hosted_cluster_name> \-o jsonpath='{.spec.release.image}'
-
-
Check that the status of the
HostedControlPlaneresource isProgressingby running the following command:$ oc get hostedcontrolplane -n <hosted_control_plane_namespace> -oyamlExample output#...- lastTransitionTime: "2024-07-28T13:07:18Z"message: HostedCluster is deploying, upgrading, or reconfiguringobservedGeneration: 5reason: Progressingstatus: "True"type: Progressing#... -
Check that the status of the
HostedClusterresource isProgressingby running the following command:$ oc get hostedcluster <hosted_cluster_name> \-n <hosted_cluster_namespace> -oyaml
Verification
-
Verify that a node pool is using the multi-architecture release image in your
HostedControlPlaneresource by running the following command:$ oc get hostedcontrolplane -n clusters-example -oyamlExample output#...version:availableUpdates: nulldesired:image: quay.io/openshift-release-dev/ocp-release:4.20.0-multiurl: https://access.redhat.com/errata/RHBA-2024:4855version: 4.20.0history:- completionTime: "2024-07-28T13:10:58Z"image: quay.io/openshift-release-dev/ocp-release:4.20.0-multistartedTime: "2024-07-28T13:10:27Z"state: Completedverified: falseversion: 4.20.0noteThe multi-architecture OpenShift Container Platform release image is updated in your
HostedCluster,HostedControlPlaneresources, and hosted control plane pods. However, your existing node pools do not transition with the multi-architecture image automatically, because the release image transition is decoupled between the hosted cluster and node pools. You must create new node pools on your new multi-architecture hosted cluster.
Next steps
- Create node pools on the multi-architecture hosted cluster.
Creating node pools on the multi-architecture hosted cluster
After you transition your hosted cluster from single-architecture to multi-architecture, create node pools on compute machines based on 64-bit AMD and 64-bit ARM architectures.
Procedure
-
Create node pools based on 64-bit ARM architecture by entering the following command:
$ hcp create nodepool aws \--cluster-name <hosted_cluster_name> \--name <nodepool_name> \--node-count=<node_count> \--arch arm64- Replace
<hosted_cluster_name>with your hosted cluster name. - Replace
<nodepool_name>with your node pool name. - Replace
<node_count>with integer for your node count, for example,2.
- Replace
-
Create node pools based on 64-bit AMD architecture by entering the following command:
$ hcp create nodepool aws \--cluster-name <hosted_cluster_name> \--name <nodepool_name> \--node-count=<node_count> \--arch amd64- Replace
<hosted_cluster_name>with your hosted cluster name. - Replace
<nodepool_name>with your node pool name. - Replace
<node_count>with integer for your node count, for example,2.
- Replace
Verification
-
Verify that a node pool is using the multi-architecture release image by entering the following command:
$ oc get nodepool/<nodepool_name> -oyamlExample output for 64-bit AMD node pools#...spec:arch: amd64#...release:image: quay.io/openshift-release-dev/ocp-release:4.20.0-multiExample output for 64-bit ARM node pools#...spec:arch: arm64#...release:image: quay.io/openshift-release-dev/ocp-release:4.20.0-multi
Adding or updating AWS tags for a hosted cluster
As a cluster instance administrator, you can add or update Amazon Web Services (AWS) tags without needing to re-create your hosted cluster. Tags are key-value pairs that are attached to AWS resources for management and automation.
You might want to use tags for the following purposes:
- Managing access controls.
- Tracking chargeback or showback.
- Managing cloud Identity and Access Management (IAM) conditional permissions.
- Aggregating resources based on tags. For example, you can query tags to calculate resource usage and billing costs.
You can add or update tags for several different types of resources, including Amazon Elastic File System (EFS) access points, load balancer resources, Amazon Elastic Block Storage (EBS) volumes, IAM users, and AWS S3.
On network load balancers, tags cannot be added or updated. The AWS load balancer reconciles whatever tags are in the HostedCluster resource. If you try to add or update a tag, the load balancer overwrites the tag.
In addition, tags cannot be updated on the default security group resource that is created directly by hosted control planes.
Prerequisites
- You must have cluster administrator permissions for your hosted cluster on AWS.
Procedure
-
If you want to add or update tags for EFS access points, complete steps 1 and 2. If you are adding or updating tags for other types of resources, complete only step 2.
- In the
aws-efs-csi-driver-operatorservice account, add two annotations, as shown in the following example. These annotations are required so that the Amazon Elastic Kubernetes Service (EKS) pod identity webhook that runs on the cluster can correctly assign AWS roles to the pods that the EFS Operator uses.apiVersion: v1kind: ServiceAccountmetadata:name: <service_account_name>namespace: <project_name>annotations:eks.amazonaws.com/role-arn:<role_arn>eks.amazonaws.com/audience:sts.amazonaws.com - Delete the Operator pod or roll out a restart of the
aws-efs-csi-driver-operatordeployment.
- In the
-
In the
HostedClusterresource, enter information in theresourceTagsfields, as shown in the following example:Example HostedCluster resourceapiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:#...spec:autoscaling: {}clusterID: <cluster_id>configuration: {}controllerAvailabilityPolicy: SingleReplicadns:#...etcd:#...fips: falseinfraID: <infra_id>infrastructureAvailabilityPolicy: SingleReplicaissuerURL: https://<issuer_url>.s3.<region>.amazonaws.comnetworking:#...olmCatalogPlacement: managementplatform:aws:#...resourceTags:- key: kubernetes.io/cluster/<tag>value: ownedrolesRef:#...type: AWSReplace
<tag>with the tag that you want to add to your resource.
Configuring node pool capacity blocks on AWS
After you create a hosted cluster, you can configure node pool capacity blocks for graphics processing unit (GPU) reservations on Amazon Web Services (AWS).
Procedure
-
Create GPU reservations on AWS by running a command similar to the following example:
warningThe zone of the GPU reservation must match your hosted cluster zone.
$ aws ec2 describe-capacity-block-offerings \--instance-type "p4d.24xlarge"\--instance-count "1" \--start-date-range "$(date -u +"2025-07-21T10:14:39Z")" \--end-date-range "$(date -u -d "2 day" +"2025-07-22T10:16:36Z")" \--capacity-duration-hours 24 \--output json--instance-typedefines the type of your AWS instance.--instance-countdefines your instance purchase quantity. Valid values are integers ranging from1to64.--start-date-rangedefines the start date range.--end-date-rangedefines the end date range.--capacity-duration-hoursdefines the duration of capacity blocks in hours.
-
Purchase the minimum fee capacity block by running the following command:
$ aws ec2 purchase-capacity-block \--capacity-block-offering-id "${MIN_FEE_ID}" \--instance-platform "Linux/UNIX"\--tag-specifications 'ResourceType=capacity-reservation,Tags=[{Key=usage-cluster-type,Value=hypershift-hosted}]' \--output json > "${CR_OUTPUT_FILE}"--capacity-block-offering-iddefines the ID of the capacity block offering.--instance-platformdefines the platform of your instance.--tag-specificationsdefines the tag for your instance.
-
Create an environment variable to set the capacity reservation ID by running the following command:
$ CB_RESERVATION_ID=$(jq -r '.CapacityReservation.CapacityReservationId' "${CR_OUTPUT_FILE}")Wait for a couple of minutes for the GPU reservation to become available.
-
Add a node pool to use the GPU reservation by running the following command:
$ hcp create nodepool aws \--cluster-name <hosted_cluster_name> \--name <node_pool_name> \--node-count <node_pool_count> \--instance-type <instance_type> \--arch <arch_type> \--release-image <release_image> \--render > /tmp/np.yaml--cluster-namespecifies the name of your hosted cluster.--namespecifies the name of your node pool.--node-countdefines the node pool count, for example,1.--instance-typedefines the instance type, for example,p4d.24xlarge.--archdefines an architecture type, for example,amd64.--release-imagespecifies the release image you want to use.
-
Add the
capacityReservationsetting in yourNodePoolresource by using the following example configuration:# ...spec:arch: amd64clusterName: cb-np-hcpmanagement:autoRepair: falseupgradeType: Replaceplatform:aws:instanceProfile: cb-np-hcp-dqppw-workerinstanceType: p4d.24xlargerootVolume:size: 120type: gp3subnet:id: subnet-00000placement:capacityReservation:id: ${CB_RESERVATION_ID}marketType: CapacityBlockstype: AWS# ... -
Apply the node pool configuration by running the following command:
$ oc apply -f /tmp/np.yaml
Verification
-
Verify that your new node pool is created successfully by running the following command:
$ oc get np -n clustersExample outputNAMESPACE NAME CLUSTER DESIRED NODES CURRENT NODES AUTOSCALING AUTOREPAIR VERSION UPDATINGVERSION UPDATINGCONFIG MESSAGEclusters cb-np cb-np-hcp 1 1 False False 4.22.0-0.nightly-2025-06-05-224220 False False -
Verify that your new compute nodes are created in the hosted cluster by running the following command:
$ oc get nodesExample outputNAME STATUS ROLES AGE VERSIONip-10-0-132-74.ec2.internal Ready worker 17m v1.35.4ip-10-0-134-183.ec2.internal Ready worker 4h5m v1.35.4
Deleting a hosted cluster after configuring node pool capacity blocks
After you configure node pool capacity blocks, you can optionally delete a hosted cluster and uninstall the HyperShift Operator.
Procedure
- To delete a hosted cluster, run a command similar to the following example:
$ hcp destroy cluster aws \--name cb-np-hcp \--aws-creds $HOME/.aws/credentials \--namespace clusters \--region us-east-2
- To uninstall the HyperShift Operator, run the following command:
$ hcp install render --format=yaml | oc delete -f -
Amazon Spot Instance support for node pools
To reduce cloud infrastructure costs for non-critical and fault-tolerant workloads, you can use Amazon Spot Instances for your compute nodes in hosted clusters.
Spot Instances use spare Amazon Elastic Compute Cloud (Amazon EC2) capacity at reduced prices compared to on-demand instances. When Amazon EC2 needs the capacity block, it can interrupt a Spot Instance with a 2-minute warning. You can use Spot Instances for node pools in hosted clusters on AWS. However, you cannot combine Spot Instances with Amazon EC2 Capacity Reservations.
Spot Instances are suitable for fault-tolerant, stateless, and flexible workloads. Do not use Spot Instances for workloads that cannot tolerate interruptions. Spot Instances require default tenancy. Dedicated tenancy is not supported.
Configuring Amazon SQS and Amazon EventBridge to receive Amazon EC2 interruption events
Before you enable Amazon Spot Instances, you must set up Amazon Simple Queue Service (SQS) and Amazon EventBridge to receive Amazon EC2 interruption events.
For hosted control planes to be able to provide a graceful shut down, it needs to poll the SQS queue and cordon or drain nodes before they are deleted. For more information about Amazon SQS and Amazon EventBridge, see "Getting started with Amazon SQS" and "Getting started: Create an Amazon EventBridge bus rule".
Procedure
-
Create an Amazon SQS queue to receive Spot Instance notifications by entering the following command:
CLUSTER_NAME="my-cluster"AWS_REGION="us-east-1"$ aws sqs create-queue \--queue-name "${CLUSTER_NAME}-spot-interruption-queue" \--region "${AWS_REGION}"- Be sure to replace
my-clusterandus-east-1with your cluster name and AWS region. - Take note of the URL from the output. You need the URL when you create the hosted cluster.
- Be sure to replace
-
Create Amazon EventBridge rules to route Amazon EC2 Spot interruption warnings and rebalance recommendations to the SQS queue.
- Set the attributes by entering the following command:
ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)QUEUE_URL="https://sqs.${AWS_REGION}.amazonaws.com/${ACCOUNT_ID}/${CLUSTER_NAME}-spot-interruption-queue"QUEUE_ARN=$(aws sqs get-queue-attributes \--queue-url "${QUEUE_URL}" \--attribute-names QueueArn \--query 'Attributes.QueueArn' --output text)
- Set the interruption warning by entering the following command:
$ aws events put-rule \--name "${CLUSTER_NAME}-spot-interruption-warning" \--event-pattern '{"source":["aws.ec2"],"detail-type":["EC2 Spot Instance Interruption Warning"]}' \--region "${AWS_REGION}"
- Set the target for the interruption warning by entering the following command:
$ aws events put-targets \--rule "${CLUSTER_NAME}-spot-interruption-warning" \--targets "Id=1,Arn=${QUEUE_ARN}" \--region "${AWS_REGION}"
- Set the recommendations by entering the following command:
$ aws events put-rule \--name "${CLUSTER_NAME}-rebalance-recommendation" \--event-pattern '{"source":["aws.ec2"],"detail-type":["EC2 Instance Rebalance Recommendation"]}' \--region "${AWS_REGION}"
- Set the target for the recommendations by entering the following command:
$ aws events put-targets \--rule "${CLUSTER_NAME}-rebalance-recommendation" \--targets "Id=1,Arn=${QUEUE_ARN}" \--region "${AWS_REGION}"
- Set the attributes by entering the following command:
-
Configure the Amazon SQS queue policy to allow EventBridge to send messages to the queue by entering the following command:
ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)QUEUE_URL="https://sqs.${AWS_REGION}.amazonaws.com/${ACCOUNT_ID}/${CLUSTER_NAME}-spot-interruption-queue"$ aws sqs set-queue-attributes \--queue-url "${QUEUE_URL}" \--attributes '{"Policy": "{\"Version\":\"2012-10-17\",\"Statement\":[{\"Effect\":\"Allow\",\"Principal\":{\"Service\":\"events.amazonaws.com\"},\"Action\":\"sqs:SendMessage\",\"Resource\":\"'${QUEUE_ARN}'\"}]}"}'
Additional resources
Enabling Amazon Spot Instance support for node pools
Enable Amazon Spot Instances for your compute nodes in hosted cluster node pools to reduce cloud infrastructure costs for non-critical and fault-tolerant workloads.
Prerequisites
- You set up Amazon Simple Queue Service (SQS) and Amazon EventBridge to receive Amazon EC2 interruption events.
- You have the URL of your Amazon SQS queue, which you created in "Configuring Amazon SQS and Amazon EventBridge to receive Amazon EC2 interruption events".
Procedure
-
Set the Amazon SQS queue URL on your hosted cluster.
- If you are setting the SQS queue URL on a new hosted cluster, take the following steps:
-
Configure the
HostedClusterresource as follows:apiVersion: hypershift.openshift.io/v1beta1kind: HostedClustermetadata:name: <my_hosted_cluster>namespace: <my_hosted_cluster_namespace>spec:platform:type: AWSaws:region: us-east-1terminationHandlerQueueURL: "https://sqs.us-east-1.amazonaws.com/123456789012/my-cluster-spot-interruption-queue"...where:
spec.platform.aws.terminationHandlerQueueURL- Specifies the SQS queue URL.
-
Apply the configuration by entering the following command:
$ oc apply -f <hosted_cluster_config>.yaml
-
- If you are setting the SQS queue URL on an existing hosted cluster, patch the
HostedClusterresource as follows:$ oc patch hostedcluster my-cluster -n clusters --type merge -p '{"spec": {"platform": {"aws": {"terminationHandlerQueueURL": "https://sqs.us-east-1.amazonaws.com/123456789012/my-cluster-spot-interruption-queue"}}}}'
- If you are setting the SQS queue URL on a new hosted cluster, take the following steps:
-
Create a
NodePoolobject that has the Spot market type configured, as shown in the following example:apiVersion: hypershift.openshift.io/v1beta1kind: NodePoolmetadata:name: spot-workersnamespace: <my_hosted_cluster_namespace>spec:clusterName: <my_hosted_cluster>replicas: 3release:image: quay.io/openshift-release-dev/ocp-release:4.22.0-x86_64management:autoRepair: trueupgradeType: Replaceplatform:type: AWSaws:instanceType: m5.xlargeinstanceProfile: my-cluster-workerrootVolume:size: 120type: gp3placement:marketType: SpotmarketType: Spot- Specifies that all
Machineobjects andNodeobjects have thehypershift.openshift.io/interruptible-instancelabel and that the EC2 instances have theaws-node-termination-handler/managedtag so they can be identified by the termination-handling components. You can specifyspotoptions only when themarketTypefield is set toSpot.
-
Apply the configuration by entering the following command:
$ oc apply -f <node_pool_config>.yaml