Managing hosted control planes on OpenShift Virtualization
After you deploy a hosted cluster on OpenShift Virtualization, you can manage the cluster.
Accessing the hosted cluster
You can access the hosted cluster by either getting the kubeconfig file and kubeadmin credential directly from resources, or by using the hcp command-line interface to generate a kubeconfig file.
Prerequisites
To access the hosted cluster by getting the kubeconfig file and credentials directly from resources, you must be familiar with the access secrets for hosted clusters. The hosted cluster (hosting) namespace has hosted cluster resources and the access secrets. The hosted control plane namespace is where the hosted control plane runs.
The secret name formats are as follows:
kubeconfigsecret:<hosted_cluster_namespace>-<name>-admin-kubeconfig(clusters-hypershift-demo-admin-kubeconfig)kubeadminpassword secret:<hosted_cluster_namespace>-<name>-kubeadmin-password(clusters-hypershift-demo-kubeadmin-password)
The kubeconfig secret has a Base64-encoded kubeconfig field, which you can decode and save into a file to use with the following command:
$ oc --kubeconfig <hosted_cluster_name>.kubeconfig get nodesThe kubeadmin password secret is also Base64-encoded. You can decode it and use the password to log in to the API server or console of the hosted cluster.
To access the hosted cluster by using the hcp CLI to generate the kubeconfig file, take the following steps.
Procedure
- Generate the
kubeconfigfile by entering the following command:terminal$ hcp create kubeconfig --namespace <hosted_cluster_namespace> \ --name <hosted_cluster_name> > <hosted_cluster_name>.kubeconfig - After you save the
kubeconfigfile, you can access the hosted cluster by entering the following example command:terminal$ oc --kubeconfig <hosted_cluster_name>.kubeconfig get nodes
Enabling node auto-scaling for the hosted cluster
When you need more capacity in your hosted cluster and spare agents are available, you can enable auto-scaling to install new worker nodes.
Procedure
- To enable auto-scaling, enter the following command:terminal
$ oc -n <hosted_cluster_namespace> patch nodepool <hosted_cluster_name> \ --type=json \ -p '[{"op": "remove", "path": "/spec/replicas"},{"op":"add", "path": "/spec/autoScaling", "value": { "max": 5, "min": 2 }}]'NoteIn the example, the minimum number of nodes is 2, and the maximum is 5. The maximum number of nodes that you can add might be bound by your platform. For example, if you use the Agent platform, the maximum number of nodes is bound by the number of available agents.
- Create a workload that requires a new node.
- Create a YAML file that has the workload configuration, by using the following example:yaml
apiVersion: apps/v1 kind: Deployment metadata: creationTimestamp: null labels: app: reversewords name: reversewords namespace: default spec: replicas: 40 selector: matchLabels: app: reversewords strategy: {} template: metadata: creationTimestamp: null labels: app: reversewords spec: containers: - image: quay.io/mavazque/reversewords:latest name: reversewords resources: requests: memory: 2Gi status: {} - Save the file as
workload-config.yaml. - Apply the YAML by entering the following command:terminal
$ oc apply -f workload-config.yaml
- Create a YAML file that has the workload configuration, by using the following example:
- Extract the
admin-kubeconfigsecret by entering the following command:terminal$ oc extract -n <hosted_cluster_namespace> \ secret/<hosted_cluster_name>-admin-kubeconfig \ --to=./hostedcluster-secrets --confirmExample outputhostedcluster-secrets/kubeconfig - You can check if new nodes are in the
Readystatus by entering the following command:terminal$ oc --kubeconfig ./hostedcluster-secrets get nodes - To remove the node, delete the workload by entering the following command:terminal
$ oc --kubeconfig ./hostedcluster-secrets -n <namespace> \ delete deployment <deployment_name> - Wait for several minutes to pass without requiring the additional capacity. On the Agent platform, the agent is decommissioned and can be reused. You can confirm that the node was removed by entering the following command:terminal
$ oc --kubeconfig ./hostedcluster-secrets get nodesNoteFor IBM Z(R) agents, if you are using an OSA network device in Processor Resource/Systems Manager (PR/SM) mode, auto scaling is not supported. You must delete the old agent manually and scale up the node pool because the new agent joins during the scale down process.
Storage for hosted control planes on OpenShift Virtualization
If you do not provide any advanced storage configuration, the default storage class is used for the KubeVirt virtual machine (VM) images, the KubeVirt Container Storage Interface (CSI) mapping, and the etcd volumes.
The following table lists the capabilities that the infrastructure must provide to support persistent storage in a hosted cluster:
Persistent storage modes in a hosted cluster
| Infrastructure CSI provider | Hosted cluster CSI provider | Hosted cluster capabilities |
|---|---|---|
Any RWX Block CSI provider |
kubevirt-csi |
|
Any RWX Block CSI provider |
Red Hat OpenShift Data Foundation | Red Hat OpenShift Data Foundation feature set. External mode has a smaller footprint and uses a standalone Red Hat Ceph Storage. Internal mode has a larger footprint, but is self-contained and suitable for use cases that require expanded capabilities such as RWX File. |
OpenShift Virtualization handles storage on hosted clusters, which especially helps customers whose requirements are limited to block storage.
Mapping KubeVirt CSI storage classes
You can map the infrastructure storage class to the hosted storage class during cluster creation.
KubeVirt CSI supports mapping only an infrastructure storage class that is capable of ReadWriteMany (RWX) access.
The following table shows how volume and access mode capabilities map to KubeVirt CSI storage classes:
Mapping KubeVirt CSI storage classes to access and volume modes
| Infrastructure CSI capability | Hosted cluster CSI capability | VM live migration support | Notes |
|---|---|---|---|
RWX: Block or Filesystem | ReadWriteOnce (RWO) Block or Filesystem RWX Block only | Supported | Use Block mode because Filesystem volume mode results in degraded hosted Block mode performance. RWX Block volume mode is supported only when the hosted cluster is OpenShift Container Platform 4.16 or later. |
RWO Block storage | RWO Block storage or Filesystem | Not supported | Lack of live migration support affects the ability to update the underlying infrastructure cluster that hosts the KubeVirt VMs. |
RWO FileSystem | RWO Block or Filesystem | Not supported | Lack of live migration support affects the ability to update the underlying infrastructure cluster that hosts the KubeVirt VMs. Use of the infrastructure Filesystem volume mode results in degraded hosted Block mode performance. |
Procedure
- To map the infrastructure storage class to the hosted storage class, use the
--infra-storage-class-mappingargument as shown in the following example:terminal$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 2 \ --pull-secret /user/name/pullsecret \ --memory 8Gi \ --cores 2 \ --infra-storage-class-mapping=<infrastructure_storage_class>/<hosted_storage_class>--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--infra-storage-class-mappingspecifies the storage class names for the infrastructure and hosted cluster. Replace<infrastructure_storage_class>with the infrastructure storage class name and<hosted_storage_class>with the hosted cluster storage class name. You can use the--infra-storage-class-mappingargument multiple times within thehcp create clustercommand.After you create the hosted cluster, the infrastructure storage class is visible within the hosted cluster. When you create a Persistent Volume Claim (PVC) within the hosted cluster that uses one of those storage classes, KubeVirt CSI provisions that volume by using the infrastructure storage class mapping that you configured during cluster creation.
Mapping a single KubeVirt CSI volume snapshot class
You can expose your infrastructure volume snapshot class to the hosted cluster by using KubeVirt CSI.
Procedure
- To map your volume snapshot class to the hosted cluster, use the
--infra-volumesnapshot-class-mappingargument when creating a hosted cluster as shown in the following example:terminal$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 2 \ --pull-secret /user/name/pullsecret \ --memory 8Gi \ --cores 2 \ --infra-storage-class-mapping=<infrastructure_storage_class>/<hosted_storage_class> \ --infra-volumesnapshot-class-mapping=<infrastructure_volume_snapshot_class>/<hosted_volume_snapshot_class>--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--infra-storage-class-mappingspecifies the storage classes. Replace<infrastructure_storage_class>with the storage class in the infrastructure cluster, and replace<hosted_storage_class>with the storage class in the hosted cluster.--infra-volumesnapshot-class-mappingspecifies the volume snapshot classes. Replace<infrastructure_volume_snapshot_class>with the volume snapshot class in the infrastructure cluster, and replace<hosted_volume_snapshot_class>with the volume snapshot class in the hosted cluster.NoteIf you do not use the
--infra-storage-class-mappingand--infra-volumesnapshot-class-mappingarguments, a hosted cluster is created with the default storage class and the volume snapshot class. Therefore, you must set the default storage class and the volume snapshot class in the infrastructure cluster.
Mapping multiple KubeVirt CSI volume snapshot classes
You can map multiple volume snapshot classes to the hosted cluster by assigning them to a specific group. The infrastructure storage class and the volume snapshot class are compatible with each other only if they belong to a same group.
Procedure
- To map multiple volume snapshot classes to the hosted cluster, use the
groupoption when creating a hosted cluster, as shown in the following example:terminal$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 2 \ --pull-secret /user/name/pullsecret \ --memory 8Gi \ --cores 2 \ --infra-storage-class-mapping=<infrastructure_storage_class>/<hosted_storage_class>,group=<group_name> \ --infra-storage-class-mapping=<infrastructure_storage_class>/<hosted_storage_class>,group=<group_name> \ --infra-storage-class-mapping=<infrastructure_storage_class>/<hosted_storage_class>,group=<group_name> \ --infra-volumesnapshot-class-mapping=<infrastructure_volume_snapshot_class>/<hosted_volume_snapshot_class>,group=<group_name> \ --infra-volumesnapshot-class-mapping=<infrastructure_volume_snapshot_class>/<hosted_volume_snapshot_class>,group=<group_name>--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--infra-storage-class-mappingspecifies the storage classes. Replace<infrastructure_storage_class>with the storage class in the infrastructure cluster,<hosted_storage_class>with the storage class in the hosted cluster, and<group_name>with the group name. For example,infra-storage-class-mygroup/hosted-storage-class-mygroup,group=mygroupandinfra-storage-class-mymap/hosted-storage-class-mymap,group=mymap.--infra-volumesnapshot-class-mappingspecifies the volume snapshot classes. Replace<infrastructure_volume_snapshot_class>with the volume snapshot class in the infrastructure cluster and<hosted_volume_snapshot_class>with the volume snapshot class in the hosted cluster. For example,infra-vol-snap-mygroup/hosted-vol-snap-mygroup,group=mygroupandinfra-vol-snap-mymap/hosted-vol-snap-mymap,group=mymap.
Configuring KubeVirt VM root volume
At cluster creation time, you can configure the storage class that is used to host the KubeVirt virtual machine (VM) root volumes by using the --root-volume-storage-class argument.
Procedure
- To set a custom storage class and volume size for KubeVirt VMs, run a command as shown in the following example:terminal
$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 2 \ --pull-secret /user/name/pullsecret \ --memory 8Gi \ --cores 2 \ --root-volume-storage-class ocs-storagecluster-ceph-rbd \ --root-volume-size 64--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--root-volume-storage-classspecifies a name of the storage class to host the KubeVirt VM root volumes.--root-volume-sizespecifies the volume size.As a result, you get a hosted cluster created with VMs hosted on persistent volume claims (PVCs).
Enabling KubeVirt VM image caching
To optimize both cluster startup time and storage usage, you can use KubeVirt virtual machine (VM) image caching.
KubeVirt VM image caching supports the use of a storage class that is capable of smart cloning and the ReadWriteMany access mode. For more information about smart cloning, see "Cloning a data volume using smart-cloning".
Image caching works as follows:
- The VM image is imported to a persistent volume claim (PVC) that is associated with the hosted cluster.
- A unique clone of that PVC is created for every KubeVirt VM that is added as a worker node to the cluster.
Image caching reduces VM startup time by requiring only a single image import. It can further reduce overall cluster storage usage when the storage class supports copy-on-write cloning.
Procedure
- To enable image caching, during cluster creation, use the
--root-volume-cache-strategy=PVCargument by running a command as shown in the following example:terminal$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 2 \ --pull-secret /user/name/pullsecret \ --memory 8Gi \ --cores 2 \ --root-volume-cache-strategy=PVC--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--root-volume-cache-strategyspecifies a strategy for image caching.
KubeVirt CSI storage security and isolation
KubeVirt Container Storage Interface (CSI) extends the storage capabilities of the underlying infrastructure cluster to hosted clusters.
The CSI driver ensures secure and isolated access to the infrastructure storage classes and hosted clusters by using the following security constraints:
- The storage of a hosted cluster is isolated from the other hosted clusters.
- Compute nodes in a hosted cluster do not have a direct API access to the infrastructure cluster. The hosted cluster can provision storage on the infrastructure cluster only through the controlled KubeVirt CSI interface.
- The hosted cluster does not have access to the KubeVirt CSI cluster controller. As a result, the hosted cluster cannot access arbitrary storage volumes on the infrastructure cluster that are not associated with the hosted cluster. The KubeVirt CSI cluster controller runs in a pod in the hosted control plane namespace.
- Role-based access control (RBAC) of the KubeVirt CSI cluster controller limits the persistent volume claim (PVC) access to only the hosted control plane namespace. Therefore, KubeVirt CSI components cannot access storage from the other namespaces.
Configuring etcd storage
At cluster creation time, you can configure the storage class that is used to host etcd data by using the --etcd-storage-class argument.
Procedure
- To configure a storage class for etcd, run a command similar to the following example:terminal
$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 2 \ --pull-secret /user/name/pullsecret \ --memory 8Gi \ --cores 2 \ --etcd-storage-class=lvm-storageclass--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--etcd-storage-classspecifies the etcd storage class name. If you do not provide an--etcd-storage-classargument, the default storage class is used.
Attaching NVIDIA GPU devices by using the hcp CLI
You can attach one or more NVIDIA graphics processing unit (GPU) devices to node pools by using the hcp command-line interface (CLI) in a hosted cluster on OpenShift Virtualization.
Attaching NVIDIA GPU devices to node pools is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
Prerequisites
- You have exposed the NVIDIA GPU device as a resource on the node where the GPU device resides. For more information, see NVIDIA GPU Operator with OpenShift Virtualization.
- You have exposed the NVIDIA GPU device as an extended resource on the node to assign it to node pools.
Procedure
- You can attach the GPU device to node pools during cluster creation by running a command similar to the following example:terminal
$ hcp create cluster kubevirt \ --name my-hosted-cluster \ --node-pool-replicas 3 \ --pull-secret /user/name/pullsecret \ --memory 16Gi \ --cores 2 \ --host-device-name="nvidia-a100,count:2"--namespecifies the name of your hosted cluster.--node-pool-replicasspecifies the worker count.--pull-secretspecifies the path to your pull secret.--memoryspecifies a value for memory.--coresspecifies a value for CPU.--host-device-namespecifies the GPU device name and the count. The--host-device-nameargument takes the name of the GPU device from the infrastructure node and an optional count that represents the number of GPU devices you want to attach to each virtual machine (VM) in node pools. The default count is1. For example, if you attach 2 GPU devices to 3 node pool replicas, all 3 VMs in the node pool are attached to the 2 GPU devices.TipYou can use the
--host-device-nameargument multiple times to attach multiple devices of different types.
Attaching NVIDIA GPU devices by using the NodePool resource
You can attach one or more NVIDIA graphics processing unit (GPU) devices to node pools by configuring the nodepool.spec.platform.kubevirt.hostDevices field in the NodePool resource.
Attaching NVIDIA GPU devices to node pools is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
Procedure
- To attach a single GPU device, configure the
NodePoolresource by using the following example configuration:yamlapiVersion: hypershift.openshift.io/v1beta1 kind: NodePool metadata: name: <hosted_cluster_name> namespace: <hosted_cluster_namespace> spec: arch: amd64 clusterName: <hosted_cluster_name> management: autoRepair: false upgradeType: Replace nodeDrainTimeout: 0s nodeVolumeDetachTimeout: 0s platform: kubevirt: attachDefaultNetwork: true compute: cores: <cpu> memory: <memory> hostDevices: - count: <count> deviceName: <gpu_device_name> networkInterfaceMultiqueue: Enable rootVolume: persistent: size: 32Gi type: Persistent type: KubeVirt replicas: <worker_node_count><hosted_cluster_name>specifies the name of your hosted cluster; for example,my-hosted-cluster.<hosted_cluster_namespace>specifies the name of the hosted cluster namespace; for example,my-hc-namespace.<cpu>specifies a value for CPU; for example,2.<memory>specifies a value for memory; for example,16Gi.<count>specifies the number of GPU devices you want to attach to each virtual machine (VM) in node pools. For example, if you attach 2 GPU devices to 3 node pool replicas, all 3 VMs in the node pool are attached to the 2 GPU devices. The default count is1. ThehostDevicesfield defines a list of different types of GPU devices that you can attach to node pools.<gpu_device_name>specifies the GPU device name; for example,nvidia-a100.<worker_node_count>specifies the worker count; for example,3.
- To attach multiple GPU devices, configure the
NodePoolresource by using the following example configuration:yamlapiVersion: hypershift.openshift.io/v1beta1 kind: NodePool metadata: name: <hosted_cluster_name> namespace: <hosted_cluster_namespace> spec: arch: amd64 clusterName: <hosted_cluster_name> management: autoRepair: false upgradeType: Replace nodeDrainTimeout: 0s nodeVolumeDetachTimeout: 0s platform: kubevirt: attachDefaultNetwork: true compute: cores: <cpu> memory: <memory> hostDevices: - count: <count> deviceName: <gpu_device_name> - count: <count> deviceName: <gpu_device_name> - count: <count> deviceName: <gpu_device_name> - count: <count> deviceName: <gpu_device_name> networkInterfaceMultiqueue: Enable rootVolume: persistent: size: 32Gi type: Persistent type: KubeVirt replicas: <worker_node_count><hosted_cluster_name>specifies the name of your hosted cluster; for example,my-hosted-cluster.<hosted_cluster_namespace>specifies the name of the hosted cluster namespace; for example,my-hc-namespace.<cpu>specifies a value for CPU; for example,2.<memory>specifies a value for memory; for example,16Gi.<count>specifies the number of GPU devices you want to attach to each VM in node pools. For example, if you attach 2 GPU devices to 3 node pool replicas, all 3 VMs in the node pool are attached to the 2 GPU devices. The default count is1. ThehostDevicesfield defines a list of different types of GPU devices that you can attach to node pools.<gpu_device_name>specifies the GPU device name; for example,nvidia-a100.<worker_node_count>specifies the worker count; for example,3.
Evicting KubeVirt virtual machines
In cases where KubeVirt virtual machines (VMs) cannot be live migrated, such as when you use GPU passthrough, the VMs must be evicted at the same time as the NodePool resource of the hosted cluster.
Otherwise, the compute nodes might be shut down without being drained from the workload. This might also happen when you are upgrading the OpenShift Virtualization Operator.
To achieve a synchronized restart, you can set the evictionStrategy parameter on the hyperconverged resource to ensure that only VMs that are drained from workloads are rebooted.
Procedure
- To learn more about the
hyperconvergedresource and the allowed values for theevictionStrategyparameter, enter the following command:terminal$ oc explain --api-version=hco.kubevirt.io/v1beta1 hyperconverged.spec.evictionStrategy - Patch the
hyperconvergedresource by entering the following command:terminal$ oc patch -n openshift-cnv hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged \ --type=merge \ -p '{"spec": {"evictionStrategy": "External"}}' - Patch the workload update strategy and the workload update methods by entering the following command:terminal
$ oc patch -n openshift-cnv hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged \ --type=merge \ -p '{"spec": {"workloadUpdateStrategy": {"workloadUpdateMethods": ["LiveMigrate","Evict"]}}}'By applying this patch, you specify that VMs should be live-migrated if possible, and that only the VMs that cannot be live-migrated should be evicted.
Verification
- Check whether the patch command was applied properly by entering the following command:terminal
$ oc get -n openshift-cnv hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged -ojsonpath='{.spec.evictionStrategy}'Example outputExternal
Spreading node pool VMs by using topologySpreadConstraint
In some scenarios, node pool virtual machines (VMs) might run on the same node, which can cause availability issues. To avoid distribution of VMs on a single node, use the descheduler to continuously honor the topologySpreadConstraint constraint to spread VMs on multiple nodes.
By default, KubeVirt VMs created by a node pool are scheduled on any available nodes that have the capacity to run the VMs. The topologySpreadConstraint constraint is set to schedule VMs on multiple nodes.
Prerequisites
- You installed the Kube Descheduler Operator. For more information, see "Installing the descheduler".
Procedure
- Open the
KubeDeschedulercustom resource (CR) by entering the following command, and then modify theKubeDeschedulerCR to use theSoftTopologyAndDuplicatesandKubeVirtRelieveAndMigrateprofiles so that you maintain thetopologySpreadConstraintconstraint settings.The
KubeDeschedulerCR namedclusterruns in theopenshift-kube-descheduler-operatornamespace.terminal$ oc edit kubedescheduler cluster -n openshift-kube-descheduler-operatorExample KubeDescheduler configurationapiVersion: operator.openshift.io/v1 kind: KubeDescheduler metadata: name: cluster namespace: openshift-kube-descheduler-operator spec: mode: Automatic managementState: Managed deschedulingIntervalSeconds: 30 profiles: - SoftTopologyAndDuplicates - KubeVirtRelieveAndMigrate profileCustomizations: devDeviationThresholds: AsymmetricLow devActualUtilizationProfile: PrometheusCPUCombined # ...where:
spec.deschedulingIntervalSecondsSets the number of seconds between the descheduler running cycles.
spec.profilesThe
SoftTopologyAndDuplicatesprofile evicts pods that follow thewhenUnsatisfiable: ScheduleAnywaysoft topology constraint. TheKubeVirtRelieveAndMigrateprofile balances resource usage between nodes and enables strategies, such asRemovePodsHavingTooManyRestartsandLowNodeUtilization.