Managing nodes
OpenShift Container Platform uses a KubeletConfig custom resource (CR) to manage the configuration of nodes. By creating an instance of a KubeletConfig object, a managed machine config is created to override setting on the node.
Logging in to remote machines for the purpose of changing their configuration is not supported.
Modifying nodes
To make configuration changes to a cluster, or machine pool, you must create a custom resource definition (CRD), or kubeletConfig object. OpenShift Container Platform uses the Machine Config Controller to watch for changes introduced through the CRD to apply the changes to the cluster.
Most Kubelet Configuration options can be set by the user. However, you cannot overwrite the following options:
- CgroupDriver
- ClusterDNS
- ClusterDomain
- StaticPodPath
If a single node contains more than 50 images, pod scheduling might be imbalanced across nodes. This is because the list of images on a node is shortened to 50 by default. You can disable the image limit by editing the KubeletConfig object and setting the value of nodeStatusMaxImages to -1.
Because the fields in a kubeletConfig object are passed directly to the kubelet from upstream Kubernetes, the validation of those fields is handled directly by the kubelet itself. Please refer to the relevant Kubernetes documentation for the valid values for these fields. Invalid values in the kubeletConfig object can render cluster nodes unusable.
Procedure
-
Obtain the label associated with the static CRD, Machine Config Pool, for the type of node you want to configure. Perform one of the following steps:
-
Check current labels of the desired machine config pool. For example:
$ oc get machineconfigpool --show-labelsExample outputNAME CONFIG UPDATED UPDATING DEGRADED LABELSmaster rendered-master-e05b81f5ca4db1d249a1bf32f9ec24fd True False False operator.machineconfiguration.openshift.io/required-for-upgrade=worker rendered-worker-f50e78e1bc06d8e82327763145bfcf62 True False False -
Add a custom label to the desired machine config pool. For example:
$ oc label machineconfigpool worker custom-kubelet=enabled
-
-
Create a
kubeletconfigcustom resource (CR) for your configuration change, as demonstrated in the following sample configuration for acustom-configCR:apiVersion: machineconfiguration.openshift.io/v1kind: KubeletConfigmetadata:name: custom-configspec:machineConfigPoolSelector:matchLabels:custom-kubelet: enabledkubeletConfig:podsPerCore: 10maxPods: 250systemReserved:cpu: 2000mmemory: 1Gi#...where:
name- Assign a name to CR.
custom-kubelet- Specify the label to apply the configuration change, this is the label you added to the machine config pool.
kubeletConfig- Specify the new value(s) you want to change.
-
Create the CR object:
$ oc create -f <file-name>For example:
$ oc create -f master-kube-config.yaml
Configuring control plane nodes as schedulable
You can configure control plane nodes to be schedulable, meaning that new pods are allowed for placement on the control plane nodes.
By default, control plane nodes are not schedulable. You can set the control plane nodes to be schedulable, but you must retain the compute nodes.
You can deploy OpenShift Container Platform with no compute nodes on a bare-metal cluster. In this case, the control plane nodes are marked schedulable by default.
You can allow or disallow control plane nodes to be schedulable by configuring the mastersSchedulable field.
When you configure control plane nodes from the default unschedulable to schedulable, additional subscriptions are required. This is because control plane nodes then become compute nodes.
Procedure
-
Edit the
schedulers.config.openshift.ioresource.$ oc edit schedulers.config.openshift.io cluster -
Configure the
mastersSchedulablefield.apiVersion: config.openshift.io/v1kind: Schedulermetadata:creationTimestamp: "2019-09-10T03:04:05Z"generation: 1name: clusterresourceVersion: "433"selfLink: /apis/config.openshift.io/v1/schedulers/clusteruid: a636d30a-d377-11e9-88d4-0a60097bee62spec:mastersSchedulable: falsestatus: {}#...where:
spec.mastersSchedulable- Specifies whether the control plane nodes are schedulable. Set to
trueto allow control plane nodes to be schedulable, orfalseto disallow control plane nodes from being schedulable.
-
Save the file to apply the changes.
Setting SELinux booleans
OpenShift Container Platform allows you to enable and disable an SELinux boolean on a Red Hat Enterprise Linux CoreOS (RHCOS) node. The following procedure explains how to modify SELinux booleans on nodes using the Machine Config Operator (MCO). This procedure uses container_manage_cgroup as the example boolean. You can modify this value to whichever boolean you need.
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Create a new YAML file with a
MachineConfigobject, displayed in the following example:apiVersion: machineconfiguration.openshift.io/v1kind: MachineConfigmetadata:labels:machineconfiguration.openshift.io/role: workername: 99-worker-setseboolspec:config:ignition:version: 3.2.0systemd:units:- contents: |[Unit]Description=Set SELinux booleansBefore=kubelet.service[Service]Type=oneshotExecStart=/sbin/setsebool container_manage_cgroup=onRemainAfterExit=true[Install]WantedBy=multi-user.target graphical.targetenabled: truename: setsebool.service#... -
Create the new
MachineConfigobject by running the following command:$ oc create -f 99-worker-setsebool.yamlnoteApplying any changes to the
MachineConfigobject causes all affected nodes to gracefully reboot after the change is applied.
Add kernel arguments to nodes
In some special cases, you can add kernel arguments to a set of nodes in your cluster to customize the kernel behavior to meet specific needs you might have.
You should add kernel arguments with caution and a clear understanding of the implications of the arguments you set.
Improper use of kernel arguments can result in your systems becoming unbootable.
Examples of kernel arguments you could set include:
-
nosmt: Disables symmetric multithreading (SMT) in the kernel. Multithreading allows multiple logical threads for each CPU. You could consider
nosmtin multi-tenant environments to reduce risks from potential cross-thread attacks. By disabling SMT, you essentially choose security over performance. -
enforcing=0: Configures Security Enhanced Linux (SELinux) to run in permissive mode. In permissive mode, the system acts as if SELinux is enforcing the loaded security policy, including labeling objects and emitting access denial entries in the logs, but it does not actually deny any operations. While not supported for production systems, permissive mode can be helpful for debugging.
warningDisabling SELinux on RHCOS in production is not supported. After SELinux has been disabled on a node, it must be re-provisioned before re-inclusion in a production cluster.
See Kernel.org kernel parameters for a list and descriptions of kernel arguments.
In the following procedure, you create a MachineConfig object that identifies:
- A set of machines to which you want to add the kernel argument. In this case, machines with a worker role.
- Kernel arguments that are appended to the end of the existing kernel arguments.
- A label that indicates where in the list of machine configs the change is applied.
Prerequisites
- You have
cluster-adminprivileges. - Your cluster is running.
Procedure
-
List existing
MachineConfigobjects for your OpenShift Container Platform cluster to determine how to label your machine config:$ oc get MachineConfigExample outputNAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE00-master 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m00-worker 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-master-container-runtime 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-master-kubelet 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-worker-container-runtime 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-worker-kubelet 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m99-master-generated-registries 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m99-master-ssh 3.2.0 40m99-worker-generated-registries 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m99-worker-ssh 3.2.0 40mrendered-master-23e785de7587df95a4b517e0647e5ab7 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33mrendered-worker-5d596d9293ca3ea80c896a1191735bb1 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m -
Create a
MachineConfigobject file that identifies the kernel argument (for example,05-worker-kernelarg-selinuxpermissive.yaml)apiVersion: machineconfiguration.openshift.io/v1kind: MachineConfigmetadata:labels:machineconfiguration.openshift.io/role: workername: 05-worker-kernelarg-selinuxpermissivespec:kernelArguments:- enforcing=0where:
machineconfiguration.openshift.io/role- Specifies a label to apply changes to specific nodes.
name- Specifies a name to identify where it fits among the machine configs (05) and what it does (adds a kernel argument to configure SELinux permissive mode).
kernelArguments- Specifies the exact kernel argument as
enforcing=0.
-
Create the new machine config:
$ oc create -f 05-worker-kernelarg-selinuxpermissive.yaml -
Check the machine configs to see that the new one was added:
$ oc get MachineConfigExample outputNAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE00-master 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m00-worker 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-master-container-runtime 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-master-kubelet 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-worker-container-runtime 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m01-worker-kubelet 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m05-worker-kernelarg-selinuxpermissive 3.5.0 105s99-master-generated-registries 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m99-master-ssh 3.2.0 40m99-worker-generated-registries 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m99-worker-ssh 3.2.0 40mrendered-master-23e785de7587df95a4b517e0647e5ab7 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33mrendered-worker-5d596d9293ca3ea80c896a1191735bb1 52dd3ba6a9a527fc3ab42afac8d12b693534c8c9 3.5.0 33m -
Check the nodes:
$ oc get nodesExample outputNAME STATUS ROLES AGE VERSIONip-10-0-136-161.ec2.internal Ready worker 28m v1.35.4ip-10-0-136-243.ec2.internal Ready master 34m v1.35.4ip-10-0-141-105.ec2.internal Ready,SchedulingDisabled worker 28m v1.35.4ip-10-0-142-249.ec2.internal Ready master 34m v1.35.4ip-10-0-153-11.ec2.internal Ready worker 28m v1.35.4ip-10-0-153-150.ec2.internal Ready master 34m v1.35.4You can see that scheduling on each worker node is disabled as the change is being applied.
-
Check that the kernel argument worked by going to one of the worker nodes and listing the kernel command-line arguments (in
/proc/cmdlineon the host):$ oc debug node/ip-10-0-141-105.ec2.internalExample outputStarting pod/ip-10-0-141-105ec2internal-debug ...To use host binaries, run `chroot /host`sh-4.2# cat /host/proc/cmdlineBOOT_IMAGE=/ostree/rhcos-... console=tty0 console=ttyS0,115200n8rootflags=defaults,prjquota rw root=UUID=fd0... ostree=/ostree/boot.0/rhcos/16...coreos.oem.id=qemu coreos.oem.id=ec2 ignition.platform.id=ec2 enforcing=0sh-4.2# exitYou should see the
enforcing=0argument added to the other kernel arguments.
About configuring parallel container image pulls
To help control bandwidth issues, you can configure the number of workload images that can be pulled at the same time.
By default, the cluster pulls images in parallel, which allows multiple workloads to pull images at the same time. Pulling multiple images in parallel can improve workload start-up time because workloads can pull needed images without waiting for each other. However, pulling too many images at the same time can use excessive network bandwidth and cause latency issues throughout your cluster.
The default setting allows unlimited simultaneous image pulls. But, you can configure the maximum number of images that can be pulled in parallel. You can also force serial image pulling, which means that only one image can be pulled at a time.
To control the number of images that can be pulled simultaneously, use a kubelet configuration to set the maxParallelImagePulls to specify a limit. Additional image pulls above this limit are held until one of the current pulls is complete.
To force serial image pulls, use a kubelet configuration to set serializeImagePulls field to true.
Configure parallel container image pulls
You can control the number of images that can be pulled by your workload simultaneously by using a kubelet configuration. You can set a maximum number of images that can be pulled or force workloads to pull images one at a time.
Prerequisites
- You have a running OpenShift Container Platform cluster.
- You are logged in to the cluster as a user with administrative privileges.
Procedure
-
Apply a custom label to the machine config pool where you want to configure parallel pulls by running a command similar to the following.
$ oc label machineconfigpool <mcp_name> parallel-pulls=set -
Create a custom resource (CR) to configure parallel image pulling.
apiVersion: machineconfiguration.openshift.io/v1kind: KubeletConfigmetadata:name: parallel-image-pulls# ...spec:machineConfigPoolSelector:matchLabels:parallel-pulls: setkubeletConfig:serializeImagePulls: falsemaxParallelImagePulls: 3# ...where:
serializeImagePulls- Specifies whether parallel pulling is enabled for the nodes in the associated machine config pool. Set to
falseto enable parallel image pulls. Set totrueto force serial image pulling. The default isfalse. maxParallelImagePulls- Specifies the maximum number of images that can be pulled in parallel. Enter a number or set to
nilto specify no limit. This field cannot be set ifSerializeImagePullsistrue. The default isnil.
-
Create the new machine config by running a command similar to the following:
$ oc create -f <file_name>.yaml
Verification
-
Check the machine configs to see that a new one was added by running the following command:
$ oc get MachineConfigNAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE00-master 70025364a114fc3067b2e82ce47fdb0149630e4b 3.5.0 133m00-worker 70025364a114fc3067b2e82ce47fdb0149630e4b 3.5.0 133m# ...99-parallel-generated-kubelet 70025364a114fc3067b2e82ce47fdb0149630e4b 3.5.0 15s# ...rendered-parallel-c634a80f644740974ceb40c054c79e50 70025364a114fc3067b2e82ce47fdb0149630e4b 3.5.0 10swhere:
99-parallel-generated-kubelet- Specifies the new machine config. In this example, the machine config is for the
parallelcustom machine config pool. rendered-parallel-<sha_numnber>- Specifies the new rendered machine config. In this example, the machine config is for the
parallelcustom machine config pool.
-
Check to see that the nodes in the
parallelmachine config pool are being updated by running the following command:$ oc get machineconfigpoolNAME CONFIG UPDATED UPDATING DEGRADED MACHINECOUNT READYMACHINECOUNT UPDATEDMACHINECOUNT DEGRADEDMACHINECOUNT AGEparallel rendered-parallel-3904f0e69130d125b3b5ef0e981b1ce1 False True False 1 0 0 0 65mmaster rendered-master-7536834c197384f3734c348c1d957c18 True False False 3 3 3 0 140mworker rendered-worker-c634a80f644740974ceb40c054c79e50 True False False 2 2 2 0 140m -
When the nodes are updated, verify that the parallel pull maximum was configured:
-
Open an
oc debugsession to a node by running a command similar to the following:$ oc debug node/<node_name> -
Set
/hostas the root directory within the debug shell by running the following command:sh-5.1# chroot /host -
Examine the
kubelet.conffile by running the following command:sh-5.1# cat /etc/kubernetes/kubelet.conf | grep -i maxParallelImagePullsmaxParallelImagePulls: 3
-
Migrating control plane nodes from one RHOSP host to another manually
If control plane machine sets are not enabled on your cluster, you can run a script that moves a control plane node from one Red Hat OpenStack Platform (RHOSP) node to another.
Control plane machine sets are not enabled on clusters that run on user-provisioned infrastructure.
For information about control plane machine sets, see "Managing control plane machines with control plane machine sets".
Prerequisites
- The environment variable
OS_CLOUDrefers to acloudsentry that has administrative credentials in aclouds.yamlfile. - The environment variable
KUBECONFIGrefers to a configuration that contains administrative OpenShift Container Platform credentials.
Procedure
-
From a command line, run the following script:
#!/usr/bin/env bashset -Eeuo pipefailif [ $# -lt 1 ]; thenecho "Usage: '$0 node_name'"exit 64fi# Check for admin OpenStack credentialsopenstack server list --all-projects >/dev/null || { >&2 echo "The script needs OpenStack admin credentials. Exiting"; exit 77; }# Check for admin OpenShift credentialsoc adm top node >/dev/null || { >&2 echo "The script needs OpenShift admin credentials. Exiting"; exit 77; }set -xdeclare -r node_name="$1"declare server_idserver_id="$(openstack server list --all-projects -f value -c ID -c Name | grep "$node_name" | cut -d' ' -f1)"readonly server_id# Drain the nodeoc adm cordon "$node_name"oc adm drain "$node_name" --delete-emptydir-data --ignore-daemonsets --force# Power off the serveroc debug "node/${node_name}" -- chroot /host shutdown -h 1# Verify the server is shut offuntil openstack server show "$server_id" -f value -c status | grep -q 'SHUTOFF'; do sleep 5; done# Migrate the nodeopenstack server migrate --wait "$server_id"# Resize the VMopenstack server resize confirm "$server_id"# Wait for the resize confirm to finishuntil openstack server show "$server_id" -f value -c status | grep -q 'SHUTOFF'; do sleep 5; done# Restart the VMopenstack server start "$server_id"# Wait for the node to show up as Ready:until oc get node "$node_name" | grep -q "^${node_name}[[:space:]]\+Ready"; do sleep 5; done# Uncordon the nodeoc adm uncordon "$node_name"# Wait for cluster operators to stabilizeuntil oc get co -o go-template='statuses: {{ range .items }}{{ range .status.conditions }}{{ if eq .type "Degraded" }}{{ if ne .status "False" }}DEGRADED{{ end }}{{ else if eq .type "Progressing"}}{{ if ne .status "False" }}PROGRESSING{{ end }}{{ else if eq .type "Available"}}{{ if ne .status "True" }}NOTAVAILABLE{{ end }}{{ end }}{{ end }}{{ end }}' | grep -qv '\(DEGRADED\|PROGRESSING\|NOTAVAILABLE\)'; do sleep 5; doneIf the script completes, the control plane machine is migrated to a new RHOSP node.
Enabling Pressure Stall Information (PSI) monitoring
You can enable Linux Pressure Stall Information (PSI) monitoring by using a MachineConfig object. Enabling PSI monitoring makes PSI metrics for CPU, memory, and I/O available for your cluster. You can use the metrics to help identify disruptions caused by CPU, memory, and I/O resource issues for better scheduling and eviction decisions.
The performance impact of enabling PSI is negligible. However, for latency-sensitive environments, you can enable PSI in a test environment to determine its impact before enabling it on production clusters. Or, you can enable the PSI on one worker node and monitor for the performance impact. Then, gradually enable the feature on more worker nodes. Enabling PSI introduces new container metrics that can increase the RSS memory usage of prometheus-k8s pods. On a cluster with a large number of containers, monitor the memory usage of the prometheus-k8s pods and adjust resource as needed.
Prerequisites
- You have obtained the label associated with the static
MachineConfigPoolCR for the type of node you want to configure.
Procedure
-
Create a YAML file similar to the following that contains the machine configuration:
apiVersion: machineconfiguration.openshift.io/v1kind: MachineConfigmetadata:labels:machineconfiguration.openshift.io/role: <label>name: 99-openshift-machineconfig-<label>-kargsspec:kernelArguments:- psi=1where:
metadata.labels.machineconfiguration.openshift.io/role:<label>- Specifies the role of the nodes where you want to enable PSI.
spec.kernelArguments.psi.1- Specifies that Pressure Stall Information (PSI) monitoring is to be enabled.
-
Create the
MachineConfigobject by running a command similar to the following:$ oc create -f <file-name>.yamlThe nodes with the specified role reboot while the changes are applied.
Verification
- After the nodes return to the
Readystate, start a debug session for a node by running the following command:$ oc debug node/<node_name> - Set
/hostas the root directory within the debug shell by running the following command:sh-5.1# chroot /host - Use any or all of the following methods to verify that PSI is enabled on the node:
-
Check that PSI is enabled by running the following command:
sh-5.1# cat /proc/cmdline | grep psi=1Example outputBOOT_IMAGE=(hd0,gpt3)/boot/ostree/rhcos-462f0c234e81f17224d7dbd4dc22d3f7046f70179e259c41a0a807e8b2f99428/vmlinuz-5.14.0-570.80.1.el9_6.x86_64 rw ostree=/ostree/boot.0/rhcos/462f0c234e81f17224d7dbd4dc22d3f7046f70179e259c41a0a807e8b2f99428/0 ignition.platform.id=gcp console=tty0 console=ttyS0,115200n8 root=UUID=d8254ef4-668e-49c5-b430-a45eefb834d7 rw rootflags=prjquota boot=UUID=22972ae8-ed08-43ff-931d-49e6d737c542 systemd.unified_cgroup_hierarchy=1 cgroup_no_v1=all psi=1where:
psi=1- Specifies that PSI is enabled on the node.
-
Check that the PSI files are present by running the following command:
sh-5.1# ls /proc/pressure/Example outputcpu io irq memorywhere:
cpu,io,irq, andmemory- Specifies the pressure files, indicating the PSI is enabled.
-
Check the PSI metrics to show the percentage of time processes were stalled waiting for CPU resources by running the following command:
# cat /proc/pressure/cpuExample outputsome avg10=0.91 avg60=1.05 avg300=1.28 total=61523749 full avg10=0.00 avg60=0.00 avg300=0.00 total=0where:
some- Specifies the share of time in which at least some tasks are stalled while waiting for CPU.
full- Specifies the share of time in which all non-idle tasks are stalled while waiting for CPU.
For more information on using PSI metrics, see "PSI - Pressure Stall Information" in the Linux Kernel documentation.
-
Additional resources