Tuning nodes for low latency with the performance profile¶
Tune nodes for low latency by using the cluster performance profile. You can restrict CPUs for infra and application containers, configure huge pages, Hyper-Threading, and configure CPU partitions for latency-sensitive processes.
Creating a performance profile¶
You can create a cluster performance profile by using the Performance Profile Creator (PPC) tool. The PPC is a function of the Node Tuning Operator.
The PPC combines information about your cluster with user-supplied configurations to generate a performance profile that is appropriate to your hardware, topology and use-case.
Note
Performance profiles are applicable only to bare-metal environments where the cluster has direct access to the underlying hardware resources. You can configure performances profiles for both single-node OpenShift and multi-node clusters.
The following is a high-level workflow for creating and applying a performance profile in your cluster:
-
Create a machine config pool (MCP) for nodes that you want to target with performance configurations. In single-node OpenShift clusters, you must use the
masterMCP because there is only one node in the cluster. -
Gather information about your cluster using the
must-gathercommand. -
Use the PPC tool to create a performance profile by using either of the following methods:
- Run the PPC tool by using Podman as described in Running the Performance Profile Creator using Podman. .
- Run the PPC tool by using a wrapper script as described in Running the Performance Profile Creator wrapper script..
-
Configure the performance profile for your use case and apply the performance profile to your cluster.
About the Performance Profile Creator¶
The Performance Profile Creator (PPC) is a command-line tool and is delivered with the Node Tuning Operator. You can use the PPC CLI to create a performance profile for your cluster.
Initially, you can use the PPC tool to process the must-gather data to display key performance configurations for your cluster, including the following information:
- NUMA cell partitioning with the allocated CPU IDs
- Hyper-Threading node configuration
You can use this information to help you configure the performance profile.
Specify performance configuration arguments to the PPC tool to generate a proposed performance profile that is appropriate for your hardware, topology, and use-case.
You can run the PPC by using one of the following methods:
- Run the PPC by using Podman
- Run the PPC by using the wrapper script
Note
Using the wrapper script abstracts some of the more granular Podman tasks into an executable script. For example, the wrapper script handles tasks such as pulling and running the required container image, mounting directories into the container, and providing parameters directly to the container through Podman. Both methods achieve the same result.
Create a machine config pool to target nodes for performance tuning¶
For multi-node clusters, you can define a machine config pool (MCP) to identify the target nodes that you want to configure with a performance profile.
In single-node OpenShift clusters, you must use the master MCP because there is only one node in the cluster. You do not need to create a separate MCP for single-node OpenShift clusters.
Prerequisites
- You have
cluster-adminrole access. - You installed the OpenShift CLI (
oc).
Procedure
-
Label the target nodes for configuration by running the following command:
<node_name>: Specifies the name of your node. This example applies theworker-cnflabel.
-
Create a
MachineConfigPoolresource containing the target nodes:-
Create a YAML file that defines the
MachineConfigPoolresource:Example mcp-worker-cnf.yaml fileapiVersion: machineconfiguration.openshift.io/v1 kind: MachineConfigPool metadata: name: worker-cnf labels: machineconfiguration.openshift.io/role: worker-cnf spec: machineConfigSelector: matchExpressions: - { key: machineconfiguration.openshift.io/role, operator: In, values: [worker, worker-cnf], } paused: false nodeSelector: matchLabels: node-role.kubernetes.io/worker-cnf: ""where:
metadata.name- Specifies a name for the
MachineConfigPoolresource. machineconfiguration.openshift.io/role- Specifes a unique label for the machine config pool.
node-role.kubernetes.io/worker-cnf- Specifies the nodes with the target label that you defined.
-
Apply the
MachineConfigPoolresource by running the following command:
-
Verification
-
Check the machine config pools in your cluster by running the following command:
Example outputNAME CONFIG UPDATED UPDATING DEGRADED MACHINECOUNT READYMACHINECOUNT UPDATEDMACHINECOUNT DEGRADEDMACHINECOUNT AGE master rendered-master-58433c7c3c1b4ed5ffef95234d451490 True False False 3 3 3 0 6h46m worker rendered-worker-168f52b168f151e4f853259729b6azc4 True False False 2 2 2 0 6h46m worker-cnf rendered-worker-cnf-168f52b168f151e4f853259729b6azc4 True False False 1 1 1 0 73s
Gather data about your cluster for the PPC¶
The Performance Profile Creator (PPC) tool requires must-gather data. As a cluster administrator, run the must-gather command to capture information about your cluster.
Prerequisites
- Access to the cluster as a user with the
cluster-adminrole. - You installed the OpenShift CLI (
oc). - You identified a target MCP that you want to configure with a performance profile.
Procedure
-
Navigate to the directory where you want to store the
must-gatherdata. -
Collect cluster information by running the following command:
The command creates a folder with the
must-gatherdata in your local directory with a naming format similar to the following:must-gather.local.1971646453781853027. -
Optional: Create a compressed file from the
must-gatherdirectory:-
<must_gather_folder>: Specifies the name of themust-gatherdata folder.Note
Compressed output is required if you are running the Performance Profile Creator wrapper script.
-
Additional resources
Run the Performance Profile Creator using Podman¶
As a cluster administrator, you can use Podman with the Performance Profile Creator (PPC) to create a performance profile.
For more information about the PPC arguments, see the section "Performance Profile Creator arguments".
Warning
The PPC uses the must-gather data from your cluster to create the performance profile. If you make any changes to your cluster, such as relabeling a node targeted for performance configuration, you must re-create the must-gather data before running PPC again.
Prerequisites
- Access to the cluster as a user with the
cluster-adminrole. - A cluster installed on bare-metal hardware.
- You installed
podmanand the OpenShift CLI (oc). - Access to the Node Tuning Operator image.
- You identified a machine config pool containing target nodes for configuration.
- You have access to the
must-gatherdata for your cluster.
Procedure
-
Check the machine config pool by running the following command:
The following output lists the available machine config pools:
NAME CONFIG UPDATED UPDATING DEGRADED MACHINECOUNT READYMACHINECOUNT UPDATEDMACHINECOUNT DEGRADEDMACHINECOUNT AGE master rendered-master-58433c8c3c0b4ed5feef95434d455490 True False False 3 3 3 0 8h worker rendered-worker-668f56a164f151e4a853229729b6adc4 True False False 2 2 2 0 8h worker-cnf rendered-worker-cnf-668f56a164f151e4a853229729b6adc4 True False False 1 1 1 0 79m -
Use Podman to authenticate to
registry.redhat.ioby running the following command: -
Optional: Display help for the PPC tool by running the following command:
$ podman run --rm --entrypoint performance-profile-creator registry.redhat.io/openshift4/ose-cluster-node-tuning-rhel9-operator:v4.22 -hThe following output shows the available flags and commands for the PPC tool:
A tool that automates creation of Performance Profiles Available Commands: completion Generate the autocompletion script for the specified shell help Help about any command info requires --must-gather-dir-path, ignores other arguments. [Valid values: log,json] Usage: performance-profile-creator [flags] performance-profile-creator [command] Flags: --disable-ht Disable Hyperthreading --enable-hardware-tuning Enable setting maximum cpu frequencies -h, --help help for performance-profile-creator --mcp-name string MCP name corresponding to the target machines (required) --must-gather-dir-path string Must gather directory path (default "must-gather") --offlined-cpu-count int Number of offlined CPUs --per-pod-power-management Enable Per Pod Power Management --power-consumption-mode string The power consumption mode. [Valid values: default, low-latency, ultra-low-latency] (default "default") --profile-name string Name of the performance profile to be created (default "performance") --reserved-cpu-count int Number of reserved CPUs (required) --rt-kernel Enable Real Time Kernel (required) --split-reserved-cpus-across-numa Split the Reserved CPUs across NUMA nodes --topology-manager-policy string Kubelet Topology Manager Policy of the performance profile to be created. [Valid values: single-numa-node, best-effort, restricted] (default "restricted") --user-level-networking Run with User level Networking(DPDK) enabled Use "performance-profile-creator [command] --help" for more information about a command. -
To display information about the cluster, run the PPC tool with the
infocommand by running the following command:$ podman run --entrypoint performance-profile-creator -v <path_to_must_gather>:/must-gather:z registry.redhat.io/openshift4/ose-cluster-node-tuning-rhel9-operator:v4.22 info --must-gather-dir-path /must-gather-
--entrypoint performance-profile-creatordefines the performance profile creator as a new entry point topodman. -
-v <path_to_must_gather>specifies the path to either of the following components:-
The directory containing the
must-gatherdata. -
An existing directory containing the
must-gatherdecompressed .tar file.The following output is generated from running the command:
PPC: Nodes names targeted by master pool are: PPC: Nodes names targeted by worker-cnf pool are: host2.example.com PPC: Nodes names targeted by worker pool are: host.example.com host1.example.com PPC: Cluster info: MCP 'master' nodes: --- MCP 'worker' nodes: Node: host.example.com (NUMA cells: 1, HT: true) NUMA cell 0 : [0 1 2 3] CPU(s): 4 Node: host1.example.com (NUMA cells: 1, HT: true) NUMA cell 0 : [0 1 2 3] CPU(s): 4 --- MCP 'worker-cnf' nodes: Node: host2.example.com (NUMA cells: 1, HT: true) NUMA cell 0 : [0 1 2 3] CPU(s): 4 ---
-
-
-
Create a performance profile by running the following command. The example uses sample PPC arguments and values:
$ podman run --entrypoint performance-profile-creator -v <path_to_must_gather>:/must-gather:z registry.redhat.io/openshift4/ose-cluster-node-tuning-rhel9-operator:v4.22 --mcp-name=worker-cnf --reserved-cpu-count=2 --rt-kernel=true --split-reserved-cpus-across-numa=false --must-gather-dir-path /must-gather --power-consumption-mode=ultra-low-latency > my-performance-profile.yaml-
-v <path_to_must_gather>specifies the path to either of the following components:- The directory containing the
must-gatherdata. - The directory containing the
must-gatherdecompressed .tar file.
- The directory containing the
-
--mcp-name=worker-cnfspecifies theworker-cnfmachine config pool. -
--reserved-cpu-count=2specifies two reserved CPUs. -
--rt-kernel=trueenables the real-time kernel. -
--split-reserved-cpus-across-numa=falsedisables reserved CPUs splitting across NUMA nodes. -
--power-consumption-mode=ultra-low-latencyspecifies minimal latency at the cost of increased power consumption.Note
The
mcp-nameargument in this example is set toworker-cnfbased on the output of the commandoc get mcp. For single-node OpenShift use--mcp-name=master.The following output shows the CPU and NUMA allocation for the performance profile:
-
-
Review the created YAML file by running the following command:
The following example shows the contents of the generated performance profile:
--- apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: performance spec: cpu: isolated: 2-3 reserved: "0-1" machineConfigPoolSelector: machineconfiguration.openshift.io/role: worker-cnf net: userLevelNetworking: false nodeSelector: node-role.kubernetes.io/worker-cnf: "" numa: topologyPolicy: restricted realTimeKernel: enabled: true workloadHints: highPowerConsumption: true perPodPowerManagement: false realTime: true -
Apply the generated profile:
The following output confirms that the performance profile is created:
Run the Performance Profile Creator wrapper script¶
The wrapper script simplifies the process of creating a performance profile with the Performance Profile Creator (PPC) tool.
The script handles tasks such as pulling and running the required container image, mounting directories into the container, and providing parameters directly to the container through Podman.
For more information about the Performance Profile Creator arguments, see the section "Performance Profile Creator arguments".
Warning
The PPC uses the must-gather data from your cluster to create the performance profile. If you make any changes to your cluster, such as relabeling a node targeted for performance configuration, you must re-create the must-gather data before running PPC again.
Prerequisites
- Access to the cluster as a user with the
cluster-adminrole. - A cluster installed on bare-metal hardware.
- You installed
podmanand the OpenShift CLI (oc). - Access to the Node Tuning Operator image.
- You identified a machine config pool containing target nodes for configuration.
- Access to the
must-gathertarball.
Procedure
-
Create a file on your local machine named, for example,
run-perf-profile-creator.sh: -
Paste the following code into the file:
#!/bin/bash readonly CONTAINER_RUNTIME=${CONTAINER_RUNTIME:-podman} readonly CURRENT_SCRIPT=$(basename "$0") readonly CMD="${CONTAINER_RUNTIME} run --entrypoint performance-profile-creator" readonly IMG_EXISTS_CMD="${CONTAINER_RUNTIME} image exists" readonly IMG_PULL_CMD="${CONTAINER_RUNTIME} image pull" readonly MUST_GATHER_VOL="/must-gather" NTO_IMG="registry.redhat.io/openshift4/ose-cluster-node-tuning-rhel9-operator:v4.22" MG_TARBALL="" DATA_DIR="" usage() { print "Wrapper usage:" print " ${CURRENT_SCRIPT} [-h] [-p image][-t path] -- [performance-profile-creator flags]" print "" print "Options:" print " -h help for ${CURRENT_SCRIPT}" print " -p Node Tuning Operator image" print " -t path to a must-gather tarball" ${IMG_EXISTS_CMD} "${NTO_IMG}" && ${CMD} "${NTO_IMG}" -h } function cleanup { [ -d "${DATA_DIR}" ] && rm -rf "${DATA_DIR}" } trap cleanup EXIT exit_error() { print "error: $*" usage exit 1 } print() { echo "$*" >&2 } check_requirements() { ${IMG_EXISTS_CMD} "${NTO_IMG}" || ${IMG_PULL_CMD} "${NTO_IMG}" || \ exit_error "Node Tuning Operator image not found" [ -n "${MG_TARBALL}" ] || exit_error "Must-gather tarball file path is mandatory" [ -f "${MG_TARBALL}" ] || exit_error "Must-gather tarball file not found" DATA_DIR=$(mktemp -d -t "${CURRENT_SCRIPT}XXXX") || exit_error "Cannot create the data directory" tar -zxf "${MG_TARBALL}" --directory "${DATA_DIR}" || exit_error "Cannot decompress the must-gather tarball" chmod a+rx "${DATA_DIR}" return 0 } main() { while getopts ':hp:t:' OPT; do case "${OPT}" in h) usage exit 0 ;; p) NTO_IMG="${OPTARG}" ;; t) MG_TARBALL="${OPTARG}" ;; ?) exit_error "invalid argument: ${OPTARG}" ;; esac done shift $((OPTIND - 1)) check_requirements || exit 1 ${CMD} -v "${DATA_DIR}:${MUST_GATHER_VOL}:z" "${NTO_IMG}" "$@" --must-gather-dir-path "${MUST_GATHER_VOL}" echo "" 1>&2 } main "$@" -
Add execute permissions for everyone on this script:
-
Use Podman to authenticate to
registry.redhat.ioby running the following command: -
Optional: Display help for the PPC tool by running the following command:
Wrapper usage: run-perf-profile-creator.sh [-h] [-p image][-t path] -- [performance-profile-creator flags] Options: -h help for run-perf-profile-creator.sh -p Node Tuning Operator image -t path to a must-gather tarball A tool that automates creation of Performance Profiles Available Commands: completion Generate the autocompletion script for the specified shell help Help about any command info requires --must-gather-dir-path, ignores other arguments. [Valid values: log,json] Usage: performance-profile-creator [flags] performance-profile-creator [command] Flags: --disable-ht Disable Hyperthreading --enable-hardware-tuning Enable setting maximum cpu frequencies -h, --help help for performance-profile-creator --mcp-name string MCP name corresponding to the target machines (required) --must-gather-dir-path string Must gather directory path (default "must-gather") --offlined-cpu-count int Number of offlined CPUs --per-pod-power-management Enable Per Pod Power Management --power-consumption-mode string The power consumption mode. [Valid values: default, low-latency, ultra-low-latency] (default "default") --profile-name string Name of the performance profile to be created (default "performance") --reserved-cpu-count int Number of reserved CPUs (required) --rt-kernel Enable Real Time Kernel (required) --split-reserved-cpus-across-numa Split the Reserved CPUs across NUMA nodes --topology-manager-policy string Kubelet Topology Manager Policy of the performance profile to be created. [Valid values: single-numa-node, best-effort, restricted] (default "restricted") --user-level-networking Run with User level Networking(DPDK) enabled Use "performance-profile-creator [command] --help" for more information about a command.Note
You can optionally set a path for the Node Tuning Operator image using the
-poption. If you do not set a path, the wrapper script uses the default image:registry.redhat.io/openshift4/ose-cluster-node-tuning-rhel9-operator:v4.22. -
To display information about the cluster, run the PPC tool with the
infocommand by running the following command:-
-t /<path_to_must_gather_dir>/must-gather.tar.gz: Specifies the path to directory containing the must-gather tarball. This is a required argument for the wrapper script.The following example shows the cluster information output:
PPC: Cluster info: MCP 'master' nodes: --- MCP 'worker' nodes: Node: host.example.com (NUMA cells: 1, HT: true) NUMA cell 0 : [0 1 2 3] CPU(s): 4 Node: host1.example.com (NUMA cells: 1, HT: true) NUMA cell 0 : [0 1 2 3] CPU(s): 4 --- MCP 'worker-cnf' nodes: Node: host2.example.com (NUMA cells: 1, HT: true) NUMA cell 0 : [0 1 2 3] CPU(s): 4 ---
-
-
Create a performance profile by running the following command. The example command uses sample PPC arguments and values.
$ ./run-perf-profile-creator.sh -t /path-to-must-gather/must-gather.tar.gz -- --mcp-name=worker-cnf --reserved-cpu-count=2 --rt-kernel=true --split-reserved-cpus-across-numa=false --power-consumption-mode=ultra-low-latency > my-performance-profile.yaml-
--mcp-name=worker-cnfspecifies theworker-cnfmachine config pool. -
--reserved-cpu-count=2specifies two reserved CPUs. -
--rt-kernel=trueenables the real-time kernel. -
--split-reserved-cpus-across-numa=falsedisables reserved CPUs splitting across NUMA nodes. -
--power-consumption-mode=ultra-low-latencyspecifies minimal latency at the cost of increased power consumption.Note
The
mcp-nameargument in this example is set toworker-cnfbased on the output of the commandoc get mcp. For single-node OpenShift use--mcp-name=master.
-
-
Review the created YAML file by running the following command:
The generated YAML file should contain the following output:
apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: performance spec: cpu: isolated: 2-3 reserved: "0-1" machineConfigPoolSelector: machineconfiguration.openshift.io/role: worker-cnf nodeSelector: node-role.kubernetes.io/worker-cnf: "" numa: topologyPolicy: restricted realTimeKernel: enabled: true workloadHints: highPowerConsumption: true perPodPowerManagement: false realTime: true -
Apply the generated profile:
The performance profile is successfully applied and created as shown in the following output:
Performance Profile Creator arguments¶
To customize the generation of performance profiles, review the arguments for the Performance Profile Creator.
Required Performance Profile Creator arguments
| Argument | Description |
|---|---|
mcp-name |
Name for MCP; for example, worker-cnf corresponding to the target machines. |
must-gather-dir-path |
The path of the must gather directory. This argument is only required if you run the PPC tool by using Podman. If you use the PPC with the wrapper script, do not use this argument. Instead, specify the directory path to the must-gather tarball by using the -t option for the wrapper script. |
reserved-cpu-count |
Number of reserved CPUs. Use a natural number greater than zero. |
rt-kernel |
Enables real-time kernel. Possible values: true or false. |
Optional Performance Profile Creator arguments
| Argument | Description |
|---|---|
disable-ht |
Disable Hyper-Threading. Possible values: true or false.Default: false.Warning If this argument is set to |
| enable-hardware-tuning | Enable the setting of maximum CPU frequencies. To enable this feature, set the maximum frequency for applications running on isolated and reserved CPUs for both of the following fields:
PerformanceProfile includes warnings and guidance on how to set frequency settings. |
info |
This captures cluster information. This argument also requires the must-gather-dir-path argument. If any other arguments are set they are ignored.Possible values:
log. |
offlined-cpu-count |
Number of offlined CPUs. Note Use a natural number greater than zero. If not enough logical processors are offlined, then error messages are logged. The messages are: Error: failed to compute the reserved and isolated CPUs: please ensure that reserved-cpu-count plus offlined-cpu-count should be in the range [0,1] Error: failed to compute the reserved and isolated CPUs: please specify the offlined CPU count in the range [0,1] |
power-consumption-mode |
The power consumption mode. Possible values:
default. |
per-pod-power-management |
Enable per pod power management. You cannot use this argument if you configured ultra-low-latency as the power consumption mode.Possible values: true or false.Default: false. |
profile-name |
Name of the performance profile to create. Default: performance. |
split-reserved-cpus-across-numa |
Split the reserved CPUs across NUMA nodes. Possible values: true or false.Default: false. |
topology-manager-policy |
Kubelet Topology Manager policy of the performance profile to be created. Possible values:
restricted. |
user-level-networking |
Run with user level networking (DPDK) enabled. Possible values: true or false.Default: false. |
Reference performance profiles¶
Use the following reference performance profiles as the basis to develop your own custom profiles.
Performance profile template for clusters that use OVS-DPDK on OpenStack¶
To maximize machine performance in a cluster that uses Open vSwitch with the Data Plane Development Kit (OVS-DPDK) on Red Hat OpenStack Platform (RHOSP), you can use a performance profile.
You can use the following performance profile template to create a profile for your deployment.
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
name: cnf-performanceprofile
spec:
additionalKernelArgs:
- nmi_watchdog=0
- audit=0
- mce=off
- processor.max_cstate=1
- idle=poll
- intel_idle.max_cstate=0
- default_hugepagesz=1GB
- hugepagesz=1G
- intel_iommu=on
cpu:
isolated: <CPU_ISOLATED>
reserved: <CPU_RESERVED>
hugepages:
defaultHugepagesSize: 1G
pages:
- count: <HUGEPAGES_COUNT>
node: 0
size: 1G
nodeSelector:
node-role.kubernetes.io/worker: ''
realTimeKernel:
enabled: false
globallyDisableIrqLoadBalancing: true
Insert values that are appropriate for your configuration for the CPU_ISOLATED, CPU_RESERVED, and HUGEPAGES_COUNT keys.
Telco RAN DU reference design performance profile¶
You can use a pre-configured design performance profile that configures node-level performance settings for OpenShift Container Platform clusters on commodity hardware to host telco RAN DU workloads.
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
# if you change this name make sure the 'include' line in TunedPerformancePatch.yaml
# matches this name: include=openshift-node-performance-${PerformanceProfile.metadata.name}
# Also in file 'validatorCRs/informDuValidator.yaml':
# name: 50-performance-${PerformanceProfile.metadata.name}
name: openshift-node-performance-profile
annotations:
ran.openshift.io/reference-configuration: "ran-du.redhat.com"
spec:
additionalKernelArgs:
- "rcupdate.rcu_normal_after_boot=0"
- "efi=runtime"
- "vfio_pci.enable_sriov=1"
- "vfio_pci.disable_idle_d3=1"
- "module_blacklist=irdma"
cpu:
isolated: $isolated
reserved: $reserved
hugepages:
defaultHugepagesSize: $defaultHugepagesSize
pages:
- size: $size
count: $count
node: $node
machineConfigPoolSelector:
pools.operator.machineconfiguration.openshift.io/$mcp: ""
nodeSelector:
node-role.kubernetes.io/$mcp: ''
numa:
topologyPolicy: "restricted"
# To use the standard (non-realtime) kernel, set enabled to false
realTimeKernel:
enabled: true
workloadHints:
# WorkloadHints defines the set of upper level flags for different type of workloads.
# See https://github.com/openshift/cluster-node-tuning-operator/blob/master/docs/performanceprofile/performance_profile.md#workloadhints
# for detailed descriptions of each item.
# The configuration below is set for a low latency, performance mode.
realTime: true
highPowerConsumption: false
perPodPowerManagement: false
Telco core reference design performance profile¶
You can use a pre-configured design performance profile that configures node-level performance settings for OpenShift Container Platform clusters on commodity hardware to host telco core workloads.
# required
# count: 1
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
name: $name
annotations:
# Some pods want the kernel stack to ignore IPv6 router Advertisement.
kubeletconfig.experimental: |
{"allowedUnsafeSysctls":["net.ipv6.conf.all.accept_ra"]}
spec:
cpu:
# node0 CPUs: 0-17,36-53
# node1 CPUs: 18-34,54-71
# siblings: (0,36), (1,37)...
# we want to reserve the first Core of each NUMA socket
#
# no CPU left behind! all-cpus == isolated + reserved
isolated: $isolated # eg 1-17,19-35,37-53,55-71
reserved: $reserved # eg 0,18,36,54
# Guaranteed QoS pods will disable IRQ balancing for cores allocated to the pod.
# default value of globallyDisableIrqLoadBalancing is false
globallyDisableIrqLoadBalancing: false
hugepages:
defaultHugepagesSize: 1G
pages:
# 32GB per numa node
- count: $count # eg 64
size: 1G
#machineConfigPoolSelector: {}
# pools.operator.machineconfiguration.openshift.io/worker: ''
nodeSelector: {}
#node-role.kubernetes.io/worker: ""
workloadHints:
realTime: false
highPowerConsumption: false
perPodPowerManagement: true
realTimeKernel:
enabled: false
numa:
# All guaranteed QoS containers get resources from a single NUMA node
topologyPolicy: "single-numa-node"
net:
userLevelNetworking: false
Supported performance profile API versions¶
The Node Tuning Operator supports v2, v1, and v1alpha1 for the performance profile apiVersion field. The v1 and v1alpha1 APIs are identical. The v2 API includes an optional boolean field globallyDisableIrqLoadBalancing with a default value of false.
- Upgrading the performance profile to use device interrupt processing
-
When you upgrade the Node Tuning Operator performance profile custom resource definition (CRD) from v1 or v1alpha1 to v2,
globallyDisableIrqLoadBalancingis set totrueon existing profiles.Note
globallyDisableIrqLoadBalancingtoggles whether IRQ load balancing will be disabled for the Isolated CPU set. When the option is set totrueit disables IRQ load balancing for the Isolated CPU set. Setting the option tofalseallows the IRQs to be balanced across all CPUs. - Upgrading Node Tuning Operator API from v1alpha1 to v1
- When upgrading Node Tuning Operator API version from v1alpha1 to v1, the v1alpha1 performance profiles are converted on-the-fly using a "None" Conversion strategy and served to the Node Tuning Operator with API version v1.
- Upgrading Node Tuning Operator API from v1alpha1 or v1 to v2
- When upgrading from an older Node Tuning Operator API version, the existing v1 and v1alpha1 performance profiles are converted using a conversion webhook that injects the
globallyDisableIrqLoadBalancingfield with a value oftrue.
Node power consumption and realtime processing with workload hints¶
You can create a performance profile appropriate for the hardware and topology of an environment by using the Performance Profile Creator (PPC) tool.
The following table describes the possible values set for the power-consumption-mode flag associated with the PPC tool and the workload hint that is applied.
Impact of combinations of power consumption and real-time settings on latency
| Performance Profile creator setting | Hint | Environment | Description |
|---|---|---|---|
| Default | workloadHints: highPowerConsumption: false realTime: false |
High throughput cluster without latency requirements | Performance achieved through CPU partitioning only. |
| Low-latency | workloadHints: highPowerConsumption: false realTime: true |
Regional data-centers | Both energy savings and low-latency are desirable: compromise between power management, latency and throughput. |
| Ultra-low-latency | workloadHints: highPowerConsumption: true realTime: true |
Far edge clusters, latency critical workloads | Optimized for absolute minimal latency and maximum determinism at the cost of increased power consumption. |
| Per-pod power management | workloadHints: realTime: true highPowerConsumption: false perPodPowerManagement: true |
Critical and non-critical workloads | Allows for power management per pod. |
The following configuration is commonly used in a telco RAN DU deployment:
apiVersion: performance.openshift.io/v2
kind: PerformanceProfile
metadata:
name: workload-hints
spec:
...
workloadHints:
realTime: true
highPowerConsumption: false
perPodPowerManagement: false
perPodPowerManagement- Specifies to disable some debugging and monitoring features that can affect system latency.
Note
When the realTime workload hint flag is set to true in a performance profile, add the cpu-quota.crio.io: disable annotation to every guaranteed pod with pinned CPUs. This annotation is necessary to prevent the degradation of the process performance within the pod. If the realTime workload hint is not explicitly set, it defaults to true.
For more information how combinations of power consumption and real-time settings impact latency, see "Understanding workload hints".
Additional resources
Configure power saving for nodes that run colocated high and low priority workloads¶
You can enable power savings for a node that has low priority workloads that are colocated with high priority workloads without impacting the latency or throughput of the high priority workloads. Power saving is possible without modifications to the workloads themselves.
Warning
The feature is supported on Intel Ice Lake and later generations of Intel CPUs. The capabilities of the processor might impact the latency and throughput of the high priority workloads.
Prerequisites
- You enabled C-states and operating system controlled P-states in the BIOS
Procedure
-
Generate a
PerformanceProfilewith theper-pod-power-managementargument set totrue:$ podman run --entrypoint performance-profile-creator -v \ /must-gather:/must-gather:z registry.redhat.io/openshift4/ose-cluster-node-tuning-rhel9-operator:v4.22 \ --mcp-name=worker-cnf --reserved-cpu-count=20 --rt-kernel=true \ --split-reserved-cpus-across-numa=false --topology-manager-policy=single-numa-node \ --must-gather-dir-path /must-gather --power-consumption-mode=low-latency \ --per-pod-power-management=true > my-performance-profile.yamlThe
power-consumption-modeargument must bedefaultorlow-latencywhen theper-pod-power-managementargument is set totrue. -
Set the default
cpufreqgovernor as an additional kernel argument in thePerformanceProfilecustom resource (CR):apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: performance spec: ... additionalKernelArgs: - cpufreq.default_governor=schedutil # ...where:
cpufreq.default_governor=schedutil- Specifies using the
schedutilgovernor. You can use other governors, such as theondemandorpowersavegovernors.
-
Set the maximum CPU frequency in the
TunedPerformancePatchCR:where:
/sys/devices/system/cpu/intel_pstate/max_perf_pct- Specifies the
max_perf_pctthat controls the maximum frequency that thecpufreqdriver is allowed to set as a percentage of the maximum supported cpu frequency. This value applies to all CPUs. You can check the maximum supported frequency in/sys/devices/system/cpu/cpu0/cpufreq/cpuinfo_max_freq. As a starting point, you can use a percentage that caps all CPUs at theAll Cores Turbofrequency. TheAll Cores Turbofrequency is the frequency that all cores will run at when the cores are all fully occupied.
Protecting low latency workloads from exec operations¶
How ExecCPUAffinity prevents latency spikes from exec operations¶
When you run exec operations such as oc exec or shell access on a container with isolated CPUs, those processes can interrupt your time-sensitive workloads. The ExecCPUAffinity feature automatically pins these secondary processes to a specific CPU within the container’s isolated set.
This ensures that your primary low-latency applications, such as Telco RAN DU or 5G Core, maintain deterministic performance without resource contention.
ExecCPUAffinity is enabled by default whenever you apply a PerformanceProfile to a node. The feature operates at the container level and requires the following conditions:
-
Runtime Class: The pod must use the
PerformanceProfileruntime class, for example<PP-name>-performance. -
QoS Class: The pod must belong to the Guaranteed QoS class and request whole integer CPUs.
-
CPU Selection Logic: The system automatically selects the first available CPU to host the executed process. It prioritizes a shared CPU if one is configured; otherwise, it uses the first exclusive CPU in the container’s set.
Note
If a Pod contains multiple containers, only the container requesting an integer number of CPUs uses
ExecCPUAffinity. Any container with fractional CPU requests will follow the default behavior, allowing processes to run on any CPU within the container’s cgroup set.
If you need the previous behavior where executed processes can run on any allocated core, you can disable the feature for all workloads by using that performance runtime class.
To disable ExecCPUAffinity, add the following annotation to your PerformanceProfile:
Note
Use this annotation only as a temporary fallback and is expected to be removed in future releases.
Isolate exec processes from latency-sensitive workloads¶
You can prevent oc exec and shell processes from interrupting latency-sensitive workloads by applying a PerformanceProfile to a node.
The Node Tuning Operator (NTO) automatically enables the ExecCPUAffinity feature, which pins exec processes to a designated CPU so that your primary workload CPUs remain undisturbed.
Prerequisites
- You have access to an OpenShift Container Platform cluster using an account with
cluster-adminpermissions.
Procedure
-
Apply a
PerformanceProfiletailored to your workload tuning requirements, such as reserved CPU counts or real-time kernel settings. For detailed instructions on generating a profile, see Creating a performance profile. -
Create a namespace for testing the performance configuration by running the following command:
Note
You can create the workload in any namespace. This
performance-profile-testingis used for testing purposes in this example. -
Deploy a
GuaranteedQoS pod that requests whole integer CPUs and uses the generatedRuntimeClass.-
Retrieve the
RuntimeClasscreated by the PerformanceProfile by running the following command:The following example shows the expected output:
-
Create a YAML file for the pod, for example
my-pod.yaml. Ensure it requests whole integer CPUs and references theruntimeClassNamegenerated by yourPerformanceProfile:apiVersion: v1 kind: Pod metadata: name: test namespace: performance-profile-testing spec: runtimeClassName: performance-performance containers: - name: test image: "quay.io/openshift-kni/cnf-tests:4.22" command: ["sleep", "10h"] resources: requests: cpu: "2" memory: "256Mi" limits: cpu: "2" memory: "256Mi"Note
The following requirements apply for CPU pinning:
-
Guaranteed QoS Class: To trigger CPU pinning, the Pod must belong to the Guaranteed Quality of Service class. This requires that every container in the pod has both CPU and memory limits defined, and those limits must exactly equal their corresponding requests.
-
Integer CPUs: The CPU request must be a whole integer.
- SMT/Hyperthreading: On systems where Simultaneous Multithreading (SMT) is enabled, the request should typically be a multiple of the threads per core, which is usually 2, to ensure exclusive core allocation.
- Fractional Requests: Pods with fractional CPU requests do not trigger the
ExecCPUAffinitylogic and follow the previous scheduling behavior.
-
Capacity: Ensure your cluster has sufficient isolated CPU capacity available in the node’s allocated pool to satisfy the request.
-
-
Apply the pod definition by running the following command:
-
-
Verify that the pod is running by running the following command:
Wait until the pod status shows
Runningbefore proceeding to the verification steps.
Verification
-
Verify
execprocess pinning as follows:-
Run the following command to see the specific cores exclusively assigned to this container by the CPU Manager:
Note
In a single-container pod, the command defaults to that container. In a multi-container pod, you must specify the container name using the
-c <container_name>flag to ensure you are checking the correct context.The following example shows the expected output:
The output shows exactly 2 CPUs, such as indexes 4 and 5, because the container requested
cpu: "2". The output represents the baseline exclusive CPUs for comparison. -
Start an exec process in the pod by running the following command:
-
-
Identify the node where the test pod is running:
-
Start a debug session on the node and set the root directory:
-
Find the PID of the sleep 3600 process by running the following command:
The following example shows the expected output:
-
Identify the dedicated CPU set for the container by running the following command:
Note
In a multi-container pod, you must specify the container name using the
-c <container_name>flag.The following example shows the expected output for example the specific cores exclusively assigned to this container by the CPU Manager:
-
Check the CPU affinity of the exec process by running the following command, replacing
<PID>with the PID identified in the previous step:Compare the taskset output against the container’s dedicated CPU set from Step 1.
In this example, while the container has access to CPUs 4-5, the
ExecCPUAffinitylogic has pinned the exec process specifically to CPU 4. This confirms the pinning logic is active and that your primary workload’s exclusive CPUs remain undisturbed.
-
Disable CPU isolation for executed processes¶
If you have a high-performance workload that requires executed processes to use any available core rather than being pinned to the first core, you can opt out of the default behavior by following this procedure.
Note
Adding or removing the performance.openshift.io/exec-cpu-affinity annotation triggers a MachineConfig rollout that reboots the affected nodes.
Prerequisites
- You have installed the OpenShift CLI (
oc). - You have logged in as a user with
cluster-adminprivileges.
Procedure
-
List existing profiles by running the following command:
Identify the PerformanceProfile applied to the nodes running your high-performance workload.
-
Edit the identified PerformanceProfile to add the annotation that disables CPU isolation for executed processes by running the following command, replacing
<profile-name>with the name of your PerformanceProfile: -
In the editor that opens, add the following annotation under the
metadatasection: -
Save and close the editor to apply the changes.
-
Wait for the MachineConfigPool (MCP) to finish updating:
All pools should show
UPDATED=TrueandUPDATING=Falsebefore proceeding.
Verification
-
Verify that the annotation has been added by running the following command, replacing
<profile-name>with the name of your PerformanceProfile:Expected output:
Troubleshoot ExecCPUAffinity configuration¶
If a process initiated by using oc exec is not being pinned correctly despite the pod meeting the Guaranteed QoS and integer CPU requirements, use the following procedure to verify the configuration at the node level.
Prerequisites
- You have access to the cluster as a user with
cluster-adminpermissions. - You have the OpenShift CLI (
oc) installed. - You have identified the node where the pod is running.
Procedure
-
Start a debug session for the targeted node:
-
Set
/hostas the root directory for the debug shell: -
Inspect the performance runtime configuration file:
The following example shows the expected output:
Note
If
exec_cpu_affinity = "first"is missing, ensure thePerformanceProfiledoes not contain theperformance.openshift.io/exec-cpu-affinity: "disable"annotation. If you recently changed the annotation, verify that the MachineConfigPool (MCP) has finished updating. -
Verify the live CRI-O configuration to ensure the setting is loaded into memory:
The following example shows the expected output:
Verification
-
If both the runtime configuration file and live CRI-O configuration show
exec_cpu_affinity = "first", the ExecCPUAffinity feature is correctly configured. Processes initiated byoc execon Guaranteed QoS pods with integer CPU requests are pinned to the first available core. -
If
exec_cpu_affinity = "first"is missing from either output, check that thePerformanceProfiledoes not contain theperformance.openshift.io/exec-cpu-affinity: "disable"annotation and verify that the MachineConfigPool (MCP) has finished updating by running:All pools should show
UPDATED=TrueandUPDATING=False.
Troubleshooting ExecCPUAffinity configuration common issues and resolutions¶
Troubleshoot ExecCPUAffinity configuration issues by identifying common symptoms and their resolutions. Use this information to ensure processes are correctly pinned to the intended CPUs and that the MachineConfigPool updates successfully.
The following table describes common issues and resolutions for ExecCPUAffinity configuration:
| Symptom | Potential Cause | Resolution |
|---|---|---|
| Process runs on all CPUs in the container set. | The pod does not belong to the Guaranteed QoS class. For more information see Pod Quality of Service Classes. | Ensure the container has limits and requests defined for both CPU and memory, and that they are equal. |
| Process runs on all CPUs in the container set. | The container uses fractional CPU requests for example 500m. |
Update the pod specification to request a whole integer number of CPUs. |
The 99-runtimes.conf file does not exist or is not updated. |
The MachineConfigPool (MCP) is still updating or has failed. | Check the MCP status using oc get mcp. Changing the ExecCPUAffinity status triggers a node reboot; ensure the update has completed. |
The exec process is pinned to an exclusive CPU instead of a shared CPU. |
Shared CPUs are not defined in the PerformanceProfile. |
When the MixedCPUsAllocation Technology Preview feature is enabled through the TechPreviewNoUpgrade feature set, the system’s CPU pinning logic for exec processes changes. If shared CPUs are defined in the PerformanceProfile under spec.cpu.shared and workloadHints.mixedCpus is set to true, the system prioritizes the first shared CPU. If no shared CPUs are defined, it defaults to the first exclusive (isolated) CPU. Enabling this feature set cannot be undone and is not recommended for production clusters. Additionally, the container must request shared CPUs by including workload.openshift.io/enable-shared-cpus: "1" in the resource limits. |
Additional resources
- About the Performance Profile Creator
- Disabling power saving mode for high priority pods
- Managing device interrupt processing for guaranteed pod isolated CPUs
CPUs for infra and application containers¶
Generic housekeeping and workload tasks use CPUs in a way that might impact latency-sensitive processes. By default, the container runtime uses all online CPUs to run all containers together, which can result in context switches and spikes in latency.
Partitioning the CPUs prevents noisy processes from interfering with latency-sensitive processes by separating them from each other. The following table describes how processes run on a CPU after you have tuned the node using the Node Tuning Operator:
Process' CPU assignments
| Process type | Details |
|---|---|
Burstable and BestEffort pods |
Runs on any CPU except where low latency workload is running |
| Infrastructure pods | Runs on any CPU except where low latency workload is running |
| Interrupts | Redirects to reserved CPUs (optional in OpenShift Container Platform 4.7 and later) |
| Kernel processes | Pins to reserved CPUs |
| Latency-sensitive workload pods | Pins to a specific set of exclusive CPUs from the isolated pool |
| OS processes/systemd services | Pins to reserved CPUs |
The allocatable capacity of cores on a node for pods of all QoS process types, Burstable, BestEffort, or Guaranteed, is equal to the capacity of the isolated pool. The capacity of the reserved pool is removed from the node’s total core capacity for use by the cluster and operating system housekeeping duties.
- Example 1
- A node features a capacity of 100 cores. Using a performance profile, the cluster administrator allocates 50 cores to the isolated pool and 50 cores to the reserved pool. The cluster administrator assigns 25 cores to QoS
Guaranteedpods and 25 cores forBestEffortorBurstablepods. This matches the capacity of the isolated pool. - Example 2
- A node features a capacity of 100 cores. Using a performance profile, the cluster administrator allocates 50 cores to the isolated pool and 50 cores to the reserved pool. The cluster administrator assigns 50 cores to QoS
Guaranteedpods and one core forBestEffortorBurstablepods. This exceeds the capacity of the isolated pool by one core. Pod scheduling fails because of insufficient CPU capacity.
The exact partitioning pattern to use depends on many factors like hardware, workload characteristics and the expected system load. Some sample use cases are as follows:
- If the latency-sensitive workload uses specific hardware, such as a network interface controller (NIC), ensure that the CPUs in the isolated pool are as close as possible to this hardware. At a minimum, you should place the workload in the same Non-Uniform Memory Access (NUMA) node.
- The reserved pool is used for handling all interrupts. When depending on system networking, allocate a sufficiently-sized reserve pool to handle all the incoming packet interrupts. In 4.22 and later versions, workloads can optionally be labeled as sensitive.
The decision regarding which specific CPUs should be used for reserved and isolated partitions requires detailed analysis and measurements. Factors like NUMA affinity of devices and memory play a role. The selection also depends on the workload architecture and the specific use case.
Warning
The reserved and isolated CPU pools must not overlap and together must span all available cores in the worker node.
To ensure that housekeeping tasks and workloads do not interfere with each other, specify two groups of CPUs in the spec section of the performance profile.
isolated- Specifies the CPUs for the application container workloads. These CPUs have the lowest latency. Processes in this group have no interruptions and can, for example, reach much higher DPDK zero packet loss bandwidth.reserved- Specifies the CPUs for the cluster and operating system housekeeping duties. Threads in thereservedgroup are often busy. Do not run latency-sensitive applications in thereservedgroup. Latency-sensitive applications run in theisolatedgroup.
Partition CPUs for infra and application containers¶
By partitioning CPUs, you can prevent noisy processes from interfering with latency-sensitive processes by separating the processes from each other.
Procedure
-
Create a performance profile appropriate for the environment’s hardware and topology. The following example adds the
reservedandisolatedparameters with the CPUs you want reserved and isolated for the infra and application containers:apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: infra-cpus spec: cpu: reserved: "0-4,9" isolated: "5-8" nodeSelector: node-role.kubernetes.io/worker: "" # ...where:
spec.cpu.reserved- Specifies which CPUs are for infra containers to perform cluster and operating system housekeeping duties.
spec.cpu.isolated- Specifies which CPUs are for application containers to run workloads.
spec.nodeSelector- Specifies a node selector to apply the performance profile to specific nodes. Optional parameter.
Configure Hyper-Threading for a cluster¶
To configure Hyper-Threading for an OpenShift Container Platform cluster, set the CPU threads in the performance profile to the same cores that are configured for the reserved or isolated CPU pools.
Note
If you configure a performance profile, and subsequently change the Hyper-Threading configuration for the host, ensure that you update the CPU isolated and reserved fields in the PerformanceProfile YAML to match the new configuration.
Warning
Disabling a previously enabled host Hyper-Threading configuration can cause the CPU core IDs listed in the PerformanceProfile YAML to be incorrect. This incorrect configuration can cause the node to become unavailable because the listed CPUs can no longer be found.
Prerequisites
- Access to the cluster as a user with the
cluster-adminrole. - Install the OpenShift CLI (
oc).
Procedure
-
Ascertain which threads are running on what CPUs for the host you want to configure.
You can view which threads are running on the host CPUs by logging in to the cluster and running the following command:
Example outputCPU NODE SOCKET CORE L1d:L1i:L2:L3 ONLINE MAXMHZ MINMHZ 0 0 0 0 0:0:0:0 yes 4800.0000 400.0000 1 0 0 1 1:1:1:0 yes 4800.0000 400.0000 2 0 0 2 2:2:2:0 yes 4800.0000 400.0000 3 0 0 3 3:3:3:0 yes 4800.0000 400.0000 4 0 0 0 0:0:0:0 yes 4800.0000 400.0000 5 0 0 1 1:1:1:0 yes 4800.0000 400.0000 6 0 0 2 2:2:2:0 yes 4800.0000 400.0000 7 0 0 3 3:3:3:0 yes 4800.0000 400.0000In this example, there are eight logical CPU cores running on four physical CPU cores. CPU0 and CPU4 are running on physical Core0, CPU1 and CPU5 are running on physical Core 1, and so on. Alternatively, to view the threads that are set for a particular physical CPU core (
cpu0in the example below), open a shell prompt and run the following: -
Apply the isolated and reserved CPUs in the
PerformanceProfileYAML. For example, you can set logical cores CPU0 and CPU4 asisolated, and logical cores CPU1 to CPU3 and CPU5 to CPU7 asreserved. When you configure reserved and isolated CPUs, the infra containers in pods use the reserved CPUs and the application containers use the isolated CPUs.Note
The reserved and isolated CPU pools must not overlap and together must span all available cores in the worker node.
Warning
Hyper-Threading is enabled by default on most Intel processors. If you enable Hyper-Threading, all threads processed by a particular core must be isolated or processed on the same core.
When Hyper-Threading is enabled, all guaranteed pods must use multiples of the simultaneous multi-threading (SMT) level to avoid a "noisy neighbor" situation that can cause the pod to fail. See Static policy options for more information.
Disable Hyper-Threading for low latency applications¶
When configuring clusters for low latency processing, consider whether you want to disable Hyper-Threading before you deploy the cluster.
To disable Hyper-Threading, perform the following steps:
Procedure
-
Create a performance profile that is appropriate for your hardware and topology. The following example sets
nosmtas an additional kernel argument:Example performance profileapiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: example-performanceprofile spec: additionalKernelArgs: - nmi_watchdog=0 - audit=0 - mce=off - processor.max_cstate=1 - idle=poll - intel_idle.max_cstate=0 - nosmt cpu: isolated: 2-3 reserved: 0-1 hugepages: defaultHugepagesSize: 1G pages: - count: 2 node: 0 size: 1G nodeSelector: node-role.kubernetes.io/performance: '' realTimeKernel: enabled: trueNote
When you configure reserved and isolated CPUs, the infra containers in pods use the reserved CPUs and the application containers use the isolated CPUs.
Managing device interrupt processing for guaranteed pod isolated CPUs¶
The Node Tuning Operator can manage host CPUs by dividing them into reserved CPUs for cluster and operating system housekeeping duties, including pod infra containers, and isolated CPUs for application containers to run the workloads.
By completing these tasks, you can set CPUs for low-latency workloads as isolated workloads.
Device interrupts are load balanced between all isolated and reserved CPUs to avoid CPUs being overloaded, with the exception of CPUs where there is a guaranteed pod running. Guaranteed pod CPUs are prevented from processing device interrupts when the relevant annotations are set for the pod.
In the performance profile, globallyDisableIrqLoadBalancing is used to manage whether device interrupts are processed or not. For certain workloads, the reserved CPUs are not always sufficient for dealing with device interrupts, and for this reason, device interrupts are not globally disabled on the isolated CPUs. By default, Node Tuning Operator does not disable device interrupts on isolated CPUs.
Finding the effective IRQ affinity setting for a node¶
Some IRQ controllers lack support for an IRQ affinity setting and might always expose all online CPUs as the IRQ mask. Because these IRQ controllers effectively run on CPU 0, you must find the effective IRQ affinity setting for a node.
The following are examples of drivers and hardware that Red Hat are aware lack support for IRQ affinity setting. The list is, by no means, exhaustive:
- Some RAID controller drivers, such as
megaraid_sas - Many non-volatile memory express (NVMe) drivers
- Some LAN on motherboard (LOM) network controllers
- The driver uses
managed_irqs
Note
The reason they do not support IRQ affinity setting might be associated with factors such as the type of processor, the IRQ controller, or the circuitry connections in the motherboard.
If the effective affinity of any IRQ is set to an isolated CPU, it might be a sign of some hardware or driver not supporting IRQ affinity setting. To find the effective affinity, log in to the host and run the following command:
/proc/irq/0/effective_affinity: 1
/proc/irq/1/effective_affinity: 8
/proc/irq/2/effective_affinity: 0
/proc/irq/3/effective_affinity: 1
/proc/irq/4/effective_affinity: 2
/proc/irq/5/effective_affinity: 1
/proc/irq/6/effective_affinity: 1
/proc/irq/7/effective_affinity: 1
/proc/irq/8/effective_affinity: 1
/proc/irq/9/effective_affinity: 2
/proc/irq/10/effective_affinity: 1
/proc/irq/11/effective_affinity: 1
/proc/irq/12/effective_affinity: 4
/proc/irq/13/effective_affinity: 1
/proc/irq/14/effective_affinity: 1
/proc/irq/15/effective_affinity: 1
/proc/irq/24/effective_affinity: 2
/proc/irq/25/effective_affinity: 4
/proc/irq/26/effective_affinity: 2
/proc/irq/27/effective_affinity: 1
/proc/irq/28/effective_affinity: 8
/proc/irq/29/effective_affinity: 4
/proc/irq/30/effective_affinity: 4
/proc/irq/31/effective_affinity: 8
/proc/irq/32/effective_affinity: 8
/proc/irq/33/effective_affinity: 1
/proc/irq/34/effective_affinity: 2
Some drivers use managed_irqs, whose affinity is managed internally by the kernel and userspace cannot change the affinity. In some cases, these IRQs might be assigned to isolated CPUs. For more information about managed_irqs, see "Affinity of managed interrupts cannot be changed even if they target isolated CPU".
Additional resources
Configure node interrupt affinity¶
Configure a cluster node for IRQ dynamic load balancing to control which cores can receive device interrupt requests (IRQ).
Prerequisites
- For core isolation, all server hardware components must support IRQ affinity. To check if the hardware components of your server support IRQ affinity, view the hardware specifications of the server or contact your hardware provider.
Procedure
-
Log in to the OpenShift Container Platform cluster as a user with cluster-admin privileges.
-
Set the performance profile
apiVersionto useperformance.openshift.io/v2. -
Remove the
globallyDisableIrqLoadBalancingfield or set it tofalse. -
Set the appropriate isolated and reserved CPUs. The following snippet illustrates a profile that reserves 2 CPUs. IRQ load-balancing is enabled for pods running on the
isolatedCPU set:apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: dynamic-irq-profile spec: cpu: isolated: 2-5 reserved: 0-1 ...Note
When you configure reserved and isolated CPUs, operating system processes, kernel processes, and systemd services run on reserved CPUs. Infrastructure pods run on any CPU except where the low latency workload is running. Low latency workload pods run on exclusive CPUs from the isolated pool. For more information, see "Partitioning CPUs for infra and application containers".
Configuring memory page sizes¶
By configuring memory page sizes, system administrators can implement more efficient memory management on a specific node to suit workload requirements. The Node Tuning Operator provides a method for configuring huge pages and kernel page sizes by using a performance profile.
Configure kernel page sizes¶
Use the kernelPageSize specification in a performance profile to configure the kernel page size on a specific node. Specify larger kernel page sizes for memory-intensive, high-performance workloads.
Note
For nodes with an x86_64 or AMD64 architecture, you can only specify 4k for the kernelPageSize specification. For nodes with an AArch64 architecture, you can specify 4k or 64k for the kernelPageSize specification. You must disable the realtime kernel before you can use the 64k option. The default value is 4k.
Prerequisites
- Access to the cluster as a user with the
cluster-adminrole. - Install the OpenShift CLI (
oc).
Procedure
-
Create a performance profile to target nodes where you want to configure the kernel page size by creating a YAML file that defines the
PerformanceProfileresource:Example pp-kernel-pages.yaml fileapiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: example-performance-profile #... spec: kernelPageSize: "64k" realTimeKernel: enabled: false nodeSelector: node-role.kubernetes.io/worker: ""where:
spec.kernelPageSize- Specifies a kernel page size of
64k. You can only specify64kfor nodes with an AArch64 architecture. The default value is4k. spec.realTimeKernel.enabled:false- Specifies whether to disable the realtime kernel. A setting of
falsedisables the kernel. You must disable the realtime kernel to use the64kkernel page size option. spec.nodeSelector.node-role.kubernetes.io/worker- Specifies targets nodes with the
workerrole.
-
Apply the performance profile to the cluster:
Verification
-
Start a debug session on the node where you applied the performance profile by running the following command:
<node_name>: Replace<node_name>with the name of the node with the performance profile applied.
-
Verify that the kernel page size is set to the value you specified in the performance profile by running the following command:
Configure huge pages¶
Because nodes must pre-allocate huge pages used in an OpenShift Container Platform cluster, use the Node Tuning Operator to allocate huge pages on a specific node.
OpenShift Container Platform provides a method for creating and allocating huge pages. Node Tuning Operator provides an easier method for doing this using the performance profile.
Procedure
-
In the
hugepages.pagessection of the performance profile, specify multiple blocks ofsize,count, and, optionally,node:Example configurationhugepages: defaultHugepagesSize: "1G" pages: - size: "1G" count: 4 node: 0 # ...where:
hugepages.pages.node- Specifies the
nodethat is the NUMA node in which the huge pages are allocated. If you omitnode, the pages are evenly spread across all NUMA nodes.
Note
Wait for the relevant machine config pool status that indicates the update is finished.
These are the only configuration steps you need to do to allocate huge pages.
Verification
-
To verify the configuration, see the
/proc/meminfofile on the node: -
Use
oc describeto report the new size:
Allocate multiple huge page sizes¶
You can request huge pages with different sizes under the same container. By doing this task, you can define more complicated pods consisting of containers with different huge page size needs.
The following example, shows you how to define sizes 1G and 2M. The Node Tuning Operator configures both sizes on the node.
Procedure
-
Edit the
PerformanceProfileobject to define1Gand2Msizes for the huge pages. The Node Tuning Operator configures both sizes on the node.
Reducing NIC queues using the Node Tuning Operator¶
The Node Tuning Operator facilitates reducing NIC queues for enhanced performance. Adjustments are made using the performance profile, allowing customization of queues for different network devices.
Adjust the NIC queues with the performance profile¶
You can use a performance profile to adjust the queue count for each network device. By using the Node Tuning Operator, you can reduce NIC queues for enhanced performance.
Supported network devices:
- Non-virtual network devices
- Network devices that support multiple queues (channels)
Unsupported network devices:
- Pure software network interfaces
- Block devices
- Intel DPDK virtual functions
Prerequisites
- Access to the cluster as a user with the
cluster-adminrole. - Install the OpenShift CLI (
oc).
Procedure
-
Log in to the OpenShift Container Platform cluster running the Node Tuning Operator as a user with
cluster-adminprivileges. -
Create and apply a performance profile appropriate for your hardware and topology. For guidance on creating a profile, see the "Creating a performance profile" section.
-
Edit this created performance profile:
-
Populate the
specfield with thenetobject. The object list can contain two fields:-
userLevelNetworkingis a required field specified as a boolean flag. IfuserLevelNetworkingistrue, the queue count is set to the reserved CPU count for all supported devices. The default isfalse. -
devicesis an optional field specifying a list of devices that will have the queues set to the reserved CPU count. If the device list is empty, the configuration applies to all network devices. The configuration is as follows:-
interfaceName: This field specifies the interface name, and it supports shell-style wildcards, which can be positive or negative.- Example wildcard syntax is as follows:
<string> .* - Negative rules are prefixed with an exclamation mark. To apply the net queue changes to all devices other than the excluded list, use
!<device>, for example,!eno1.
- Example wildcard syntax is as follows:
-
vendorID: The network device vendor ID represented as a 16-bit hexadecimal number with a0xprefix. -
deviceID: The network device ID (model) represented as a 16-bit hexadecimal number with a0xprefix.Note
When a
deviceIDis specified, thevendorIDmust also be defined. A device that matches all of the device identifiers specified in a device entryinterfaceName,vendorID, or a pair ofvendorIDplusdeviceIDqualifies as a network device. This network device then has its net queues count set to the reserved CPU count.When two or more devices are specified, the net queues count is set to any net device that matches one of them.
-
-
-
Set the queue count to the reserved CPU count for all devices by using this example performance profile:
-
Set the queue count to the reserved CPU count for all devices matching any of the defined device identifiers by using this example performance profile:
apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: manual spec: cpu: isolated: 3-51,55-103 reserved: 0-2,52-54 net: userLevelNetworking: true devices: - interfaceName: "eth0" - interfaceName: "eth1" - vendorID: "0x1af4" deviceID: "0x1000" nodeSelector: node-role.kubernetes.io/worker-cnf: "" # ... -
Set the queue count to the reserved CPU count for all devices starting with the interface name
ethby using this example performance profile: -
Set the queue count to the reserved CPU count for all devices with an interface named anything other than
eno1by using this example performance profile: -
Set the queue count to the reserved CPU count for all devices that have an interface name
eth0,vendorIDof0x1af4, anddeviceIDof0x1000by using this example performance profile:apiVersion: performance.openshift.io/v2 kind: PerformanceProfile metadata: name: manual spec: cpu: isolated: 3-51,55-103 reserved: 0-2,52-54 net: userLevelNetworking: true devices: - interfaceName: "eth0" - vendorID: "0x1af4" deviceID: "0x1000" nodeSelector: node-role.kubernetes.io/worker-cnf: "" # ... -
Apply the updated performance profile:
Verifying the queue status¶
To ensure that your performance profile changes are active, verify the queue status.
By reviewing the provided examples, you can confirm that specific tuning configurations are successfully applied to your environment.
In this section, several examples illustrate different performance profiles and how to verify the changes are applied.
- Example 1
-
Example 1 demonstrates that the net queue count that is set to the reserved CPU count (2) for all supported devices.
The relevant section from the performance profile is:
apiVersion: performance.openshift.io/v2 metadata: name: performance spec: kind: PerformanceProfile spec: cpu: reserved: 0-1 #total = 2 isolated: 2-8 net: userLevelNetworking: true # ...The following command displays the status of the queues associated with a device:
Note
Run this command on the node where the performance profile was applied.
The following command verifies the queue status before the profile is applied:
Example outputChannel parameters for ens4: Pre-set maximums: RX: 0 TX: 0 Other: 0 Combined: 4 Current hardware settings: RX: 0 TX: 0 Other: 0 Combined: 4The following command verifies the queue status after the profile is applied:
Example outputChannel parameters for ens4: Pre-set maximums: RX: 0 TX: 0 Other: 0 Combined: 4 Current hardware settings: RX: 0 TX: 0 Other: 0 Combined: 2Combined: Specifies the combined channel that shows the total count of reserved CPUs for all supported devices is 2. This matches what is configured in the performance profile.
- Example 2
-
Example 2 demonstrates that the net queue count is set to the reserved CPU count (2) for all supported network devices with a specific
vendorID.The relevant section from the performance profile is:
apiVersion: performance.openshift.io/v2 metadata: name: performance spec: kind: PerformanceProfile spec: cpu: reserved: 0-1 isolated: 2-8 net: userLevelNetworking: true devices: - vendorID = 0x1af4 # ...The following command displays the status of the queues associated with a device:
Note
Run this command on the node where the performance profile was applied.
The following command verifies the queue status after the profile is applied:
Example outputChannel parameters for ens4: Pre-set maximums: RX: 0 TX: 0 Other: 0 Combined: 4 Current hardware settings: RX: 0 TX: 0 Other: 0 Combined: 2Combined: Specifies that the total count of reserved CPUs for all supported devices withvendorID=0x1af4is 2. For example, if there is another network deviceens2withvendorID=0x1af4it will also have total net queues of 2. This matches what is configured in the performance profile.
- Example 3
-
Example 3 shows that the net queue count is set to the reserved CPU count (2) for all supported network devices that match any of the defined device identifiers. The command
udevadm infoprovides a detailed report on a device. In this example the devices are:# udevadm info -p /sys/class/net/ens4 ... E: ID_MODEL_ID=0x1000 E: ID_VENDOR_ID=0x1af4 E: INTERFACE=ens4 ...# udevadm info -p /sys/class/net/eth0 ... E: ID_MODEL_ID=0x1002 E: ID_VENDOR_ID=0x1001 E: INTERFACE=eth0 ...Set the net queues to 2 for a device with
interfaceNameequal toeth0and any devices that have avendorID=0x1af4with the following performance profile:apiVersion: performance.openshift.io/v2 metadata: name: performance spec: kind: PerformanceProfile spec: cpu: reserved: 0-1 #total = 2 isolated: 2-8 net: userLevelNetworking: true devices: - interfaceName = eth0 - vendorID = 0x1af4 # ...The following command verifies the queue status after the profile is applied:
Example outputChannel parameters for ens4: Pre-set maximums: RX: 0 TX: 0 Other: 0 Combined: 4 Current hardware settings: RX: 0 TX: 0 Other: 0 Combined: 2Combined: Specifies that the total count of reserved CPUs for all supported devices withvendorID=0x1af4is set to 2.
For example, if there is another network device
ens2withvendorID=0x1af4, it will also have the total net queues set to 2. Similarly, a device withinterfaceNameequal toeth0will have total net queues set to 2.
Logging associated with adjusting NIC queues¶
To verify NIC queue adjustments, review the Tuned daemon logs. These log messages detail the assigned devices that are recorded in the respective Tuned daemon logs.
The following messages might be recorded to the /var/log/tuned/tuned.log file:
-
An
INFOmessage is recorded detailing the successfully assigned devices: -
A
WARNINGmessage is recorded if none of the devices can be assigned: