Storage checkups
You can use a storage checkup to verify that the cluster storage is optimally configured for OpenShift Virtualization.
Running a predefined checkup in an existing namespace involves setting up a service account for the checkup, creating the Role and RoleBinding objects for the service account, enabling permissions for the checkup, and creating the input config map and the checkup job. You can run a checkup multiple times.
You must always:
- Verify that the checkup image is from a trustworthy source before applying it.
- Review the checkup permissions before creating the
RoleandRoleBindingobjects.
Retaining resources for troubleshooting storage checkups
The predefined storage checkup includes skipTeardown configuration options, which control resource clean up after a storage checkup runs. By default, the skipTeardown field value is Never, which means that the checkup always performs teardown steps and deletes all resources after the checkup runs.
You can retain resources for further inspection in case a failure occurs by setting the skipTeardown field to onfailure.
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
- Run the following command to edit the
storage-checkup-configconfig map:$ oc edit configmap storage-checkup-config -n <checkup_namespace> - Configure the
skipTeardownfield to use theonfailurevalue. You can do this by modifying thestorage-checkup-configconfig map, stored in thestorage_checkup.yamlfile:apiVersion: v1kind: ConfigMapmetadata:name: storage-checkup-confignamespace: <checkup_namespace>data:spec.param.skipTeardown: onfailure# ... - Reapply the
storage-checkup-configconfig map by running the following command:$ oc apply -f storage_checkup.yaml -n <checkup_namespace>
Running a storage checkup by using the web console
You can run a storage checkup to validate that storage is working correctly for virtual machines.
Procedure
- Navigate to Virtualization → Checkups in the web console.
- Click the Storage tab.
- Click Install permissions.
- Click Run checkup.
- Enter a name for the checkup in the Name field.
- Enter a timeout value for the checkup in the Timeout (minutes) fields.
- Click Run.
Result
You can view the status of the storage checkup in the Checkups list on the Storage tab. Click on the name of the checkup for more details.
Running a storage checkup by using the CLI
Use a predefined checkup to verify that the OpenShift Container Platform cluster storage is configured optimally to run OpenShift Virtualization workloads.
Prerequisites
-
You have installed the OpenShift CLI (
oc). -
The cluster administrator has created the required
cluster-readerpermissions for the storage checkup service account and namespace, such as in the following example:apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRoleBindingmetadata:name: kubevirt-storage-checkup-clustereaderroleRef:apiGroup: rbac.authorization.k8s.iokind: ClusterRolename: cluster-readersubjects:- kind: ServiceAccountname: storage-checkup-sanamespace: <target_namespace>where:
<target_namespace>- Specifies the namespace where the checkup is to be run.
Procedure
-
Create a
ServiceAccount,Role, andRoleBindingmanifest file for the storage checkup. Example service account, role, and rolebinding manifest:---apiVersion: v1kind: ServiceAccountmetadata:name: storage-checkup-sa---apiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata:name: storage-checkup-rolerules:- apiGroups: [ "" ]resources: [ "configmaps" ]verbs: ["get", "update"]- apiGroups: [ "kubevirt.io" ]resources: [ "virtualmachines" ]verbs: [ "create", "delete" ]- apiGroups: [ "kubevirt.io" ]resources: [ "virtualmachineinstances" ]verbs: [ "get" ]- apiGroups: [ "subresources.kubevirt.io" ]resources: [ "virtualmachineinstances/addvolume", "virtualmachineinstances/removevolume" ]verbs: [ "update" ]- apiGroups: [ "kubevirt.io" ]resources: [ "virtualmachineinstancemigrations" ]verbs: [ "create" ]- apiGroups: [ "cdi.kubevirt.io" ]resources: [ "datavolumes" ]verbs: [ "create", "delete" ]- apiGroups: [ "" ]resources: [ "persistentvolumeclaims" ]verbs: [ "delete" ]---apiVersion: rbac.authorization.k8s.io/v1kind: RoleBindingmetadata:name: storage-checkup-rolesubjects:- kind: ServiceAccountname: storage-checkup-saroleRef:apiGroup: rbac.authorization.k8s.iokind: Rolename: storage-checkup-role -
Apply the
ServiceAccount,Role, andRoleBindingmanifest in the target namespace:$ oc apply -n <target_namespace> -f <storage_sa_roles_rolebinding>.yaml -
Create a
ConfigMapandJobmanifest file. The config map contains the input parameters for the checkup job. Example input config map and job manifest:---apiVersion: v1kind: ConfigMapmetadata:name: storage-checkup-confignamespace: $CHECKUP_NAMESPACEdata:spec.timeout: 10mspec.param.storageClass: ocs-storagecluster-ceph-rbd-virtualizationspec.param.vmiTimeout: 3m---apiVersion: batch/v1kind: Jobmetadata:name: storage-checkupnamespace: $CHECKUP_NAMESPACEspec:backoffLimit: 0template:spec:serviceAccount: storage-checkup-sarestartPolicy: Nevercontainers:- name: storage-checkupimage: quay.io/kiagnose/kubevirt-storage-checkup:mainimagePullPolicy: Alwaysenv:- name: CONFIGMAP_NAMESPACEvalue: $CHECKUP_NAMESPACE- name: CONFIGMAP_NAMEvalue: storage-checkup-config -
Apply the
ConfigMapandJobmanifest file in the target namespace to run the checkup:$ oc apply -n <target_namespace> -f <storage_configmap_job>.yaml -
Wait for the job to complete:
$ oc wait job storage-checkup -n <target_namespace> --for condition=complete --timeout 10m -
Review the results of the checkup by running the following command:
$ oc get configmap storage-checkup-config -n <target_namespace> -o yamlExample output config map (success):
apiVersion: v1kind: ConfigMapmetadata:name: storage-checkup-configlabels:kiagnose/checkup-type: kubevirt-storagedata:spec.timeout: 10mstatus.succeeded: "true"status.failureReason: ""status.startTimestamp: "2023-07-31T13:14:38Z"status.completionTimestamp: "2023-07-31T13:19:41Z"status.result.cnvVersion: 4.21.2status.result.defaultStorageClass: trident-nfsstatus.result.goldenImagesNoDataSource: <data_import_cron_list>status.result.goldenImagesNotUpToDate: <data_import_cron_list>status.result.ocpVersion: 4.21.0status.result.pvcBound: "true"status.result.storageProfileMissingVolumeSnapshotClass: <storage_class_list>status.result.storageProfilesWithEmptyClaimPropertySets: <storage_profile_list>status.result.storageProfilesWithSmartClone: <storage_profile_list>status.result.storageProfilesWithSpecClaimPropertySets: <storage_profile_list>status.result.storageProfilesWithRWX: |-ocs-storagecluster-ceph-rbdocs-storagecluster-ceph-rbd-virtualizationocs-storagecluster-cephfstrident-iscsitrident-miniotrident-nfswindows-vmsstatus.result.vmBootFromGoldenImage: VMI "vmi-under-test-dhkb8" successfully bootedstatus.result.vmHotplugVolume: |-VMI "vmi-under-test-dhkb8" hotplug volume readyVMI "vmi-under-test-dhkb8" hotplug volume removedstatus.result.vmLiveMigration: VMI "vmi-under-test-dhkb8" migration completedstatus.result.vmVolumeClone: 'DV cloneType: "csi-clone"'status.result.vmsWithNonVirtRbdStorageClass: <vm_list>status.result.vmsWithUnsetEfsStorageClass: <vm_list>
data.status.succeededdefines if the checkup is successful (true) or not (false).data.status.failureReasondefines the reason for failure if the checkup fails.data.status.startTimestampdefines the time when the checkup started, in RFC 3339 time format.data.status.completionTimestampdefines the time when the checkup has completed, in RFC 3339 time format.data.status.result.cnvVersiondefines the OpenShift Virtualization version.data.status.result.defaultStorageClassdefines if there is a default storage class.data.status.result.goldenImagesNoDataSourcedefines the list of golden images whose data source is not ready.data.status.result.goldenImagesNotUpToDatedefines the list of golden images whose data import cron is not up-to-date.data.status.result.ocpVersiondefines the OpenShift Container Platform version.data.status.result.pvcBounddefines if a PVC of 10Mi has been created and bound by the provisioner.data.status.result.storageProfileMissingVolumeSnapshotClassdefines the list of storage profiles using snapshot-based clone but missing VolumeSnapshotClass.data.status.result.storageProfilesWithEmptyClaimPropertySetsdefines the list of storage profiles with unknown provisioners.data.status.result.storageProfilesWithSmartClonedefines the list of storage profiles with smart clone support (CSI/snapshot).data.status.result.storageProfilesWithSpecClaimPropertySetsdefines the list of storage profiles spec-overriden claimPropertySets.data.status.result.vmsWithNonVirtRbdStorageClassdefines the list of virtual machines that use the Ceph RBD storage class when the virtualization storage class exists.data.status.result.vmsWithUnsetEfsStorageClassdefines the list of virtual machines that use an Elastic File Store (EFS) storage class where the GID and UID are not set in the storage class.-
Delete the job and config map that you previously created by running the following commands:
$ oc delete job -n <target_namespace> storage-checkup$ oc delete config-map -n <target_namespace> storage-checkup-config -
Optional: If you do not plan to run another checkup, delete the
ServiceAccount,Role, andRoleBindingmanifest:$ oc delete -f <storage_sa_roles_rolebinding>.yaml
-
Troubleshooting a failed storage checkup
If a storage checkup fails, there are steps that you can take to identify the reason for failure.
Prerequisites
- You have installed the OpenShift CLI (
oc). - You have downloaded the directory provided by the
must-gathertool.
Procedure
-
Review the
status.failureReasonfield in thestorage-checkup-configconfig map by running the following command and observing the output:$ oc get configmap storage-checkup-config -n <namespace> -o yamlExample output config map:
apiVersion: v1kind: ConfigMapmetadata:name: storage-checkup-configlabels:kiagnose/checkup-type: kubevirt-storagedata:spec.timeout: 10mstatus.succeeded: "false"status.failureReason: "ErrNoDefaultStorageClass"# ...- If the checkup has failed, the
status.succeededvalue isfalse. - If the checkup has failed, the
status.failureReasonfield contains an error message. In this example output, theErrNoDefaultStorageClasserror message means that no default storage class is configured.
- If the checkup has failed, the
-
Search the directory provided by the
must-gathertool for logs, events, or terms related to the error in thedata.status.failureReasonfield value.
Additional resources
Storage checkup error codes
Storage checkup error codes might appear in the storage-checkup-config config map after a storage checkup fails.
| Error code | Meaning |
|---|---|
ErrNoDefaultStorageClass | No default storage class is configured. |
ErrPvcNotBound | One or more persistent volume claims (PVCs) failed to bind. |
ErrMultipleDefaultStorageClasses | Multiple default storage classes are configured. |
ErrEmptyClaimPropertySets | There are StorageProfile objects containing empty ClaimPropertySets specs. |
ErrVMsWithUnsetEfsStorageClass | There are VMs using elastic file system (EFS) storage classes, where the GID and UID are not set in the StorageClass object. |
ErrGoldenImagesNotUpToDate | One or more golden images has a DataImportCron object that is either not up to date or has a DataSource object which is not ready. |
ErrGoldenImageNoDataSource | The DataSource object of the golden image has either no PVC or no snapshot source configured. |
ErrBootFailedOnSomeVMs | Some VMs failed to boot within the expected time. |