Managing workloads with the JobSet Operator
Use the JobSet Operator on OpenShift Container Platform to manage and run large-scale, coordinated workloads like high-performance computing (HPC) and AI training. Features like multi-template job support and stable networking can help you recover quickly and use resources efficiently.
Deploying a JobSet
You can use the JobSet Operator to deploy a JobSet to manage and run large-scale, coordinated workloads.
Prerequisites
- You have installed the JobSet Operator.
- You have a cluster with available NVIDIA GPUs.
Procedure
-
Create a new project by running the following command:
$ oc new-project <my_namespace> -
Create a file named
jobset.yaml:apiVersion: jobset.x-k8s.io/v1alpha2kind: JobSetmetadata:name: pytorchspec:replicatedJobs:- name: workerstemplate:spec:parallelism: 3completions: 3backoffLimit: 0template:spec:imagePullSecrets:- name: my-registry-secretinitContainers:- name: prepareimage: docker.io/alpine/git:v2.52.0args: ['clone', 'https://github.com/pytorch/examples']volumeMounts:- name: workdirmountPath: /gitcontainers:- name: pytorchimage: docker.io/pytorch/pytorch:2.10.0-cuda13.0-cudnn9-runtimeresources:limits:nvidia.com/gpu: "1"requests:nvidia.com/gpu: "1"ports:- containerPort: 4321env:- name: MASTER_ADDRvalue: "pytorch-workers-0-0.pytorch"- name: MASTER_PORTvalue: "4321"- name: RANKvalueFrom:fieldRef:fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']- name: PYTHONUNBUFFEREDvalue: "0"command:- /bin/sh- -c- |cd examples/distributed/ddp-tutorial-seriestorchrun --nproc_per_node=1 --nnodes=3 --rdzv_id=100 --rdzv_backend=c10d --rdzv_endpoint=$MASTER_ADDR:$MASTER_PORT multinode.py 1000 100volumeMounts:- name: workdirmountPath: /workspacevolumes:- name: workdiremptyDir: {}where:
spec.replicatedJobs.template.spec.parallelism- Specifies the number of pods running at the same time.
spec.replicatedJobs.template.spec.completions- Specifies the total number of pods that must finish successfully for the job to be marked complete.
-
Apply the JobSet configuration by running the following command:
$ oc apply -f jobset.yaml
Verification
-
Verify that pods were started by running the following command:
$ oc get pods -n <my_namespace>Example outputNAME READY STATUS RESTARTS AGEpytorch-workers-0-0-2lzwt 1/1 Running 0 2m17spytorch-workers-0-1-g2lrv 1/1 Running 0 2m17spytorch-workers-0-2-dpljq 1/1 Running 0 2m17s
Specifying a JobSet coordinator
To manage communication between JobSet pods, you can assign a specific JobSet coordinator pod. This ensures that your distributed workloads can reference a stable network endpoint as a central point of coordination for task synchronization and data exchange.
Prerequisites
- You have installed the JobSet Operator.
Procedure
-
Create a new namespace by running the following command.
$ oc new-project <new_namespace> -
Create a YAML file called
jobset-coordinator.yaml:Example YAML fileapiVersion: jobset.x-k8s.io/v1alpha2kind: JobSetmetadata:name: coordinatorspec:coordinator:replicatedJob: driverjobIndex: 0podIndex: 0replicatedJobs:- name: workerstemplate:spec:parallelism: <pods_running_number>completions: <pods_finish_number>backoffLimit: 0template:spec:containers:- name: workerenv:- name: COORDINATOR_ENDPOINTvalueFrom:fieldRef:fieldPath: metadata.labels['jobset.sigs.k8s.io/coordinator']image: quay.io/nginx/nginx-unprivileged:1.29-alpinecommand: [ "/bin/sh", "-c" ]args:- |while ! curl -s "${COORDINATOR_ENDPOINT}:8080" | grep Welcome; dosleep 3donesleep 100- name: drivertemplate:spec:parallelism: <pods_running_number>completions: <pods_finish_number>backoffLimit: 0template:spec:containers:- name: driverimage: quay.io/nginx/nginx-unprivileged:1.29-alpineports:- containerPort: 8080protocol: TCPwhere:
<pods_running_number>- Specifies the number of pods running at the same time.
<pods_finish_number>- Specifies the total number of pods that must finish successfully for the job to be marked complete.
-
Apply the
jobset-coordinator.yamlfile by running the following command:$ oc apply -f jobset-coordinator.yaml
Verification
-
Verify that pods were created by running the following command:
$ oc get pods -n <new_namespace>Example outputNAME READY STATUS RESTARTS AGEcoordinator-driver-0-0-svgk7 1/1 Running 0 67scoordinator-workers-0-0-57jvg 1/1 Running 0 67scoordinator-workers-0-1-mghvx 1/1 Running 0 67scoordinator-workers-0-2-7cnvv 1/1 Running 0 67s
Failure policy configuration for JobSet Operator
To control workload behavior in response to child job failures, you can configure a JobSet failure policy. This enables you to define specific actions, such as restarting or failing the entire JobSet, based on the failure reason or the specific replicated job affected.
Failure policy actions
These actions are available when a job failure matches a defined rule.
| Action | Description |
|---|---|
FailJobSet | Marks the entire JobSet as failed immediately. |
RestartJobSet | Restarts the JobSet by recreating all child jobs. This action counts toward the maxRestarts limit. This is the default action if no rules match. |
RestartJobSetAndIgnoreMaxRestarts | Restarts the JobSet without counting toward the maxRestarts limit. |