Create a GPU Container Job with the CLI
Create a GPU Container Job with the CLI and verify that CosmicAC created it.
Create a GPU Container Job with the CosmicAC CLI. You can create the job interactively by answering prompts, or non-interactively by passing flags.
Prerequisites
Before you start, make sure that you have the following.
- A running CosmicAC deployment. See Set up CosmicAC.
- The CosmicAC CLI installed and configured. See Install the CLI.
Steps
Create the job
Create the job interactively, or pass the job configuration as flags.
Start the interactive job setup.
cosmicac jobs createSelect GPU Container as the job type, and then answer the prompts.
Configure the following fields.
- Job name: a name that identifies the job.
- Tags: comma-separated labels for the job.
- Location: the region where the job runs.
- GPU type: the GPU to use. The CLI lists the GPUs available in the selected location.
- GPU count: the number of GPUs for the job. One of 1, 2, 4, or 8. The CLI lists only the counts with enough free GPUs.
- Base image: the CLI shows the base image it uses, such as
Ubuntu22.04/CUDA12.9. - Root disk size: the disk size in GB. Interactive mode offers 250, 500, or 1000. With
--root-disk-size-gb, the minimum is 100. - Notifications: the job lifecycle events this job reports. All four are on by default, and interactive mode prompts for them with a checkbox.
An event reaches your webhook only if it's also turned on in Settings > Notifications. See What controls delivery.
In interactive mode, the CLI shows a job summary and asks you to confirm. If a team is active, the prompt names it. Answer yes to create the job, and the CLI prints the job ID.
GPU Container Job configuration describes each field and its CLI flag.
Confirm the job
List your jobs to confirm that CosmicAC created the job.
cosmicac jobs listThe job appears in the table with its ID, name, tags, and status. Wait for it to provision. You can connect to the job once its status is running. See Connect to a GPU Container Job.
Help and troubleshooting
Job stuck in Creating or Starting
If a job stays in Creating or Starting, check the status of its KubeVirt virtual machine instance (VMI).
-
Find the job's container ID.
cosmicac jobs detail <jobId>The output lists the Container ID for each container.
-
Find the VMI for the container.
CosmicAC creates one VMI for each container and names it
<container-id>-n0. A multi-node job has one VMI per node.From a machine with
kubectlaccess to your Kubernetes cluster, run the following command.kubectl get vmi -n <namespace>Replace
<namespace>with the namespace configured inK8S_NAMESPACE. -
Check the VMI status.
-
If the VMI is not Running, inspect its events.
kubectl describe vmi <container-id>-n0 -n <namespace> -
If the VMI is Running but the job stays in Creating or Starting, cosmicac-wrk-agent-instance cannot reach cosmicac-wrk-server-k8s-nvidia. These two components connect directly, and some cluster network configurations can block the connection.
To route the connection through a relay, see Set up a relay for CosmicAC.
-