CosmicAC Logo

Welcome

Run GPU Container Jobs and Managed Inference Jobs for language and speech-to-text models on your Kubernetes cluster.

CosmicAC is a self-hosted platform for running GPU workloads on your Kubernetes cluster. You deploy CosmicAC on your host machine, where it connects to your cluster and runs GPU Container Jobs and Managed Inference Jobs for language and speech-to-text models.

Get started

To deploy CosmicAC and create your first job, complete these steps in order.

  1. Check the requirements.
  2. Deploy CosmicAC.
  3. Set up recommended model configurations.
  4. Install the CLI.
  5. Create your first job with a GPU Container Job or a vLLM Managed Inference Job.

What you can run

CosmicAC runs two types of jobs. Choose a GPU Container Job to run code on a GPU machine with shell access. Choose a Managed Inference Job to serve a model through an API.

Understand and manage CosmicAC

See how CosmicAC works, then find guides for operating the platform, managing teams and access, and managing models.

Call a model

Each Managed Inference Job exposes an endpoint for its model. Create an API key, send requests to the endpoint, and check its health.

Reference

Look up exact commands, routes, fields, and values, and see what changed in each release.

On this page