Managed Inference architecture
How a Managed Inference Job runs on your cluster, and how CosmicAC authenticates and routes an inference request.
A Managed Inference Job runs an open source model inside a KubeVirt virtual machine instance (VMI). It uses vLLM for language models or Parakeet for speech-to-text. cosmicac-proxy-inference exposes that model as an OpenAI-compatible endpoint, authenticates requests, and balances load. You reach the model through that endpoint from any OpenAI-compatible client, or by running inference directly with cosmicac-cli.
The following diagram shows how a job request reaches your cluster and how an inference request reaches the model. The example is a multi-node replica on two nodes. A replica that fits on one node runs a single VMI.
The inference agent
Every VMI in a replica runs an inference agent, cosmicac-wrk-agent-inference. The agent runs the model server in a second container that's built from the vLLM or Parakeet runtime image. If that container exits, the agent restarts it, up to three times by default. After that, the agent marks the model server as failed.
The agent isn't the model server. Every VMI runs one of each, and because the agent owns the server's container, it reads what the server prints and forwards it into the job's Application logs.
How a job starts
When you create a Managed Inference Job from cosmicac-ui or cosmicac-cli, cosmicac-app-node authenticates the request and forwards it to cosmicac-wrk-ork. The orchestrator allocates GPUs for each replica on nodes that the job's team can use, and then hands the job to cosmicac-wrk-server-k8s-nvidia. That worker creates the job's Kubernetes resources through your cluster's API server, and Kubernetes runs each replica as a pod that holds a VMI.
A multi-node replica runs one pod and one VMI on each of its nodes. The first VMI runs the model server and answers requests, and the others add their GPUs to that server. Ray forms the cluster over the overlay network, and the NVIDIA Collective Communications Library (NCCL) carries the GPU-to-GPU traffic over InfiniBand.
As a replica starts, the inference agent on its serving node registers in the distributed hash table (DHT), so cosmicac-proxy-inference can discover the replica.
How CosmicAC serves a request
Serving traffic follows a separate path from job creation. You send a request to the inference endpoint from any OpenAI-compatible client, or run inference from cosmicac-cli. cosmicac-proxy-inference checks the API key when the endpoint requires one. When the deployment turns on team scoping, the proxy also checks that the key belongs to the endpoint's team. The proxy searches the DHT by topic to discover the model servers, and sends requests to them in turn. The inference agent then runs the request on the job's model server and returns the response.
The proxy also probes each replica, and stops sending requests to a replica that's unhealthy. See Model health.