CosmicAC Logo

Check the health of a Managed Inference endpoint

Check the status, success rate, and latency of a Managed Inference endpoint from the web interface or the CLI.

Check the status, success rate, and latency of your Managed Inference endpoints, and probe an endpoint when you need new results.

Prerequisites

Before you start, make sure that you have the following.

Steps

Open the health view

CosmicAC probes the replicas every five minutes by default, and both methods show the result of the latest health check.

In the left navigation, click Model Health. While the page is open, it refreshes the results every 30 seconds.

To change the time range, select 1H, 6H, 24H, 7D, or 30D at the top of the page. The default is 1H.

The time range changes the success rate, traffic, failures, and average response. It doesn't change the status, which CosmicAC always calculates over the last 24 hours.

Read the results

Each endpoint reports the following values.

  • Status: Healthy, Degraded, or Down. For what each value means, see Model health.
  • Success rate: the percentage of requests that succeeded.
  • Traffic: the total number of requests that the endpoint handled.
  • Failures: the number of those requests that failed.
  • Avg response: the average response time in milliseconds. The CLI shows it as Latency (avg).
  • Last health check: when CosmicAC last probed the endpoint. The CLI shows it as Last Check At.
  • Last updated: when CosmicAC generated the results. The CLI shows it as Timestamp.
  • Last Check: whether the last probe succeeded. Only the CLI shows this value.
  • Last Check Latency: how long the endpoint took to answer the last probe. Only the CLI shows this value.

If the endpoint handled no requests in the time range, the success rate and the average response are empty. The web interface shows a dash, and the CLI shows N/A.

To find an unhealthy replica, check the replica details.

  • Web interface: each endpoint card shows how many replicas are healthy. Below the count, the card lists each unhealthy replica by ID. A degraded replica also shows its success rate, traffic, failures, average response, and last check. A down replica shows Down · out of rotation.
  • CLI: the Replicas table lists every replica, with its ID, status, traffic, failures, and average latency.

Probe an endpoint on demand

To get new results before the next scheduled probe, probe the endpoint yourself.

CosmicAC probes a replica by sending it a real inference request, so a probe uses serving capacity. Probing every endpoint sends a request to every replica that you run. On a large deployment, probe one endpoint at a time.

To probe every endpoint, click Run health check at the top of the Model Health page. To probe one endpoint, click Health check on its card.

The results don't change right away. They update on the next refresh, within 30 seconds.

Next steps

On this page