CosmicAC Logo

Recommended configuration values

Recommended values for a set of vLLM models and the Parakeet model.

CosmicAC can serve any model that vLLM supports. The following tables list recommended values for a set of models. The field names match the form that opens when you go to Models > Recommended configurations > Add new in the web interface. For what each vLLM environment variable does, see vLLM serving options.

Qwen3-VL-235B-A22B-Thinking-FP8

FieldValue
Model nameQwen/Qwen3-VL-235B-A22B-Thinking-FP8
Job typeManaged Inference
CUDA versionCUDA 13.0
Runtime imagevllm/vllm-openai:v0.15.1
Data typeAuto
QuantisationNone
Reasoning parserdeepseek_r1
Replicas1
GPUs per replica8
GPU memory utilisation0.9
Max model length131072
Max concurrent sequences64
Root disk (GB)500
Video and image inputOn

Environment variables

NameValue
TRUST_REMOTE_CODEtrue
SWAP_SPACE0
ENABLE_EXPERT_PARALLELtrue
ENFORCE_EAGERfalse

Qwen3.5-122B-A10B

FieldValue
Model nameQwen/Qwen3.5-122B-A10B
Job typeManaged Inference
CUDA versionCUDA 13.0
Runtime imagevllm/vllm-openai:v0.17.1
Data typeAuto
QuantisationNone
Reasoning parserdeepseek_r1
Replicas1
GPUs per replica8
GPU memory utilisation0.9
Max model length32768
Max concurrent sequences32
Root disk (GB)500
Video and image inputOn

Environment variables

NameValue
TRUST_REMOTE_CODEtrue
SWAP_SPACE0
ENABLE_EXPERT_PARALLELtrue
ENFORCE_EAGERtrue

MiniMax M2.5

FieldValue
Model nameMiniMaxAI/MiniMax-M2.5
Job typeManaged Inference
CUDA versionCUDA 13.0
Runtime imagevllm/vllm-openai:v0.15.1
Data typeAuto
QuantisationNone
Reasoning parserdeepseek_r1
Replicas1
GPUs per replica4
GPU memory utilisation0.85
Max model length131072
Max concurrent sequences32
Root disk (GB)500
Video and image inputOn

Environment variables

NameValue
TRUST_REMOTE_CODEtrue
ENABLE_EXPERT_PARALLELtrue
ENFORCE_EAGERfalse

Qwen2-VL-2B-Instruct

FieldValue
Model nameQwen/Qwen2-VL-2B-Instruct
Job typeManaged Inference
CUDA versionCUDA 13.0
Runtime imagevllm/vllm-openai:v0.15.1
Data typeAuto
QuantisationNone
Reasoning parserdefault
Replicas1
GPUs per replica1
GPU memory utilisation0.9
Max model length32768
Max concurrent sequences64
Root disk (GB)150
Video and image inputOn

Environment variables

NameValue
TRUST_REMOTE_CODEtrue
SWAP_SPACE0
ENFORCE_EAGERtrue

Parakeet

FieldValue
Model namenvidia/parakeet-tdt-0.6b-v3
Job typeParakeet Inference
Chunk duration (seconds)10
Chunk overlap (seconds)5
Maximum file size (MB)1024
Require authentication headerOff

What's next

On this page