# CYBER-FROST-3.8-BF16 deployment kit This kit reproduces the validated text-only BF16 serving profile for `CYBER-FROST-3.8-BF16`. It is a deployment recipe, not a weight-modification recipe. ## Validated profile | Item | Validated setting | |---|---| | Hardware | 4× NVIDIA B300 SXM6 | | Parallelism | tensor parallel 4 | | Runtime | vLLM `0.29.1rc1.dev13+g1cfd97281` | | Python ABI observed in the runtime | Python 3.12 | | Context exercised | 32,768 tokens | | Concurrent sequence ceiling | 4 | | GPU-memory utilization setting | 0.45 | | Attention/cache mode | prefix caching requested; Mamba SSM cache forced to BF16 | | Speculative decoding | native MTP, two draft tokens | | API | OpenAI-compatible `/v1/models` and `/v1/chat/completions` | | Validated modality | text only | The archived trial did not preserve immutable NVIDIA driver, CUDA, PyTorch, or container-image identifiers. The vLLM build identifier and launch flags are preserved exactly, but deployers must validate their own driver/framework combination before relying on the performance results. B300 is a Blackwell GPU and requires a software stack with matching architecture support. ## Prerequisites - Four visible GPUs with enough aggregate memory for the roughly 360 GB BF16 checkpoint, runtime allocations, and the requested context. - A vLLM build matching `0.29.1rc1.dev13+g1cfd97281` or a newer build independently validated with `qwen4_exp`, hybrid Mamba/full attention, native MTP, and the flags in `LAUNCH.sh`. - Hugging Face access approved for the manually gated repository and authentication through `hf auth login` or a process-level `HF_TOKEN`. - `curl` for smoke testing; `jq` is recommended. ## Launch ```bash cp DEPLOYMENT/ENV.EXAMPLE DEPLOYMENT/.env # Edit non-secret host settings if needed, then authenticate separately: hf auth login bash DEPLOYMENT/LAUNCH.sh ``` The launcher binds to `127.0.0.1:8000` by default. Keep the endpoint private or place it behind authenticated transport. Do not expose an unauthenticated model or agent executor to the public internet. Run the smoke test from another shell: ```bash bash DEPLOYMENT/SMOKE_TEST.sh ``` From the `DEPLOYMENT/` directory, verify the kit itself with `sha256sum -c SHA256SUMS`. ## Behavioral controls The packaged template defaults to thinking enabled and supports `xhigh`, `medium`, and `low` reasoning effort. The smoke test disables thinking only to keep the response short. Production callers should record the template, reasoning, sampling, and system-message settings used for each evaluation. Tool calls are model output, not authorization. If an external agent executes them, apply independent identity checks, explicit target scope, least-privilege credentials, sandboxing, command allowlists or policy checks, audit logs, rate limits, and approval gates for consequential actions. ## Runtime notes - The model config requests a float32 Mamba SSM cache. The validated speed profile overrides it to BF16. Remove `--mamba-ssm-cache-dtype bfloat16` when evaluating the configured precision instead of reproducing the speed trial. - Native MTP combined with the model's Mamba groups prevented cross-request prefix-cache reuse in the tested runtime even though prefix caching was requested. - Fused multi-step draft decode was unavailable for the experimental attention backend, so draft metadata was rebuilt between steps. - The native MTP tensors predate the final trunk-weight behavioral stage and are a provisional acceleration baseline. - The configured 262,144-token ceiling, multimodal path, tool-call correctness, high concurrency, and hardware other than the documented B300 profile are not qualified by this kit. ## Troubleshooting - **Repository access denied:** confirm that the Hugging Face account has been approved for the gate and that the runtime sees the correct token. - **Architecture is unknown:** use the pinned vLLM build or a build with explicit `qwen4_exp` support; keep `--trust-remote-code` enabled for this artifact. - **Out of memory:** verify that exactly four intended GPUs are visible, reduce `CYBER_FROST_MAX_MODEL_LEN`, lower `CYBER_FROST_MAX_NUM_SEQS`, or lower `CYBER_FROST_GPU_MEMORY_UTILIZATION` only after measuring the effect. - **Unexpected refusal or output style:** verify that the repository chat template is in use and record any caller system message, reasoning mode, and sampling overrides. - **Tool calls are plain text:** this profile validates text generation, not a particular tool parser or executor. Integrate and qualify those separately.