Blackfrost-AI's picture
Harden Cyber-Frost deployment kit
98657c2 verified
|
Raw History Blame Contribute Delete
4.57 kB

CYBER-FROST-3.8-BF16 deployment kit

This kit reproduces the validated text-only BF16 serving profile for CYBER-FROST-3.8-BF16. It is a deployment recipe, not a weight-modification recipe.

Validated profile

Item Validated setting
Hardware 4× NVIDIA B300 SXM6
Parallelism tensor parallel 4
Runtime vLLM 0.29.1rc1.dev13+g1cfd97281
Python ABI observed in the runtime Python 3.12
Context exercised 32,768 tokens
Concurrent sequence ceiling 4
GPU-memory utilization setting 0.45
Attention/cache mode prefix caching requested; Mamba SSM cache forced to BF16
Speculative decoding native MTP, two draft tokens
API OpenAI-compatible /v1/models and /v1/chat/completions
Validated modality text only

The archived trial did not preserve immutable NVIDIA driver, CUDA, PyTorch, or container-image identifiers. The vLLM build identifier and launch flags are preserved exactly, but deployers must validate their own driver/framework combination before relying on the performance results. B300 is a Blackwell GPU and requires a software stack with matching architecture support.

Prerequisites

  • Four visible GPUs with enough aggregate memory for the roughly 360 GB BF16 checkpoint, runtime allocations, and the requested context.
  • A vLLM build matching 0.29.1rc1.dev13+g1cfd97281 or a newer build independently validated with qwen4_exp, hybrid Mamba/full attention, native MTP, and the flags in LAUNCH.sh.
  • Hugging Face access approved for the manually gated repository and authentication through hf auth login or a process-level HF_TOKEN.
  • curl for smoke testing; jq is recommended.

Launch

cp DEPLOYMENT/ENV.EXAMPLE DEPLOYMENT/.env
# Edit non-secret host settings if needed, then authenticate separately:
hf auth login
bash DEPLOYMENT/LAUNCH.sh

The launcher binds to 127.0.0.1:8000 by default. Keep the endpoint private or place it behind authenticated transport. Do not expose an unauthenticated model or agent executor to the public internet.

Run the smoke test from another shell:

bash DEPLOYMENT/SMOKE_TEST.sh

From the DEPLOYMENT/ directory, verify the kit itself with sha256sum -c SHA256SUMS.

Behavioral controls

The packaged template defaults to thinking enabled and supports xhigh, medium, and low reasoning effort. The smoke test disables thinking only to keep the response short. Production callers should record the template, reasoning, sampling, and system-message settings used for each evaluation.

Tool calls are model output, not authorization. If an external agent executes them, apply independent identity checks, explicit target scope, least-privilege credentials, sandboxing, command allowlists or policy checks, audit logs, rate limits, and approval gates for consequential actions.

Runtime notes

  • The model config requests a float32 Mamba SSM cache. The validated speed profile overrides it to BF16. Remove --mamba-ssm-cache-dtype bfloat16 when evaluating the configured precision instead of reproducing the speed trial.
  • Native MTP combined with the model's Mamba groups prevented cross-request prefix-cache reuse in the tested runtime even though prefix caching was requested.
  • Fused multi-step draft decode was unavailable for the experimental attention backend, so draft metadata was rebuilt between steps.
  • The native MTP tensors predate the final trunk-weight behavioral stage and are a provisional acceleration baseline.
  • The configured 262,144-token ceiling, multimodal path, tool-call correctness, high concurrency, and hardware other than the documented B300 profile are not qualified by this kit.

Troubleshooting

  • Repository access denied: confirm that the Hugging Face account has been approved for the gate and that the runtime sees the correct token.
  • Architecture is unknown: use the pinned vLLM build or a build with explicit qwen4_exp support; keep --trust-remote-code enabled for this artifact.
  • Out of memory: verify that exactly four intended GPUs are visible, reduce CYBER_FROST_MAX_MODEL_LEN, lower CYBER_FROST_MAX_NUM_SEQS, or lower CYBER_FROST_GPU_MEMORY_UTILIZATION only after measuring the effect.
  • Unexpected refusal or output style: verify that the repository chat template is in use and record any caller system message, reasoning mode, and sampling overrides.
  • Tool calls are plain text: this profile validates text generation, not a particular tool parser or executor. Integrate and qualify those separately.