Instructions to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/CYBER-FROST-3.8-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-BF16") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/CYBER-FROST-3.8-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
- SGLang
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-AI/CYBER-FROST-3.8-BF16 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-BF16
Download DEPLOYMENT/README.md from Blackfrost-AI/CYBER-FROST-3.8-BF16: direct link, hf CLI and curl.
- Browser
- Download file 4.57 kB
-
https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16/resolve/main/DEPLOYMENT/README.md
- Command line
-
hf download hf://Blackfrost-AI/CYBER-FROST-3.8-BF16/DEPLOYMENT/README.md
-
curl -L -o README.md https://huggingface.co/Blackfrost-AI/CYBER-FROST-3.8-BF16/resolve/main/DEPLOYMENT/README.md
CYBER-FROST-3.8-BF16 deployment kit
This kit reproduces the validated text-only BF16 serving profile for CYBER-FROST-3.8-BF16. It is a deployment recipe, not a weight-modification recipe.
Validated profile
| Item | Validated setting |
|---|---|
| Hardware | 4× NVIDIA B300 SXM6 |
| Parallelism | tensor parallel 4 |
| Runtime | vLLM 0.29.1rc1.dev13+g1cfd97281 |
| Python ABI observed in the runtime | Python 3.12 |
| Context exercised | 32,768 tokens |
| Concurrent sequence ceiling | 4 |
| GPU-memory utilization setting | 0.45 |
| Attention/cache mode | prefix caching requested; Mamba SSM cache forced to BF16 |
| Speculative decoding | native MTP, two draft tokens |
| API | OpenAI-compatible /v1/models and /v1/chat/completions |
| Validated modality | text only |
The archived trial did not preserve immutable NVIDIA driver, CUDA, PyTorch, or container-image identifiers. The vLLM build identifier and launch flags are preserved exactly, but deployers must validate their own driver/framework combination before relying on the performance results. B300 is a Blackwell GPU and requires a software stack with matching architecture support.
Prerequisites
- Four visible GPUs with enough aggregate memory for the roughly 360 GB BF16 checkpoint, runtime allocations, and the requested context.
- A vLLM build matching
0.29.1rc1.dev13+g1cfd97281or a newer build independently validated withqwen4_exp, hybrid Mamba/full attention, native MTP, and the flags inLAUNCH.sh. - Hugging Face access approved for the manually gated repository and authentication through
hf auth loginor a process-levelHF_TOKEN. curlfor smoke testing;jqis recommended.
Launch
cp DEPLOYMENT/ENV.EXAMPLE DEPLOYMENT/.env
# Edit non-secret host settings if needed, then authenticate separately:
hf auth login
bash DEPLOYMENT/LAUNCH.sh
The launcher binds to 127.0.0.1:8000 by default. Keep the endpoint private or place it behind authenticated transport. Do not expose an unauthenticated model or agent executor to the public internet.
Run the smoke test from another shell:
bash DEPLOYMENT/SMOKE_TEST.sh
From the DEPLOYMENT/ directory, verify the kit itself with sha256sum -c SHA256SUMS.
Behavioral controls
The packaged template defaults to thinking enabled and supports xhigh, medium, and low reasoning effort. The smoke test disables thinking only to keep the response short. Production callers should record the template, reasoning, sampling, and system-message settings used for each evaluation.
Tool calls are model output, not authorization. If an external agent executes them, apply independent identity checks, explicit target scope, least-privilege credentials, sandboxing, command allowlists or policy checks, audit logs, rate limits, and approval gates for consequential actions.
Runtime notes
- The model config requests a float32 Mamba SSM cache. The validated speed profile overrides it to BF16. Remove
--mamba-ssm-cache-dtype bfloat16when evaluating the configured precision instead of reproducing the speed trial. - Native MTP combined with the model's Mamba groups prevented cross-request prefix-cache reuse in the tested runtime even though prefix caching was requested.
- Fused multi-step draft decode was unavailable for the experimental attention backend, so draft metadata was rebuilt between steps.
- The native MTP tensors predate the final trunk-weight behavioral stage and are a provisional acceleration baseline.
- The configured 262,144-token ceiling, multimodal path, tool-call correctness, high concurrency, and hardware other than the documented B300 profile are not qualified by this kit.
Troubleshooting
- Repository access denied: confirm that the Hugging Face account has been approved for the gate and that the runtime sees the correct token.
- Architecture is unknown: use the pinned vLLM build or a build with explicit
qwen4_expsupport; keep--trust-remote-codeenabled for this artifact. - Out of memory: verify that exactly four intended GPUs are visible, reduce
CYBER_FROST_MAX_MODEL_LEN, lowerCYBER_FROST_MAX_NUM_SEQS, or lowerCYBER_FROST_GPU_MEMORY_UTILIZATIONonly after measuring the effect. - Unexpected refusal or output style: verify that the repository chat template is in use and record any caller system message, reasoning mode, and sampling overrides.
- Tool calls are plain text: this profile validates text generation, not a particular tool parser or executor. Integrate and qualify those separately.