Instructions to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2
- SGLang
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2
- CYBER-FROST-3.8-NVFP4-V2
- October 2, 2026: V2 gate/up scale correction
- Cyber-Frost Harness evaluation
- Release status and contents
- Why Cyber-Frost exists
- Security corpus
- Model specifications and precision layout
- Lineage
- Artifact verification
- Evaluation status
- Prompt, tool use, and sampling
- Deployment
- Intended use
- Limitations and security responsibility
- License and disclaimer
- October 2, 2026: V2 gate/up scale correction
CYBER-FROST-3.8-NVFP4-V2
A first-party Blackfrost-AI mixed-precision NVFP4 model for security professionals conducting authorized research, assessment, engineering, and response work.
October 2, 2026: V2 gate/up scale correction
V2 is a fresh NVFP4 re-export of the same final Blackfrost-AI/CYBER-FROST-3.8-BF16 checkpoint. It corrects a gate/up quantization-scale compatibility issue in the original NVFP4 release; it is not a new training, LoRA, or DWM run.
Thank you to @sethforprivacy for identifying and investigating the issue in the original release discussion.
The original export quantized each routed expert's gate and up projections independently, producing potentially different FP32 global weight scales (weight_scale_2). In the vLLM fused-W13 path described in the feedback, a single global scale is used for both projections: the gate scale is retained and also applied to the up projection. When the original global scales differ, this mis-scales the up projection. Other fused gate/up backends that use one global scale can encounter the same issue. This depends on the selected runtime and kernel path; it does not imply that every NVFP4 serving backend handled the original checkpoint incorrectly.
For V2, each expert's gate/up pair shares the maximum absolute BF16 weight value across both projections before calculating new FP8 E4M3 block scales and packing new NVFP4 weights. The resulting gate/up global scales are exactly equal. This is a BF16 re-export with consistently recomputed block scales and packed weights, not merely an overwrite of existing scale metadata.
Validation of the saved V2 artifact confirmed:
- All 24,576 gate/up global-scale pairs are bit-identical within each pair.
- All 73,728 calibrated input scales are bit-identical to the existing NVIDIA-derived calibrated reference.
- All 1,562 non-routed source tensors retain their BF16-source shapes and dtypes. The export preserved these tensors without numerical conversion; the saved-file check compared metadata, not full weight hashes.
- The checkpoint contains 131 SafeTensors shards and 296,474 tensors.
Attention, shared experts, PLE, vision components, MTP, and other non-routed BF16 tensors remain unchanged. V2 performs no FP8 PLE/non-expert conversion, no MTP refresh, and no new fine-tuning or behavioral transformation.
At release time, no V2 inference or capability benchmark had been run. A subsequent October 4–5, 2026 agent evaluation exercised this exact V2 repository through the Cyber-Frost Harness, as documented below. Saved-scale and tensor-metadata checks remain distinct from inference measurements. Historical GB10 performance results retained below still describe the predecessor, not V2.
Cyber-Frost Harness evaluation
Cyber-Frost Harness is the public red/blue/purple runtime built around the Cyber-Frost family. It supplies eight procedural security skills, structured native-analysis tools, durable evidence handling, token-aware context management, and an isolated x86-64 vulnerable-image environment.
The measured run used this exact CYBER-FROST-3.8-NVFP4-V2 artifact with SGLang 0.5.20 across eight NVIDIA B200 GPUs in DP2 × TP4, thinking enabled and preserved, xhigh reasoning, temperature 1.0, top-p 0.95, a 131,072-token completion ceiling, and seed 38421.
| Experiment | Scaffold | Result | Disclosure |
|---|---|---|---|
| Earlier frozen four-task slice | OpenHands plus native grader | 4/4 | Custom held-out CyberGym-style slice; not an official leaderboard run |
| Level-0 hard-12 baseline | Generic OpenHands | 1/12 | Failure-analysis baseline used to design the harness |
| Two-task intervention pilot | Cyber-Frost Harness | 2/2 | Post-hoc prior failures; validates mechanisms, not generalization |
| Dynamic-harness hard-12 | Cyber-Frost Harness | 5/10 valid attempts | Five verified solves, five model misses, two infrastructure-invalid tasks; 5/12 verified lower bound |
The generic hard-12 baseline used 51,939,855 total model tokens. The dynamic harness produced five hidden-fixed verified solves using 22,043,450 tokens, approximately 42.4% of baseline usage. The baseline and harness used materially different scaffolds, so this is an intervention comparison rather than an apples-to-apples model-only benchmark. The harness changes runtime access, tools, skills, context handling, and artifact preservation; it does not change model weights.
GPAC and GDAL were infrastructure-invalid after the original client paired large histories with the full completion reservation and exceeded the 262,144-token context window. They were not rerun before the native host became unavailable. No final 12-task percentage is claimed.
- Source and quick start
- Architecture and trust boundaries
- Full evaluation disclosure
- Publication-safe result records
Release status and contents
Cyber-Frost is a public research release under active quality assessment.
This repository contains a standalone mixed-precision checkpoint derived from Blackfrost-AI/CYBER-FROST-3.8-BF16, plus its weight index, configuration, ModelOpt quantization metadata, tokenizer and processor assets, packaged chat template, and upstream license. It is not an adapter and does not require the BF16 parent checkpoint at load time.
| Field | Released artifact |
|---|---|
| Clean model name | CYBER-FROST-3.8-NVFP4-V2 |
| Predecessor | CYBER-FROST-3.8-NVFP4, formerly BLACKFROST-3.8-ICED-NVFP4-W4A4 |
| Architecture | Qwen4ExpForConditionalGeneration |
| Precision | Mixed: routed language-model experts use ModelOpt NVFP4 W4A4, group size 16; non-target tensors remain BF16 |
| Weight layout | 131 SafeTensors shards |
| Logical source architecture | approximately 180B parameters |
| Indexed tensor payload | 186,356,367,352 bytes (173.56 GiB) |
| Configured context | 262,144 tokens |
| Native speculative head | one BF16 MTP layer preserved from the BF16 source |
| Validated modality | text in the predecessor; V2 inference qualification remains pending |
The configuration includes a vision tower, but this release has not received a multimodal quality evaluation. Do not infer validated image or video capability from the presence of processor files.
Why Cyber-Frost exists
Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity; a general-purpose assistant can react to individual terms instead of the operator's legitimate scope.
Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. This is a design objective, not a measured refusal claim for this NVFP4 variant, and not a claim that every answer is safe or correct.
Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model.
Security corpus
The BF16 parent was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published.
Domain coverage includes:
- Reconnaissance and OSINT
- Social engineering, business-email compromise, and deepfake-enabled abuse
- Web application and API security
- Identity, authentication, and Active Directory security
- Network, perimeter, VPN, and protocol security
- Vulnerability research, bug bounty, and binary exploitation
- Malware analysis, ransomware, and endpoint defense
- Cloud, container, and Kubernetes security
- Software supply-chain security
- Mobile, IoT, wireless, and physical security
- Industrial-control-system and operational-technology security
- Cryptography and security protocols
- Privilege escalation, lateral movement, and data exfiltration
- Threat intelligence, APT analysis, and purple-team operations
- AI-agent, LLM, and adversarial-ML security
Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The release evidence independently binds one security subset to a Qwen3.8 2.4T teacher; it does not include a corpus-wide teacher manifest. These are therefore operator provenance statements, not independent benchmark findings.
Training-data provenance and licensing review for the mixed-source corpus remains in progress. The NVFP4 conversion added no new fine-tuning data; this section describes the BF16 parent inherited by the quantized artifact.
Model specifications and precision layout
The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Its hidden size is 2,560 with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. One native MTP layer is packaged for speculative decoding.
This is a mixed-precision ModelOpt checkpoint, not an all-tensor four-bit conversion. The routed expert blocks in all 48 language layers are configured for NVFP4 weights and input activations with group size 16. The 73,728 routed expert projection matrices use packed four-bit weights, FP8 E4M3 block scales, FP32 outer scales, and FP32 input scales. V2 uses a shared FP32 outer scale for each expert's gate/up pair, with block scales recomputed against that shared value before packing.
Attention and linear-attention state, embeddings, normalization tensors, routers, shared experts, the PLE table, vision components, and the MTP layer are excluded from NVFP4 conversion and remain BF16. KV-cache precision is chosen by the serving runtime and is not encoded in the checkpoint weights.
The BF16 PLE table is large. Memory-constrained systems may require a runtime with CPU or file-backed PLE handling. The 262,144-token configuration ceiling is not a blanket quality guarantee; NVFP4-specific long-context, high-concurrency, multimodal, and tool-heavy agent-loop qualification remains pending.
Lineage
- Foundational checkpoint:
Qwen/Qwen3.8-Flash-Nextat immutable revisionde4b8e4d43b917e7706784d8bb445c9af86a3540. - Blackfrost security adaptation: security-domain fine-tuning followed by a full BF16 merge. The merged internal stage was identified as
BLACKFROST-3.8-FLASH-BF16. - Behavioral stage: a Blackfrost-AI behaviorally modified derivative targeting lower false-refusal friction in authorized security workflows. The proprietary transformation process is not distributed.
- BF16 conversion source: the payload now published as
Blackfrost-AI/CYBER-FROST-3.8-BF16, at immutable source revision5321904427c4ef54df8a667edcbc2d1184e4286e. - Original NVFP4 conversion: the routed language-model expert projections were converted to ModelOpt NVFP4 W4A4 with group size 16. Excluded tensors were preserved in BF16. The conversion used ModelOpt commit
022767c7ab3d7d36211affd85e5c496770cde768; its installed package reports version0.47.0rc0. - Predecessor release identity: that checkpoint was formerly labeled
BLACKFROST-3.8-ICED-NVFP4-W4A4and was renamedCYBER-FROST-3.8-NVFP4. That rename was not another conversion or training run. - V2 corrected export, 2026-10-02: the same final BF16 source was freshly re-exported using the same ModelOpt codec, with per-expert shared gate/up global maxima established before new block-scale calculation and NVFP4 packing. This repository is
Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2.
The quantization process used calibrated input-scale metadata from nvidia/Qwen3.8-Flash-Next-NVFP4 at revision fc694b54fb0174e0913e6adf86691ef85a4ead47. That checkpoint is a conversion reference, not the behavioral or weight lineage of Cyber-Frost. V2 inherits those calibrated input-scale values unchanged through the existing Cyber-Frost NVFP4 reference.
Tokenizer and processor lineage comes through the pinned Qwen foundation and BF16 source. The tokenizer, processor, generation configuration, license, and packaged Blackfrost chat template are byte-identical to the current BF16 release assets.
Artifact verification
V2 saved-artifact validation
V2 validation completed on 2026-10-02. Saved-file scalar checks confirmed bit-identical global scales for all 24,576 gate/up pairs and bit-identical preservation of all 73,728 calibrated reference input scales. Saved-file metadata checks confirmed matching shapes and dtypes for all 1,562 non-routed source tensors. The V2 weight index covers 131 shards and 296,474 tensors. No full weight hash sweep or inference benchmark was performed as part of that export-validation stage; the later agent evaluation is reported separately above and below.
The bundled saved-scale verification, export provenance, and quality report describe the corrected export and its conversion-time checks, including sampled BF16 reconstruction. These checks establish export properties; they are not refusal, capability, or runtime-performance measurements.
Historical predecessor verification
The following hashes and reconstruction results are retained as historical provenance for CYBER-FROST-3.8-NVFP4, not as V2 hashes or V2 reconstruction results.
The predecessor's conversion and integrity validation completed on 2026-09-17. Its weight and configuration payload was re-audited at NVFP4 repository revision 69077d67a6274b6f7f6c6714b28f13da265bfd12; the predecessor model-card update did not alter that payload. V2 does alter the quantized expert payload and associated metadata through the fresh export described above.
| Historical predecessor artifact | SHA-256 |
|---|---|
config.json |
91d03d21273f4761e5a95c8f0c879eec7207af1d0eae8739596e98755fd8557c |
hf_quant_config.json |
8d6c3fb3ac2cfc6f6f9e494ec6c17a6a61c7caba43b9f5b0332b8035879d17ec |
model.safetensors.index.json |
fc68c1d90e62460c4757a515619d8a533a43156581ac19469424dac81a217588 |
| Qwen Community License file | a0dc422560841fd68e06d974907f8b4c709bca44a67daad2b528437bdf676c08 |
The predecessor validation indexed all 131 shards and 296,474 tensors. It compared 1,562 tensors intentionally preserved from the BF16 source and found them exact. A 144-matrix sample of the predecessor's converted expert weights had a mean relative Frobenius error of approximately 0.0949 and a maximum of approximately 0.0952 against the BF16 source.
The preserved MTP tensors predate the final trunk-weight behavioral stage. Treat the MTP head as a provisional acceleration baseline rather than a freshly adapted draft head. V2 does not refresh it.
Evaluation status
Agentic security evaluation
The October 4–5 Cyber-Frost Harness evaluation is the first published V2 inference evidence. It measured vulnerability-analysis agent performance on custom CyberGym-style slices and is fully disclosed in the harness repository. The frozen earlier four scored 4/4 under an OpenHands-based scaffold. On the distinct Level-0 hard-12 slice, the generic scaffold scored 1/12; the dynamic harness later produced five verified solves, five valid model misses, and two infrastructure-invalid tasks. Report this as 5/10 valid attempts or a 5/12 verified lower bound, not as a final 12-task percentage.
No NVFP4-specific refusal or over-refusal result is published yet. No standardized CyberMetric, SecBench, MMLU, HumanEval, or IFEval score is claimed for this exact checkpoint. The custom CyberGym-style results are not an official leaderboard submission, and refusal behavior is not a substitute for measuring security competence.
Legacy single-GB10 development measurement
Historical predecessor results only: these measurements were not rerun for V2.
On 2026-09-17, the predecessor NVFP4 payload was exercised on one NVIDIA GB10 in a DGX Spark with tensor parallelism 1, a patched vLLM development build (0.1.dev20073+g8e685d198), FP8 KV cache, native MTP with three speculative steps, and a 262,144-token configured serving context. The benchmark used warm single-stream requests with fixed 256-token outputs and two measured runs per category.
| Category | Mean decode speed | Mean end-to-end speed | Mean TTFT |
|---|---|---|---|
| Prose | 26.40 tok/s | 25.32 tok/s | 453 ms |
| Code | 29.29 tok/s | 28.08 tok/s | 413 ms |
These are narrow development measurements, not production throughput or quality guarantees. They do not predict V2, RTX, B300, multi-GPU, concurrent, or long-context performance. Runtime version, prompt length, sampling, context occupancy, cache precision, PLE handling, and MTP acceptance can materially change the result.
Prompt, tool use, and sampling
The repository includes a Blackfrost chat template that supplies a default operating prompt, supports caller-provided system context, exposes Qwen-style reasoning controls, and serializes tool calls. Its default is thinking enabled; supported reasoning-effort values are xhigh, medium, and low. Generation defaults are temperature 1.0, top-p 0.95, and top-k 20.
The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it.
Changing the template, system message, reasoning mode, sampling, quantization backend, or runtime can materially change refusal behavior and output quality. Record those settings when reporting results.
Deployment
Use a runtime and hardware path with explicit support for the Qwen4ExpForConditionalGeneration architecture and mixed ModelOpt NVFP4 checkpoints. The validated legacy GB10 path required a patched runtime and file-backed PLE offload; the BF16 deployment kit is not interchangeable with this artifact. That predecessor qualification is not a V2 deployment benchmark.
Runtime compatibility, kernel selection, memory use, throughput, and output quality are implementation-dependent. This repository does not yet include a generally qualified deployment kit. Record the runtime version, hardware, context, quantization backend, PLE strategy, speculative settings, chat template, and sampling parameters when reporting results. The clean API model identifier is CYBER-FROST-3.8-NVFP4-V2, and the Hugging Face repository identifier is Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2.
Intended use
Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows.
It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm.
Limitations and security responsibility
- Generated findings, code, commands, indicators, and remediation advice may be wrong, incomplete, outdated, or fabricated. Independently review them and execute only in isolated, authorized environments.
- Reduced over-refusal is a design objective inherited from the BF16 parent, not a measured result for this NVFP4 artifact. Lower-friction behavior can increase the chance of receiving actionable output in ambiguous or malicious contexts.
- BF16 behavior, safety observations, benchmark results, and runtime characteristics do not automatically transfer through quantization.
- The model is not a policy engine, authorization service, sandbox, malware scanner, or secrets boundary.
- The preserved MTP head is provisional and is not evidence of adaptation to the final trunk.
- Current evidence does not establish an NVFP4 refusal rate, standardized cyber competence, production readiness, long-context quality, or multimodal quality.
- Model behavior can shift substantially with prompts, sampling, runtime versions, quantization kernels, speculative settings, and agent scaffolding.
The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law.
License and disclaimer
Use and redistribution of this checkpoint are governed by the Qwen Community License 1.0. Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement.
Report reproducible model or packaging issues through the repository's Discussions page without including secrets, client data, live targets, or sensitive exploit details.
- Downloads last month
- 677
Model tree for Blackfrost-AI/CYBER-FROST-3.8-NVFP4-V2
Base model
Qwen/Qwen3.8-Flash-Next
