RED-SNOW 5.3 FLASH — Authorized Red Team Operations

RED-SNOW 5.3 FLASH — NVFP4

Red-Team Operations Model

RED-SNOW 5.3 FLASH is Blackfrost AI's focused red-team operations model. It is built for authorized offensive-security work: turning an objective into a technically grounded attack path, reasoning across identity, infrastructure, applications, endpoints, cloud, and operational technology, and translating findings into defensive action.

This is not a generic assistant with a security prompt layered on top. RED-SNOW combines a GLM-5.3-Flash foundation with targeted security adaptation and Blackfrost's internal post-training release process. The result is intended to operate like a practical red-team partner across discovery, exploitation analysis, adversary emulation, validation, and purple-team handoff.

Release status: public, manually gated NVFP4 deployment weights. Refusal-behavior and RED-SNOW-specific harness evaluations remain pending and will be added when complete.

Release Family

Format Repository
BF16 reference RED-SNOW-5.3-FLASH-BF16
FP8 deployment RED-SNOW-5.3-FLASH-FP8
NVFP4 deployment RED-SNOW-5.3-FLASH-NVFP4

Operational Profile

Operational strength Red-team impact
Attack-path reasoning Connects isolated weaknesses into realistic routes from initial access to material impact.
Cross-domain operations Reasons across identity, cloud, network, application, endpoint, mobile, wireless, physical, and OT boundaries.
Exploit and tradecraft analysis Supports vulnerability analysis, exploitability assessment, payload logic, post-exploitation planning, and controlled validation.
Detection-aware execution Accounts for EDR, logging, telemetry, and likely defender visibility when designing and reviewing exercises.
Tool-oriented workflows Produces structured commands, plans, artifacts, and tool calls for operator-controlled environments.
Purple-team conversion Converts offensive observations into detections, mitigations, hardening priorities, and retest criteria.

Red-Team Domain Coverage

The finalized training split contains 17,132 answer-only examples. 14,707 examples (85.8%) come from the four core offensive-security source families; the remaining operator-owned material reinforces reasoning, engineering, analysis, and reporting. The following domain map reflects the security categories actually present in the finalized training data.

1. Cloud Attack Paths and Control-Plane Security

Coverage includes cloud identity, permissions, exposed services, secrets, storage, workload boundaries, infrastructure-as-code, control-plane abuse, and paths from a foothold to broader tenant impact.

Training labels represented: cloud (3,092), cloud security posture, cloud infrastructure, infrastructure-as-code, container orchestration, configuration and secrets management.

2. Active Directory, Identity, and Authentication

Coverage includes enterprise identity mapping, credential attack paths, authentication weaknesses, privilege relationships, lateral movement, trust boundaries, IAM, and escalation across Windows and Active Directory estates.

Training labels represented: active_directory (2,640), Identity, Credentials & Authentication Attacks (90), Privilege Escalation: Windows, Linux, macOS, Active Directory (7), IAM and zero-trust.

3. ICS/OT and High-Consequence Environments

Coverage includes industrial networks, control-system exposure, segmented environments, protocol risk, safety-aware attack-path analysis, and validation strategies for operational technology without treating production systems like ordinary IT.

Training labels represented: ics_ot (1,760).

4. Malware, EDR, Ransomware, and Endpoint Tradecraft

Coverage includes malware behavior, trojans, backdoors, RATs, endpoint controls, EDR-aware reasoning, persistence, command and control, ransomware operations, extortion patterns, and detection-oriented tradecraft review.

Training labels represented: malware_edr (1,332), Malware: Trojans, Backdoors, RATs (90), Ransomware & Extortion Operations (94).

5. Cryptography, Secrets, and Trust Failures

Coverage includes cryptographic misuse, key and secret handling, protocol assumptions, token and signing weaknesses, certificate trust, and the operational consequences of broken confidentiality or authenticity.

Training labels represented: cryptography (959) and Kali cryptography/miscellaneous operations.

6. Web Applications, APIs, and Bug-Bounty Workflows

Coverage includes web and API attack surfaces, authentication and authorization failures, input handling, business logic, chained application flaws, exploit validation, impact demonstration, and report-quality reproduction steps.

Training labels represented: web_api (842), Web Application & API Attacks (92), bug_bounty (71), Kali web exploitation.

7. Binary Exploitation and Vulnerability Research

Coverage includes memory-safety reasoning, crash and root-cause analysis, exploitability assessment, mitigations, reverse-engineering workflows, proof-of-concept construction, and controlled zero-day validation.

Training labels represented: binary_exploitation (824), Perimeter, VPN & Zero-Day Exploitation (94), exploit reasoning, payload construction, reverse shells, and privilege-escalation workflows.

8. Network Infrastructure and Protocol Attacks

Coverage includes service discovery, protocol behavior, perimeter exposure, VPNs, segmentation, routing and trust assumptions, infrastructure misconfiguration, pivoting, and validation of network-layer controls.

Training labels represented: network_infra (653), Network-Layer & Protocol Attacks (95), Kali reconnaissance/enumeration and post-exploitation pivoting.

9. Reconnaissance, OSINT, and Social Engineering

Coverage includes attack-surface discovery, target research, relationship mapping, pretext analysis, email and BEC scenarios, deepfake-enabled threats, human-layer exposure, and defensible exercise design.

Training labels represented: osint_social_engineering (375), Reconnaissance & OSINT (97), Email, BEC & Deepfake Social Engineering (91).

10. Threat Intelligence and Purple-Team Operations

Coverage includes adversary behavior mapping, hypothesis-driven testing, IOC and TTP analysis, detection validation, SOC handoff, threat hunting, incident response, DFIR, SIEM/log analysis, and translating findings into measurable defensive improvements.

Training labels represented: threat_intel_purple_team (343), threat intelligence, threat hunting, detection engineering, SOC triage, incident response, DFIR/forensics, SIEM/log analysis, defensive detection, vulnerability management, and security hardening.

11. Mobile, Wireless, and Physical Attack Surfaces

Coverage includes mobile application and platform risk, device trust, wireless exposure, proximity-dependent weaknesses, and attack paths that cross digital and physical boundaries.

Training labels represented: mobile (210) and wireless_physical (73).

12. Software Supply Chain and AI-System Attacks

Coverage includes dependency and build-pipeline compromise, CI/CD and artifact trust, model and training-data attacks, adversarial ML, LLM and agent attack surfaces, prompt/context manipulation, retrieval boundaries, and tool-using agent risk.

Training labels represented: Software Supply Chain Attacks (96), Adversarial ML, Model & Training-Data Attacks (95), AI/LLM & Agent Attacks (90), plus AI architecture, model engineering, agentic tool use, build tooling, and CI/CD.

13. Kali and Operator Execution Workflows

Coverage includes reconnaissance and enumeration, command construction, exploit selection and reasoning, web exploitation, privilege escalation, payload and reverse-shell construction, post-exploitation pivoting, tool-assisted execution, and detection-aware review.

Training labels represented: security-kali (451), security-kali-safe (93), reconnaissance/enumeration, command construction, web exploitation, privilege escalation, exploit reasoning, post-exploitation pivoting, reverse-shell payloads, defensive detection, cryptography/miscellaneous, and tool-agentic operations.

Supporting Operator Disciplines

The supporting corpus is deliberately broader than exploit syntax. It reinforces the disciplines needed to complete an engagement and communicate its impact:

  • Software and systems engineering: Python, Go, JavaScript/TypeScript, C/C++, Rust, Java, SQL, debugging, testing, distributed systems, APIs, databases, performance, repositories, build tooling, and DevOps/SRE.
  • Data and analytical reasoning: data analysis, statistics, experimentation, anomaly detection, forecasting, causal reasoning, quantitative methods, optimization, and structured evidence review.
  • AI and automation: model architecture, ML application engineering, retrieval, fine-tuning, evaluation, deployment, monitoring, tool use, and agentic workflows.
  • Security operations and governance: incident management, resilience, privacy, compliance controls, risk analysis, policy, and control validation.
  • Organizational context: government contracting, legal and regulatory analysis, finance, procurement, supply chain, logistics, operations, product, sales, marketing, HR, insurance, and real estate. These domains support impact analysis and executive-relevant reporting rather than replacing specialist advice.
  • Deliverable quality: Markdown, schemas, tables, reports, citations, format conversion, multi-item consistency, and concise technical documentation.

Model Construction

Component RED-SNOW configuration
Foundation Blackfrost-AI/RED-SNOW-5.3-FLASH-BF16, derived from zai-org/GLM-5.3-Flash-BF16
Security adaptation BF16 LoRA, rank 16 / alpha 32, 529 target matrices
Training format Single-turn user/assistant pairs with assistant-only loss
Finalized data 17,132 train / 1,793 validation examples
Prompt policy Training system messages and conversation carryover removed; the release chat template includes the RED-SNOW operating prompt supplied by Blackfrost AI while continuing to support client system and developer messages
Reasoning hygiene Leading chain-of-thought blocks removed from 2,293 training and 116 validation answers; five pure chain-of-thought training rows dropped
Checkpoint selection Best held-out checkpoint at step 204, validation loss 0.856393520659158
Precision ModelOpt NVFP4 weights and activations, group size 16; FP8 KV-cache metadata
Quantizer NVIDIA ModelOpt 0.48.0.dev0+g6a1068746

The run was scheduled for three passes and stopped by the operator during pass three. RED-SNOW uses the best held-out checkpoint from step 204 rather than the final saved checkpoint. This choice prioritizes measured validation quality over training duration.

Evaluation Status

The current release candidate passed the matched internal coherence gate at:

  • 16/18 strict-format checks
  • 18/18 correctness-adjusted checks
  • 2/2 context checks
  • 2/2 forced structured-tool checks

This is a small internal construction gate, not a claim of comprehensive cyber capability or real-world exploit success. Dedicated security evaluations should be reported separately with their environments, harness versions, task lists, and grading rules.

Evaluation Status
Refusal-behavior test Pending
RED-SNOW Cyber-Frost Harness evaluation Pending
Public security benchmark report Pending

No refusal-rate claim is being made until the dedicated refusal test is complete.

Chat-template release note

The packaged template preserves GLM's visible-answer/reasoning split, supports thinking-on and thinking-off generation, and accepts prior tool calls whose arguments arrive either as mappings or as OpenAI-compatible JSON strings. The root template, deployment copy, and tokenizer-embedded copy are synchronized. This corrects the earlier JSON-string replay failure without changing model weights.

Cyber-Frost Harness Deployment Kit

This repository is prepared to include the Cyber-Frost Harness as a companion runtime for RED-SNOW. The harness does not alter or merge into the model weights. It sits beside an OpenAI-compatible inference server and turns model tool calls into bounded red-, blue-, and purple-team workflows with procedural skills, structured security tools, durable evidence, and isolated native execution.

The deployment kit includes:

  • a validated SGLang launch pattern for RED-SNOW BF16 and NVFP4;
  • the complete Cyber-Frost Harness source and its eight procedural security skills;
  • a RED-SNOW harness configuration template;
  • an inference-engine compatibility contract and porting checklist;
  • service examples, health checks, tool-protocol probes, and task-run instructions;
  • operational boundaries for isolated Docker targets and SSH Docker workers.

Start with Deployment-Kit/RUNBOOK.md. The bundled harness documentation is at Deployment-Kit/Cyber-Frost-Harness/README.md.

Historical Cyber-Frost Harness results belong to the harness project and its documented CYBER-FROST experiments. They are not presented as RED-SNOW scores. RED-SNOW-specific harness evaluation remains pending.

Intended Use

RED-SNOW is intended for experienced operators working in environments they own or are explicitly authorized to assess, including:

  • scoped penetration tests and red-team engagements;
  • adversary emulation and purple-team exercises;
  • cyber ranges, labs, CTFs, and security research;
  • vulnerability triage and controlled exploit validation;
  • detection engineering, threat hunting, and defensive control testing;
  • attack-path review, remediation planning, and technical reporting;
  • security-tool orchestration with human approval and auditable execution.

Operating Guidance

  • Keep a human operator in control of scope, target selection, credentials, tool execution, and destructive actions.
  • Provide the model with explicit engagement context, available tools, target boundaries, success criteria, and evidence requirements.
  • Validate commands and exploit assumptions in an isolated environment before use on production systems.
  • Preserve logs and artifacts so conclusions can be reproduced and reviewed.
  • Treat generated findings as hypotheses until verified against the actual system.

Limitations

  • The model can produce incorrect commands, invalid exploit assumptions, stale technique details, or incomplete remediation guidance.
  • Domain coverage does not guarantee equal capability across every product, version, architecture, or security control.
  • Tool calling depends on the serving stack, chat template, parser, and schema presented at runtime.
  • Vision and audio towers were frozen during the security adaptation; multimodal security performance has not been established by this card.
  • The internal coherence gate is not a substitute for held-out security benchmarks or operator evaluation.
  • Long-context capability, latency, and throughput depend on the deployed precision, runtime, parallelism, KV-cache configuration, and hardware.

Data and Release Notes

The corpus build retained accepted, non-empty records and excluded malformed carryover, known refusal-quarantine items, unresolved collection failures, superseded source versions, benchmark material reserved against contamination, and artifacts without usable raw-text provenance. Training source data is not distributed in this repository. The model artifact is released under the included MIT license; original source materials remain subject to their respective terms.

Lineage

zai-org/GLM-5.3-Flash-BF16
  -> focused security LoRA
  -> best held-out checkpoint (step 204)
  -> BF16 merge
  -> release validation and chat-template integration
  -> RED-SNOW 5.3 FLASH BF16
  -> ModelOpt NVFP4 export
  -> RED-SNOW 5.3 FLASH NVFP4

Status

Public, manually gated NVFP4 release.

Blackfrost AI will add completed refusal, harness, and security-benchmark results after review. The included deployment kit documents SGLang serving and Cyber-Frost Harness integration.

Downloads last month
9
Safetensors
Model size
161B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-AI/RED-SNOW-5.3-FLASH-NVFP4

Quantized
(5)
this model

Collection including Blackfrost-AI/RED-SNOW-5.3-FLASH-NVFP4