Agens Volundr 32B Preview โ€” DFlash2 drafter

Main model card, benchmarks and discussion: Blockway/Agens-Volundr-32B-Preview

A DFlash2 block drafter for speculative decoding with Agens Volundr 32B Preview. At each step it drafts 8 tokens at once from the target model's hidden states (taps at layers 8, 21, 35, 53 and 67), and the target verifies them in one pass. The target checks every token, so answers keep the target model's quality; with greedy decoding the output can differ from non-speculative decoding only where two tokens are numerically tied.

The drafter has 5 layers (about 1.9B parameters, 3.8 GB in bf16) and shares the target's embeddings and output head, which are not stored here.

Speed-up

Single user on two 48 GB GPUs, bf16, the Blockway sglang build, same server with and without the drafter:

Output Speed-up
JSON 3.6ร—
Code 2.0ร—
Replies with thinking on (code, Cantonese) 1.6ร—
Cantonese chat 1.3ร—

On average 2.56 drafted tokens are accepted per step. The drafter is for single-user generation: the drafter server runs 4 requests at a time, so with 8 or more concurrent users plain decoding gives more total throughput.

Use

Requires the Blockway sglang build (ghcr.io/blockwayz/agens-sglang, source at github.com/BlockWayz/agens-sglang). Add these flags to the two-GPU bf16 command from the model card, and run the container with -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e VOLUNDR_FI_WORKSPACE_MB=256. The larger FlashInfer workspace is needed on GPUs with many SMs, especially at tensor-parallel size 1 (the 2026-10-05 images default to 128 MB and can fail during CUDA-graph capture with aligned_alloc ... but only 134217728 bytes available):

    --speculative-algorithm DFLASH \
    --speculative-draft-model-path Blockway/Agens-Volundr-32B-Preview-DFlash2 \
    --speculative-num-draft-tokens 8 \
    --speculative-draft-model-quantization unquant \
    --mem-fraction-static 0.90 --context-length 16384 --chunked-prefill-size 2048 \
    --max-mamba-cache-size 4 --max-running-requests 4 --cuda-graph-max-bs 4

Speculative decoding needs --attention-backend flashinfer and serves text requests (leave out --enable-multimodal).

Licence

Apache-2.0. The licence files and NOTICE ship with the weights.

Downloads last month
564
Safetensors
Model size
2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support