Agens Volundr 32B Preview โ DFlash2 drafter
Main model card, benchmarks and discussion: Blockway/Agens-Volundr-32B-Preview
A DFlash2 block drafter for speculative decoding with Agens Volundr 32B Preview. At each step it drafts 8 tokens at once from the target model's hidden states (taps at layers 8, 21, 35, 53 and 67), and the target verifies them in one pass. The target checks every token, so answers keep the target model's quality; with greedy decoding the output can differ from non-speculative decoding only where two tokens are numerically tied.
The drafter has 5 layers (about 1.9B parameters, 3.8 GB in bf16) and shares the target's embeddings and output head, which are not stored here.
Speed-up
Single user on two 48 GB GPUs, bf16, the Blockway sglang build, same server with and without the drafter:
| Output | Speed-up |
|---|---|
| JSON | 3.6ร |
| Code | 2.0ร |
| Replies with thinking on (code, Cantonese) | 1.6ร |
| Cantonese chat | 1.3ร |
On average 2.56 drafted tokens are accepted per step. The drafter is for single-user generation: the drafter server runs 4 requests at a time, so with 8 or more concurrent users plain decoding gives more total throughput.
Use
Requires the Blockway sglang build (ghcr.io/blockwayz/agens-sglang, source at
github.com/BlockWayz/agens-sglang). Add these flags to the two-GPU bf16 command from the
model card, and run the container with
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -e VOLUNDR_FI_WORKSPACE_MB=256. The larger FlashInfer workspace is needed on GPUs with many SMs, especially at tensor-parallel size 1 (the 2026-10-05 images default to 128 MB and can fail during CUDA-graph capture with aligned_alloc ... but only 134217728 bytes available):
--speculative-algorithm DFLASH \
--speculative-draft-model-path Blockway/Agens-Volundr-32B-Preview-DFlash2 \
--speculative-num-draft-tokens 8 \
--speculative-draft-model-quantization unquant \
--mem-fraction-static 0.90 --context-length 16384 --chunked-prefill-size 2048 \
--max-mamba-cache-size 4 --max-running-requests 4 --cuda-graph-max-bs 4
Speculative decoding needs --attention-backend flashinfer and serves text requests (leave out
--enable-multimodal).
Licence
Apache-2.0. The licence files and NOTICE ship with the weights.
- Downloads last month
- 564