Agens Volundr 32B Preview — GGUF

Main model card, benchmarks and discussion: Blockway/Agens-Volundr-32B-Preview

GGUF files of Agens Volundr 32B Preview for llama.cpp: BF16, Q8_0 and Q4_K_M. Text only. For what the model is, its benchmarks and its chat format, see the main model card.

Needs the volundr branch of llama.cpp: github.com/BlockWayz/llama.cpp. Volundr is a new architecture (KDA linear attention, BCSA sparse attention, Engram memory, mHC hyper-connections). Stock llama.cpp, and apps built on it, will refuse these files with unknown model architecture: 'volundr' until support is merged upstream.

Files

file size contents
Agens-Volundr-32B-Preview-Q4_K_M.gguf 26.6 GB Q4_K_M for the feed-forward and BCSA/dense attention projections; every KDA projection, the Engram value/gate projections, the mHC dynamic projections, the token embedding and the output layer at Q8_0; Engram tables and BCSA indexer BF16. Recommended.
Agens-Volundr-32B-Preview-Q8_0.gguf 36.3 GB all matrices Q8_0; Engram tables and BCSA indexer BF16
Agens-Volundr-32B-Preview-BF16.gguf 64.5 GB reference: the text model's weights unquantised (BF16)

Norms, convolution, decay and mHC static parameters are F32 in every file. The linear-attention (KDA) weights are kept at Q8_0 in the 4-bit file on purpose: an earlier model of ours degraded badly when its linear-attention projections were quantised to 4 bit. SHA256SUMS lists the checksums.

Build llama.cpp (volundr branch)

git clone --branch volundr https://github.com/BlockWayz/llama.cpp && cd llama.cpp

# CPU
cmake -B build -DCMAKE_BUILD_TYPE=Release
# or NVIDIA GPU
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON

cmake --build build -j --target llama-server llama-cli

If CMake cannot find libcurl, add -DLLAMA_CURL=OFF. More detail in README-volundr.md.

Run

pip install -U huggingface_hub
hf download Blockway/Agens-Volundr-32B-Preview-GGUF Agens-Volundr-32B-Preview-Q4_K_M.gguf --local-dir .

# OpenAI-compatible server on :8080 (thinking, tool calls and reasoning_content parsing work out of the box)
./build/bin/llama-server -m Agens-Volundr-32B-Preview-Q4_K_M.gguf --jinja -c 32768 -ngl 99

# interactive chat in the terminal
./build/bin/llama-cli -m Agens-Volundr-32B-Preview-Q4_K_M.gguf --jinja -c 32768 -ngl 99

Drop -ngl 99 to run on the CPU only. When offloading to a GPU, the four Engram tables (4.3 GB of pure lookups) can stay in host RAM: --override-tensor "engram_embd=CPU".

Thinking is on by default. Per request:

curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "messages": [{"role": "user", "content": "What is 17 * 23 + 9?"}],
  "chat_template_kwargs": {"reasoning_effort": "low"}
}'
# "chat_template_kwargs": {"enable_thinking": false}      -> no thinking
# "chat_template_kwargs": {"reasoning_effort": "medium"}   -> low | medium | xhigh (default xhigh)

Tool calls use the standard OpenAI tools field and come back as tool_calls. Sampling defaults stored in the GGUF: temperature 1.0, top-k 20, top-p 0.95. The chat template is the model's own and is embedded in every file.

Speed

Q4_K_M on one 48 GB GPU, full offload (-ngl 99), flash attention, F16 KV cache:

context 1K 8K 32K
decode 32.3 tok/s 27.4 tok/s 25.0 tok/s

Prompt processing of a 1K-token prompt: 2,218 tok/s.

Quantisation quality

llama-perplexity on held-out chat text (mostly Cantonese and Chinese), 12 chunks of 4,096 tokens; KLD and top-1 are measured against the BF16 file.

file PPL vs BF16 mean KLD same top-1
BF16 3.3916 — — —
Q8_0 3.3922 +0.02 % 0.00056 99.10 %
Q4_K_M 3.4145 +0.68 % 0.0100 96.00 %

With an 8,192-token context scored on positions 4,096–8,191 (where the BCSA far field is active) the picture is the same: 3.1915 / 3.1886 / 3.2097.

Parity with the reference implementation

The llama.cpp graph was checked against the checkpoint's own PyTorch code (fp32) on all-position logits. The BF16 file agrees on 99.7–100 % of top-1 tokens (mean KL ≤ 5e-5) on prompts from 90 to 6,600 tokens in English, Cantonese, code and chat format, and 64-token greedy continuations match except where the reference itself has a near-exact tie. An F32 conversion matches the reference to fp32 rounding (100 % top-1 on a 6,600-token prompt), including BCSA's far-field block selection. CPU and CUDA builds give the same greedy continuations on the chat prompts tested.

Unused-token mask

The output layer has 248,320 rows but the tokenizer ends at id 248,076. The ids 248,077–248,319 are never valid output, yet the model can put real probability on some of them: where a tool call opens, sampling sometimes picks one instead of the tool-call token and the call comes back as plain text. The volundr branch therefore masks that range (a -inf logit bias, applied automatically for this architecture in llama-server and llama-cli); the main repository lists the same ids in generation_config.json suppress_tokens. Sampled tool calls with the Q4_K_M file:

with mask without mask
no thinking, temperature 0.6 16/16 7/16
no thinking, temperature 1.0 (default) 16/16 10/16
thinking, temperature 1.0 8/8 3/8

VOLUNDR_MASK_UNUSED_TOKENS=0 turns the mask off. If you call the llama.cpp library directly (bindings that do not use llama.cpp's common sampling code), apply the same bias yourself.

Known limits

  • Text only. The vision tower is not converted; use the main repository for image input.
  • No speculative decoding in llama.cpp.
  • CUDA and CPU only for the two Volundr-specific ops (Engram hashing, mHC Sinkhorn). On Metal, Vulkan or ROCm the scheduler should run those two ops on the CPU; those backends have not been tested.
  • BCSA prefill cost. Prompt processing scores every query against every cached position before pooling and selecting blocks, so it slows down with prompt length (about 620 tok/s for a 32K-token prompt on the same GPU). Decoding switches to a sparse path above 12,288 cached positions.
  • No context shift and no partial KV-cache removal (recurrent state), as for llama.cpp's other hybrid models.
  • Multi-GPU --split-mode layer works but is no faster than one 48 GB GPU; --split-mode row does not load.

Licence

Apache-2.0. These files are a format conversion of the weights in Blockway/Agens-Volundr-32B-Preview; LICENSE, NOTICE and the files they reference apply.

Downloads last month
664
GGUF
Model size
32B params
Architecture
volundr
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blockway/Agens-Volundr-32B-Preview-GGUF

Quantized
(1)
this model