Instructions to use Blockway/Agens-Volundr-32B-Preview-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blockway/Agens-Volundr-32B-Preview-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blockway/Agens-Volundr-32B-Preview-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
- Ollama
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with Ollama:
ollama run hf.co/Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with Docker Model Runner:
docker model run hf.co/Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
- Lemonade
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Agens-Volundr-32B-Preview-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Blockway/Agens-Volundr-32B-Preview-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Blockway/Agens-Volundr-32B-Preview-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Agens Volundr 32B Preview — GGUF
Main model card, benchmarks and discussion: Blockway/Agens-Volundr-32B-Preview
GGUF files of Agens Volundr 32B Preview for llama.cpp: BF16, Q8_0 and Q4_K_M. Text only. For what the model is, its benchmarks and its chat format, see the main model card.
Needs the
volundrbranch of llama.cpp: github.com/BlockWayz/llama.cpp. Volundr is a new architecture (KDA linear attention, BCSA sparse attention, Engram memory, mHC hyper-connections). Stock llama.cpp, and apps built on it, will refuse these files withunknown model architecture: 'volundr'until support is merged upstream.
Files
| file | size | contents |
|---|---|---|
Agens-Volundr-32B-Preview-Q4_K_M.gguf |
26.6 GB | Q4_K_M for the feed-forward and BCSA/dense attention projections; every KDA projection, the Engram value/gate projections, the mHC dynamic projections, the token embedding and the output layer at Q8_0; Engram tables and BCSA indexer BF16. Recommended. |
Agens-Volundr-32B-Preview-Q8_0.gguf |
36.3 GB | all matrices Q8_0; Engram tables and BCSA indexer BF16 |
Agens-Volundr-32B-Preview-BF16.gguf |
64.5 GB | reference: the text model's weights unquantised (BF16) |
Norms, convolution, decay and mHC static parameters are F32 in every file. The linear-attention (KDA) weights are kept
at Q8_0 in the 4-bit file on purpose: an earlier model of ours degraded badly when its linear-attention projections
were quantised to 4 bit. SHA256SUMS lists the checksums.
Build llama.cpp (volundr branch)
git clone --branch volundr https://github.com/BlockWayz/llama.cpp && cd llama.cpp
# CPU
cmake -B build -DCMAKE_BUILD_TYPE=Release
# or NVIDIA GPU
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j --target llama-server llama-cli
If CMake cannot find libcurl, add -DLLAMA_CURL=OFF. More detail in
README-volundr.md.
Run
pip install -U huggingface_hub
hf download Blockway/Agens-Volundr-32B-Preview-GGUF Agens-Volundr-32B-Preview-Q4_K_M.gguf --local-dir .
# OpenAI-compatible server on :8080 (thinking, tool calls and reasoning_content parsing work out of the box)
./build/bin/llama-server -m Agens-Volundr-32B-Preview-Q4_K_M.gguf --jinja -c 32768 -ngl 99
# interactive chat in the terminal
./build/bin/llama-cli -m Agens-Volundr-32B-Preview-Q4_K_M.gguf --jinja -c 32768 -ngl 99
Drop -ngl 99 to run on the CPU only. When offloading to a GPU, the four Engram tables (4.3 GB of pure lookups) can
stay in host RAM: --override-tensor "engram_embd=CPU".
Thinking is on by default. Per request:
curl http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
"messages": [{"role": "user", "content": "What is 17 * 23 + 9?"}],
"chat_template_kwargs": {"reasoning_effort": "low"}
}'
# "chat_template_kwargs": {"enable_thinking": false} -> no thinking
# "chat_template_kwargs": {"reasoning_effort": "medium"} -> low | medium | xhigh (default xhigh)
Tool calls use the standard OpenAI tools field and come back as tool_calls. Sampling defaults stored in the GGUF:
temperature 1.0, top-k 20, top-p 0.95. The chat template is the model's own and is embedded in every file.
Speed
Q4_K_M on one 48 GB GPU, full offload (-ngl 99), flash attention, F16 KV cache:
| context | 1K | 8K | 32K |
|---|---|---|---|
| decode | 32.3 tok/s | 27.4 tok/s | 25.0 tok/s |
Prompt processing of a 1K-token prompt: 2,218 tok/s.
Quantisation quality
llama-perplexity on held-out chat text (mostly Cantonese and Chinese), 12 chunks of 4,096 tokens; KLD and top-1 are
measured against the BF16 file.
| file | PPL | vs BF16 | mean KLD | same top-1 |
|---|---|---|---|---|
| BF16 | 3.3916 | — | — | — |
| Q8_0 | 3.3922 | +0.02 % | 0.00056 | 99.10 % |
| Q4_K_M | 3.4145 | +0.68 % | 0.0100 | 96.00 % |
With an 8,192-token context scored on positions 4,096–8,191 (where the BCSA far field is active) the picture is the same: 3.1915 / 3.1886 / 3.2097.
Parity with the reference implementation
The llama.cpp graph was checked against the checkpoint's own PyTorch code (fp32) on all-position logits. The BF16 file agrees on 99.7–100 % of top-1 tokens (mean KL ≤ 5e-5) on prompts from 90 to 6,600 tokens in English, Cantonese, code and chat format, and 64-token greedy continuations match except where the reference itself has a near-exact tie. An F32 conversion matches the reference to fp32 rounding (100 % top-1 on a 6,600-token prompt), including BCSA's far-field block selection. CPU and CUDA builds give the same greedy continuations on the chat prompts tested.
Unused-token mask
The output layer has 248,320 rows but the tokenizer ends at id 248,076. The ids 248,077–248,319 are never valid output,
yet the model can put real probability on some of them: where a tool call opens, sampling sometimes picks one instead of
the tool-call token and the call comes back as plain text. The volundr branch therefore masks that range (a -inf
logit bias, applied automatically for this architecture in llama-server and llama-cli); the main repository lists
the same ids in generation_config.json suppress_tokens. Sampled tool calls with the Q4_K_M file:
| with mask | without mask | |
|---|---|---|
| no thinking, temperature 0.6 | 16/16 | 7/16 |
| no thinking, temperature 1.0 (default) | 16/16 | 10/16 |
| thinking, temperature 1.0 | 8/8 | 3/8 |
VOLUNDR_MASK_UNUSED_TOKENS=0 turns the mask off. If you call the llama.cpp library directly (bindings that do not use
llama.cpp's common sampling code), apply the same bias yourself.
Known limits
- Text only. The vision tower is not converted; use the main repository for image input.
- No speculative decoding in llama.cpp.
- CUDA and CPU only for the two Volundr-specific ops (Engram hashing, mHC Sinkhorn). On Metal, Vulkan or ROCm the scheduler should run those two ops on the CPU; those backends have not been tested.
- BCSA prefill cost. Prompt processing scores every query against every cached position before pooling and selecting blocks, so it slows down with prompt length (about 620 tok/s for a 32K-token prompt on the same GPU). Decoding switches to a sparse path above 12,288 cached positions.
- No context shift and no partial KV-cache removal (recurrent state), as for llama.cpp's other hybrid models.
- Multi-GPU
--split-mode layerworks but is no faster than one 48 GB GPU;--split-mode rowdoes not load.
Licence
Apache-2.0. These files are a format conversion of the weights in
Blockway/Agens-Volundr-32B-Preview; LICENSE, NOTICE
and the files they reference apply.
- Downloads last month
- 664
4-bit
8-bit
16-bit
Model tree for Blockway/Agens-Volundr-32B-Preview-GGUF
Base model
Blockway/Agens-Volundr-32B-Preview