Accio

Occamy-1.0 ยท GGUF

17 quantizations ยท llama.cpp ยท optional vision projector

Model collection ยท Checkpoint explorer ยท Project ยท Paper

GGUF quantizations of Accio-Lab/occamy-1.0. Q3_K_S, Q5_K_S, Q3_K_L, Q4_0 and Q5_0 are direct BF16 conversions with standard recipes and no importance matrix. The earlier ten additions used calibration importance matrices; Q4_K_M and Q8_0 retain their historical recipes. Labels describe mixed precision rather than uniform bits for every tensor. Previously published weights and evidence are preserved.

Separate APEX GGUF profiles provide seven additional MoE mixed-precision allocations: three plain profiles and four calibrated I profiles, including Mini. Compare all 24 GGUF options in the checkpoint explorer.

Quickstart

Use a recent llama.cpp build with Qwen3.5 MoE support. The recorded CUDA validation used revision 972d2313bc0bf0a45f634f77d95c9fb03aeab12c; CPU-only and Apple Silicon execution were not tested in that run.

Download one language checkpoint. Q4_K_M is the example here; replace the filename with any option below.

hf download Accio-Lab/occamy-1.0-GGUF occamy-1.0-Q4_K_M.gguf \
  --local-dir ./occamy-gguf

llama-server -m ./occamy-gguf/occamy-1.0-Q4_K_M.gguf \
  -ngl 999 -c 8192 -np 1 -fa on --jinja \
  --host 127.0.0.1 --port 8000 --alias occamy

For image input, also download the projector and add --mmproj ./occamy-gguf/mmproj-occamy-1.0-F16.gguf to that command:

hf download Accio-Lab/occamy-1.0-GGUF mmproj-occamy-1.0-F16.gguf \
  --local-dir ./occamy-gguf

Normalize prompt content to NFC before submission; see tokenizer compatibility. The example uses the native qwen2 metadata. Earlier published qwen35 benchmark measurements used a different tokenizer protocol. The Q3_K_S/Q5_K_S addition was evaluated with native qwen2 metadata and NFC input.

After the server is ready:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"occamy","messages":[{"role":"user","content":"Write a Python function that adds two integers."}],"max_tokens":256}'

Downloads

Quantization Exact bytes GiB File
IQ2_M 11,659,235,328 10.859 Download
Q2_K 12,939,593,728 12.051 Download
IQ3_XS 14,484,144,128 13.489 Download
Q3_K_S 15,182,184,192 14.140 Download
IQ3_M 15,440,519,168 14.380 Download
Q3_K_M 16,764,764,160 15.613 Download
Q3_K_L 18,115,329,792 16.871 Download
IQ4_XS 18,728,777,728 17.443 Download
Q4_0 19,715,053,312 18.361 Download
IQ4_NL 19,779,278,848 18.421 Download
Q4_K_S 19,889,903,616 18.524 Download
Q4_K_M 21,166,757,696 19.713 Download
Q5_K_S 23,981,283,072 22.334 Download
Q5_0 23,981,283,072 22.334 Download
Q5_K_M 24,729,131,008 23.031 Download
Q6_K 28,514,152,448 26.556 Download
Q8_0 36,903,139,456 34.369 Download
F16 vision projector 899,282,944 0.838 Download

File size is not peak RAM/VRAM. KV cache, context, concurrency and images require additional memory. The projector is optional for text-only use. New quantizations were tested for text/code/JSON/tool use; vision was not retested for them.

Tokenizer and runtime

Read tokenizer compatibility before running. All released files store qwen2. The earlier ten additions were corrected from converter-inferred qwen35 without changing tensor payloads; the five standard-recipe additions received qwen2 metadata during quantization. These rules differ on some Unicode text. The pinned source tokenizer uses NFC normalization and a qwen2-style rule. The legacy NFC and 14-case findings are preserved in the original README.

The reported benchmark used qwen35 throughout, including --override-kv tokenizer.ggml.pre=str:qwen35 on the legacy Q4_K_M/Q8_0 files to match its BF16 GGUF reference. To reproduce those measurements on the corrected new copies, the same qwen35 override is also required. It is not an out-of-the-box score for those two files or proof of tokenizer equivalence to the Transformers source. See the bounded source comparison and precise reproduction notes in TOKENIZER.md.

For normal source-compatible text input, normalize prompt content with unicodedata.normalize("NFC", text) and use the native qwen2 metadata. For example:

llama-server -m occamy-1.0-Q4_K_M.gguf -ngl 999 -c 8192 -np 1 -fa on --jinja

The earlier ten corrected release copies each passed two short GPU load/generation checks with qwen2+NFC. Those checks were not a full quality rerun. Q3_K_S and Q5_K_S have the separate native-qwen2 validation below. Published hashes and release checks are separate from the original benchmark hashes.

Measured validation

Q3_K_L / Q4_0 / Q5_0 addition ยท 2026-10-02

New validation report ยท Results and hashes ยท Reproduction

Model Authored functional cases WikiText subset PPL โ†“
BF16 reference (reused) 19/20 8.2963 ยฑ 0.36372
Q3_K_L 20/20 9.1110 ยฑ 0.40662
Q4_0 20/20 8.4754 ยฑ 0.37025
Q5_0 19/20 8.2284 ยฑ 0.35732

All three new files use native qwen2 metadata and NFC-normalized inputs. They were checked on B200 with the same pinned runtime and inputs as the preceding batch; the BF16 reference results are reused from that batch. These are twenty authored functional cases and sixteen WikiText test chunks at context 512. All 733 tensors per file were dequantized and checked for finite values; vocabulary, merges, chat template and logical shapes match the BF16 conversion. The sixteen NFC probes matched with registered special-token parsing enabled. A better result on this small subset does not establish higher general quality. Image/projector, CPU and Mac inference were not tested.

Q3_K_S / Q5_K_S addition ยท 2026-10-02

New validation report ยท Results and hashes ยท Reproduction

Model Authored functional cases WikiText subset PPL โ†“
BF16 GGUF reference 19/20 8.2963 ยฑ 0.36372
Q3_K_S 19/20 9.8643 ยฑ 0.44896
Q5_K_S 19/20 8.5390 ยฑ 0.37867

Greedy, thinking disabled, NFC input and qwen2 throughout; NVIDIA B200, pinned llama.cpp revision above. All models passed text, arithmetic, independent Python execution, memory and native tool round trips; each failed the same one of four strict JSON cases. PPL uses 16 WikiText test chunks at context 512, with fixed source revision and input hash. All 733 tensors in each new file were dequantized and checked for finite values. Source vocabulary, merges, chat template and logical tensor shapes are preserved. NFC with registered special-token parsing matched all 16 authored tokenizer probes. These bounded checks do not establish full benchmark quality or arbitrary Unicode parity. Image/projector, CPU-only and Apple Silicon inference were not tested for these two files.

Earlier twelve recipes

Validation report and machine-readable results cover a local BF16 reference and all twelve quantizations. Each has 24 GSM8K questions, 24 HumanEval+ code outputs actually executed in a sandbox, and six strict JSON plus six tool cases. Per-question outputs and results are separate. These small subsets do not reproduce full published benchmark scores. Lower-bit PPL degradation is reported, not treated as file corruption.

Runtime: CUDA llama.cpp 972d2313bc0bf0a45f634f77d95c9fb03aeab12c on NVIDIA B200; eight CPU threads per timed job, full GPU offload, Flash Attention and F16 KV cache. Engine tests use five repetitions at 512/2048-token contexts. The additional three variants ran independent jobs across four B200 cards; cross-card and shared CPU/I/O effects limit comparisons. CPU-only and Apple Silicon runtime compatibility were not established here.

Historical Q4_K_M/Q8_0 conversion and vision smoke evidence remains in TECHNICAL-DETAILS.md, VALIDATION.json and the preserved README.

MTP and license

The Occamy MTP head is separate. These language-model GGUF files do not contain an MTP head; this update does not establish GGUF draft-head compatibility. Apache 2.0.

Occamy checkpoints

BF16 ยท GGUF ยท FP8 ยท NVFP4 ยท MLX 8-bit ยท MLX 6-bit ยท MLX 4-bit ยท MLX 3-bit ยท MTP head

Compare file sizes, validation scope and deployment commands in the checkpoint explorer.

Downloads last month
378,187
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Accio-Lab/occamy-1.0-GGUF

Quantized
(32)
this model

Spaces using Accio-Lab/occamy-1.0-GGUF 2

Collection including Accio-Lab/occamy-1.0-GGUF

Paper for Accio-Lab/occamy-1.0-GGUF