Instructions to use Accio-Lab/occamy-1.0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Accio-Lab/occamy-1.0-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Accio-Lab/occamy-1.0-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Accio-Lab/occamy-1.0-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Accio-Lab/occamy-1.0-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Accio-Lab/occamy-1.0-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Accio-Lab/occamy-1.0-GGUF:Q4_K_M
- Ollama
How to use Accio-Lab/occamy-1.0-GGUF with Ollama:
ollama run hf.co/Accio-Lab/occamy-1.0-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Accio-Lab/occamy-1.0-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Accio-Lab/occamy-1.0-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Accio-Lab/occamy-1.0-GGUF with Docker Model Runner:
docker model run hf.co/Accio-Lab/occamy-1.0-GGUF:Q4_K_M
- Lemonade
How to use Accio-Lab/occamy-1.0-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.occamy-1.0-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Accio-Lab/occamy-1.0-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Accio-Lab/occamy-1.0-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Accio-Lab/occamy-1.0-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Accio-Lab/occamy-1.0-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Occamy-1.0 ยท GGUF
17 quantizations ยท llama.cpp ยท optional vision projector
Model collection ยท Checkpoint explorer ยท Project ยท Paper
GGUF quantizations of Accio-Lab/occamy-1.0. Q3_K_S, Q5_K_S, Q3_K_L, Q4_0 and Q5_0 are direct BF16 conversions with standard recipes and no importance matrix. The earlier ten additions used calibration importance matrices; Q4_K_M and Q8_0 retain their historical recipes. Labels describe mixed precision rather than uniform bits for every tensor. Previously published weights and evidence are preserved.
Separate APEX GGUF profiles provide seven additional MoE mixed-precision allocations: three plain profiles and four calibrated I profiles, including Mini. Compare all 24 GGUF options in the checkpoint explorer.
Quickstart
Use a recent llama.cpp build with Qwen3.5 MoE support. The recorded CUDA validation used revision 972d2313bc0bf0a45f634f77d95c9fb03aeab12c; CPU-only and Apple Silicon execution were not tested in that run.
Download one language checkpoint. Q4_K_M is the example here; replace the filename with any option below.
hf download Accio-Lab/occamy-1.0-GGUF occamy-1.0-Q4_K_M.gguf \
--local-dir ./occamy-gguf
llama-server -m ./occamy-gguf/occamy-1.0-Q4_K_M.gguf \
-ngl 999 -c 8192 -np 1 -fa on --jinja \
--host 127.0.0.1 --port 8000 --alias occamy
For image input, also download the projector and add --mmproj ./occamy-gguf/mmproj-occamy-1.0-F16.gguf to that command:
hf download Accio-Lab/occamy-1.0-GGUF mmproj-occamy-1.0-F16.gguf \
--local-dir ./occamy-gguf
Normalize prompt content to NFC before submission; see tokenizer compatibility. The example uses the native qwen2 metadata. Earlier published qwen35 benchmark measurements used a different tokenizer protocol. The Q3_K_S/Q5_K_S addition was evaluated with native qwen2 metadata and NFC input.
After the server is ready:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"occamy","messages":[{"role":"user","content":"Write a Python function that adds two integers."}],"max_tokens":256}'
Downloads
| Quantization | Exact bytes | GiB | File |
|---|---|---|---|
| IQ2_M | 11,659,235,328 | 10.859 | Download |
| Q2_K | 12,939,593,728 | 12.051 | Download |
| IQ3_XS | 14,484,144,128 | 13.489 | Download |
| Q3_K_S | 15,182,184,192 | 14.140 | Download |
| IQ3_M | 15,440,519,168 | 14.380 | Download |
| Q3_K_M | 16,764,764,160 | 15.613 | Download |
| Q3_K_L | 18,115,329,792 | 16.871 | Download |
| IQ4_XS | 18,728,777,728 | 17.443 | Download |
| Q4_0 | 19,715,053,312 | 18.361 | Download |
| IQ4_NL | 19,779,278,848 | 18.421 | Download |
| Q4_K_S | 19,889,903,616 | 18.524 | Download |
| Q4_K_M | 21,166,757,696 | 19.713 | Download |
| Q5_K_S | 23,981,283,072 | 22.334 | Download |
| Q5_0 | 23,981,283,072 | 22.334 | Download |
| Q5_K_M | 24,729,131,008 | 23.031 | Download |
| Q6_K | 28,514,152,448 | 26.556 | Download |
| Q8_0 | 36,903,139,456 | 34.369 | Download |
| F16 vision projector | 899,282,944 | 0.838 | Download |
File size is not peak RAM/VRAM. KV cache, context, concurrency and images require additional memory. The projector is optional for text-only use. New quantizations were tested for text/code/JSON/tool use; vision was not retested for them.
Tokenizer and runtime
Read tokenizer compatibility before running. All released files store qwen2. The earlier ten additions were corrected from converter-inferred qwen35 without changing tensor payloads; the five standard-recipe additions received qwen2 metadata during quantization. These rules differ on some Unicode text. The pinned source tokenizer uses NFC normalization and a qwen2-style rule. The legacy NFC and 14-case findings are preserved in the original README.
The reported benchmark used qwen35 throughout, including --override-kv tokenizer.ggml.pre=str:qwen35 on the legacy Q4_K_M/Q8_0 files to match its BF16 GGUF reference. To reproduce those measurements on the corrected new copies, the same qwen35 override is also required. It is not an out-of-the-box score for those two files or proof of tokenizer equivalence to the Transformers source. See the bounded source comparison and precise reproduction notes in TOKENIZER.md.
For normal source-compatible text input, normalize prompt content with unicodedata.normalize("NFC", text) and use the native qwen2 metadata. For example:
llama-server -m occamy-1.0-Q4_K_M.gguf -ngl 999 -c 8192 -np 1 -fa on --jinja
The earlier ten corrected release copies each passed two short GPU load/generation checks with qwen2+NFC. Those checks were not a full quality rerun. Q3_K_S and Q5_K_S have the separate native-qwen2 validation below. Published hashes and release checks are separate from the original benchmark hashes.
Measured validation
Q3_K_L / Q4_0 / Q5_0 addition ยท 2026-10-02
New validation report ยท Results and hashes ยท Reproduction
| Model | Authored functional cases | WikiText subset PPL โ |
|---|---|---|
| BF16 reference (reused) | 19/20 | 8.2963 ยฑ 0.36372 |
| Q3_K_L | 20/20 | 9.1110 ยฑ 0.40662 |
| Q4_0 | 20/20 | 8.4754 ยฑ 0.37025 |
| Q5_0 | 19/20 | 8.2284 ยฑ 0.35732 |
All three new files use native qwen2 metadata and NFC-normalized inputs. They were checked on B200 with the same pinned runtime and inputs as the preceding batch; the BF16 reference results are reused from that batch. These are twenty authored functional cases and sixteen WikiText test chunks at context 512. All 733 tensors per file were dequantized and checked for finite values; vocabulary, merges, chat template and logical shapes match the BF16 conversion. The sixteen NFC probes matched with registered special-token parsing enabled. A better result on this small subset does not establish higher general quality. Image/projector, CPU and Mac inference were not tested.
Q3_K_S / Q5_K_S addition ยท 2026-10-02
New validation report ยท Results and hashes ยท Reproduction
| Model | Authored functional cases | WikiText subset PPL โ |
|---|---|---|
| BF16 GGUF reference | 19/20 | 8.2963 ยฑ 0.36372 |
| Q3_K_S | 19/20 | 9.8643 ยฑ 0.44896 |
| Q5_K_S | 19/20 | 8.5390 ยฑ 0.37867 |
Greedy, thinking disabled, NFC input and qwen2 throughout; NVIDIA B200, pinned llama.cpp revision above. All models passed text, arithmetic, independent Python execution, memory and native tool round trips; each failed the same one of four strict JSON cases. PPL uses 16 WikiText test chunks at context 512, with fixed source revision and input hash. All 733 tensors in each new file were dequantized and checked for finite values. Source vocabulary, merges, chat template and logical tensor shapes are preserved. NFC with registered special-token parsing matched all 16 authored tokenizer probes. These bounded checks do not establish full benchmark quality or arbitrary Unicode parity. Image/projector, CPU-only and Apple Silicon inference were not tested for these two files.
Earlier twelve recipes
Validation report and machine-readable results cover a local BF16 reference and all twelve quantizations. Each has 24 GSM8K questions, 24 HumanEval+ code outputs actually executed in a sandbox, and six strict JSON plus six tool cases. Per-question outputs and results are separate. These small subsets do not reproduce full published benchmark scores. Lower-bit PPL degradation is reported, not treated as file corruption.
Runtime: CUDA llama.cpp 972d2313bc0bf0a45f634f77d95c9fb03aeab12c on NVIDIA B200; eight CPU threads per timed job, full GPU offload, Flash Attention and F16 KV cache. Engine tests use five repetitions at 512/2048-token contexts. The additional three variants ran independent jobs across four B200 cards; cross-card and shared CPU/I/O effects limit comparisons. CPU-only and Apple Silicon runtime compatibility were not established here.
Historical Q4_K_M/Q8_0 conversion and vision smoke evidence remains in TECHNICAL-DETAILS.md, VALIDATION.json and the preserved README.
MTP and license
The Occamy MTP head is separate. These language-model GGUF files do not contain an MTP head; this update does not establish GGUF draft-head compatibility. Apache 2.0.
Occamy checkpoints
BF16 ยท GGUF ยท FP8 ยท NVFP4 ยท MLX 8-bit ยท MLX 6-bit ยท MLX 4-bit ยท MLX 3-bit ยท MTP head
Compare file sizes, validation scope and deployment commands in the checkpoint explorer.
- Downloads last month
- 378,187
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit