Instructions to use mu2solutions/Molmo2-4B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use mu2solutions/Molmo2-4B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf mu2solutions/Molmo2-4B-GGUF:F16 # Run inference directly in the terminal: llama cli -hf mu2solutions/Molmo2-4B-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf mu2solutions/Molmo2-4B-GGUF:F16 # Run inference directly in the terminal: llama cli -hf mu2solutions/Molmo2-4B-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf mu2solutions/Molmo2-4B-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf mu2solutions/Molmo2-4B-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf mu2solutions/Molmo2-4B-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf mu2solutions/Molmo2-4B-GGUF:F16
Use Docker
docker model run hf.co/mu2solutions/Molmo2-4B-GGUF:F16
- LM Studio
- Jan
- vLLM
How to use mu2solutions/Molmo2-4B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mu2solutions/Molmo2-4B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mu2solutions/Molmo2-4B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/mu2solutions/Molmo2-4B-GGUF:F16
- Ollama
How to use mu2solutions/Molmo2-4B-GGUF with Ollama:
ollama run hf.co/mu2solutions/Molmo2-4B-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use mu2solutions/Molmo2-4B-GGUF with Docker Model Runner:
docker model run hf.co/mu2solutions/Molmo2-4B-GGUF:F16
- Lemonade
How to use mu2solutions/Molmo2-4B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull mu2solutions/Molmo2-4B-GGUF:F16
Run and chat with the model
lemonade run user.Molmo2-4B-GGUF-F16
List all available models
lemonade list
- Atomic Chat
Molmo2-4B GGUF
GGUF conversion of allenai/Molmo2-4B, published by Mu2 Solutions.
Molmo2-4B is Allen Institute for AI's fully open multimodal model (Apache 2.0): a SigLIP vision encoder plus a pooling cross-attention projector over a Qwen3-4B text backbone (qwen3 architecture).
License
Apache 2.0, same as the source model. No additional restrictions.
Files
| File | Quant | Size | SHA-256 |
|---|---|---|---|
Molmo2-4B-text-Q4_K_M.gguf |
Q4_K_M (text) | 2.72 GB | 7a04e31b5703943f2ac3527d33e8fecfefa35adf2d8d03f4954da28c11aee684 |
Molmo2-4B-mmproj-F16.gguf |
F16 (vision projector) | 881 MB | 1f59f440bd4c52fade2e22875e0a762d0819941637e6004a222e4397be7acac7 |
Load the text quant together with the mmproj to get a working vision model.
If you already downloaded an earlier Molmo2-4B mmproj elsewhere, replace it. Earlier F16 projector builds of this model stored the pooling tensors in a fused
mm.pool.kvlayout, which current Molmo2 support in llama.cpp no longer accepts, so those files cannot load. The file listed above is the corrected build.
Conversion
- Source:
allenai/Molmo2-4B, revision042abfa7a388(frozen 2026-01-23) - Converter: llama.cpp HF-to-GGUF with Mu2's Molmo2 vision converter at commit
c3d25ece9 - Text architecture:
qwen3 - mmproj:
clip, 315 tensors (SigLIP ViT, pooling cross-attention, SwiGLU projector), with separatemm.pool.k/mm.pool.vtensors and the bakedmm.image_patch_addrow - Chat template embedded from the source repository
Verification (2026-09-28)
Measured on Mu2's CPU-only OG Rig (Intel i3-2100, no GPU), with the reconstructed projector:
- Vision encode: completed in 490,705 ms for one 378x378 view.
- Vision generation: asked
Describe this image in one sentence.on a real image, the model answered:
The image is a corporate wordmark on a white background with navy and dark red lettering, so the colours, the layout and the visible letters are all correct. A genuine reading of the image.The image displays a partially visible logo and text on a white background. The logo features large, stylized letters in blue and red, with the letters "R-E-V-I" prominently shown. Below
Command used:
llama-mtmd-cli -m Molmo2-4B-text-Q4_K_M.gguf \
--mmproj Molmo2-4B-mmproj-F16.gguf \
--image image.jpg -c 4096 -t 4 -n 40 --temp 0 \
-p "Describe this image in one sentence."
Requirements and known issues
- Needs a llama.cpp build with Molmo2 mtmd support. Upstream llama.cpp does not support the Molmo2 vision graph yet; this release was produced and verified with Mu2's fork.
- CPU-only encoding is slow. The vision tower is F16 and this machine has no F16C, so a single view took about eight minutes. A GPU or a modern CPU changes this by orders of magnitude.
- Token-type warning at load. llama.cpp may warn that the
</s>token is not control-type and overrides it. Cosmetic; generation is unaffected.
Usage
llama-mtmd-cli -m Molmo2-4B-text-Q4_K_M.gguf \
--mmproj Molmo2-4B-mmproj-F16.gguf \
--image photo.jpg -p "Describe this image."
- Downloads last month
- 148
4-bit