Molmo2-4B GGUF

GGUF conversion of allenai/Molmo2-4B, published by Mu2 Solutions.

Molmo2-4B is Allen Institute for AI's fully open multimodal model (Apache 2.0): a SigLIP vision encoder plus a pooling cross-attention projector over a Qwen3-4B text backbone (qwen3 architecture).

License

Apache 2.0, same as the source model. No additional restrictions.

Files

File Quant Size SHA-256
Molmo2-4B-text-Q4_K_M.gguf Q4_K_M (text) 2.72 GB 7a04e31b5703943f2ac3527d33e8fecfefa35adf2d8d03f4954da28c11aee684
Molmo2-4B-mmproj-F16.gguf F16 (vision projector) 881 MB 1f59f440bd4c52fade2e22875e0a762d0819941637e6004a222e4397be7acac7

Load the text quant together with the mmproj to get a working vision model.

If you already downloaded an earlier Molmo2-4B mmproj elsewhere, replace it. Earlier F16 projector builds of this model stored the pooling tensors in a fused mm.pool.kv layout, which current Molmo2 support in llama.cpp no longer accepts, so those files cannot load. The file listed above is the corrected build.

Conversion

  • Source: allenai/Molmo2-4B, revision 042abfa7a388 (frozen 2026-01-23)
  • Converter: llama.cpp HF-to-GGUF with Mu2's Molmo2 vision converter at commit c3d25ece9
  • Text architecture: qwen3
  • mmproj: clip, 315 tensors (SigLIP ViT, pooling cross-attention, SwiGLU projector), with separate mm.pool.k / mm.pool.v tensors and the baked mm.image_patch_add row
  • Chat template embedded from the source repository

Verification (2026-09-28)

Measured on Mu2's CPU-only OG Rig (Intel i3-2100, no GPU), with the reconstructed projector:

  • Vision encode: completed in 490,705 ms for one 378x378 view.
  • Vision generation: asked Describe this image in one sentence. on a real image, the model answered:

    The image displays a partially visible logo and text on a white background. The logo features large, stylized letters in blue and red, with the letters "R-E-V-I" prominently shown. Below

    The image is a corporate wordmark on a white background with navy and dark red lettering, so the colours, the layout and the visible letters are all correct. A genuine reading of the image.

Command used:

llama-mtmd-cli -m Molmo2-4B-text-Q4_K_M.gguf \
  --mmproj Molmo2-4B-mmproj-F16.gguf \
  --image image.jpg -c 4096 -t 4 -n 40 --temp 0 \
  -p "Describe this image in one sentence."

Requirements and known issues

  • Needs a llama.cpp build with Molmo2 mtmd support. Upstream llama.cpp does not support the Molmo2 vision graph yet; this release was produced and verified with Mu2's fork.
  • CPU-only encoding is slow. The vision tower is F16 and this machine has no F16C, so a single view took about eight minutes. A GPU or a modern CPU changes this by orders of magnitude.
  • Token-type warning at load. llama.cpp may warn that the </s> token is not control-type and overrides it. Cosmetic; generation is unaffected.

Usage

llama-mtmd-cli -m Molmo2-4B-text-Q4_K_M.gguf \
  --mmproj Molmo2-4B-mmproj-F16.gguf \
  --image photo.jpg -p "Describe this image."
Downloads last month
148
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mu2solutions/Molmo2-4B-GGUF

Quantized
(6)
this model