How to use from
Docker Model Runner
docker model run hf.co/minsore/Quill-Gen-1:Q4_K_M
Quick Links

🟒 Quill Gen 1

Lightweight Python code generation model built on Qwen2.5-Coder-0.5B.

Quill Gen 1 is a 0.5B parameter experimental model for generating Python functions from natural-language descriptions. It is the younger sibling of Pepper 1 Preview β€” same task, one third the size. Trained as an experiment in small-model code generation, it reaches 80% on HumanEval@50 while running comfortably on consumer hardware.

⚠️ This is NOT a FIM model. Quill Gen 1 cannot do fill-in-the-middle completion. For autocomplete use Quill 1 Preview.

Part of the Minsore family Β· minsore.com

✨ Highlights

  • 🧠 Instruction-tuned β€” writes Python functions from plain descriptions
  • πŸ“¦ Tiny β€” 0.5B params, ~400 MB in Q4_K_M
  • 🎯 Purpose-built β€” code generation only, not chat, not FIM
  • ⚑ Fast β€” designed for low-VRAM setups
  • πŸ§ͺ Experimental β€” a proof of concept for small-model generation
  • πŸ†“ Apache 2.0 β€” same license as base model

πŸ“Š Benchmarks

Quill Gen comparison

Evaluated against 0.5B–1.5B code models. All runs used temperature=0.0, max_tokens=512, --chat-template none.

Benchmark Quill Gen 1 Pepper 1 Preview Quill 1 Preview Qwen2.5-Coder 0.5B
HumanEval@50 80.0% 82.0% 54.0% 52.0%
HumanEval-Infilling (EditSim) 0.161 β€” 0.423 0.045
HumanEval-Infilling (Exact Match) 0% β€” 16% 0%
Delulu FIM (EditSim) 0.048 β€” 0.318 0.035

πŸ“Œ Key insight: Quill Gen 1 reaches 80% HumanEval@50 β€” nearly matching Pepper 1 Preview (82%) at one third the parameter count. On FIM-specific tasks it underperforms significantly, which is expected: this model was trained for generation, not autocomplete.

πŸš€ Quick Start

llama.cpp

llama-server -m quill-gen-1.Q4_K_M.gguf \
  --port 8080 \
  -ngl 99 \
  -c 768 \
  --chat-template none

Requirements: any GPU with β‰₯2 GB VRAM (full offload), or partial CPU offload as fallback.

Request

curl http://localhost:8080/completion \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "### Instruction:\nWrite a Python function that checks if a number is prime.\n\n### Response:\n",
    "n_predict": 256,
    "temperature": 0.0
  }'

Python

import requests

def generate(instruction, max_tokens=256):
    prompt = f"### Instruction:\n{instruction}\n\n### Response:\n"
    r = requests.post("http://localhost:8080/completion", json={
        "prompt": prompt,
        "n_predict": max_tokens,
        "temperature": 0.0,
        "stop": ["### Instruction:", "<|endoftext|>"],
    })
    return r.json()["content"]

print(generate("Write a Python function that reverses a string."))
# β†’ def reverse_string(s):
#       return s[::-1]

βš™οΈ Recommended Settings

Parameter Value Notes
--chat-template none Required. Quill Gen 1 is not a chat model
n_predict 128–256 Works best on short completions
temperature 0.0 Deterministic; use 0.2 for variation
repeat_penalty 1.1 Prevents repetition
-c 768 Matches training context length

⚠️ Limitations

  • Generation only β€” does not support fill-in-the-middle, tool calling, or chat
  • Python-only β€” trained exclusively on Python code
  • Small context β€” 768 tokens; long files are truncated
  • Weak on FIM β€” EditSim 0.16, Exact Match 0% (use Quill 1 instead)
  • Occasional over-explanation β€” may include comments when only code is requested
  • Number looping β€” may repeat large integers on some prompts (use repeat_penalty=1.1)

🧬 Training Details

Base model Qwen2.5-Coder-0.5B-Base
Method QLoRA (r=16, Ξ±=16)
Data FIM-converted Python instruction data
Epochs 1
Context 768 tokens
Optimizer paged_adamw_8bit

πŸ“ Files

File Size Description
quill-gen-1.Q4_K_M.gguf ~400 MB Ready to use with llama.cpp (recommended)
model.safetensors ~1 GB Full-precision merged weights
config.json β€” Model config
tokenizer.json β€” Tokenizer
tokenizer_config.json β€” Tokenizer config
generation_config.json β€” Generation defaults (optional)

πŸ—ΊοΈ Roadmap

  • Quill 2 β€” FIM autocomplete, fixed suffix handling, JS/TS/Rust support
  • Pepper 2 β€” improved MBPP and LiveCodeBench
  • Symphony β€” flagship agentic code model (3B MoE)

πŸ“œ License

Apache 2.0 β€” same as the base Qwen2.5-Coder-0.5B model.

πŸ™ Credits

πŸ“¬ Contact

Minsore β€” Ukrainian AI lab building open language models.

⭐ If Quill Gen 1 is useful, star the repo and share your results.

Downloads last month
140
Safetensors
Model size
0.5B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results