Text Generation
MLX
Safetensors
GGUF
Rust
qwen3_5_text
4b
agentic-coding
alloy-backfilled
android
apple-silicon
attested
bash
c
chain-of-custody
chinese
code
code-completion
code-generation
code-infill
coder
coding
consumer-gpu
cpp
cryptographically-verified
css
defragged
delta-forge
derivative
edge-inference
embedded
english
forge-alloy
function-calling
ggml
go
html
iphone
java
javascript
kotlin
llama-cpp
lm-studio
local-inference
macbook
mobile
multilingual
ollama
on-device
php
programming
python
q4-k-m
quantized
qwen
qwen3
qwen3.5
raspberry-pi
reproducible
ruby
software-engineering
sql
swift
typescript
conversational
Instructions to use continuum-ai/qwen3.5-4b-code-forged-defragged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use continuum-ai/qwen3.5-4b-code-forged-defragged with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("continuum-ai/qwen3.5-4b-code-forged-defragged") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use continuum-ai/qwen3.5-4b-code-forged-defragged with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "continuum-ai/qwen3.5-4b-code-forged-defragged"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "continuum-ai/qwen3.5-4b-code-forged-defragged" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use continuum-ai/qwen3.5-4b-code-forged-defragged with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "continuum-ai/qwen3.5-4b-code-forged-defragged"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "continuum-ai/qwen3.5-4b-code-forged-defragged" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "continuum-ai/qwen3.5-4b-code-forged-defragged", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use continuum-ai/qwen3.5-4b-code-forged-defragged with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "continuum-ai/qwen3.5-4b-code-forged-defragged"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default continuum-ai/qwen3.5-4b-code-forged-defragged
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use continuum-ai/qwen3.5-4b-code-forged-defragged with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "continuum-ai/qwen3.5-4b-code-forged-defragged"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "continuum-ai/qwen3.5-4b-code-forged-defragged" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Correct qwen3.5-4b-code-forged-defragged.alloy.json pass@1 to canonical evalplus convention (v1.0.0)
518aceb verified Download qwen3.5-4b-code-forged-defragged.alloy.json from continuum-ai/qwen3.5-4b-code-forged-defragged: direct link, hf CLI and curl.
- Browser
- Download file 3.42 kB
-
https://huggingface.co/continuum-ai/qwen3.5-4b-code-forged-defragged/resolve/main/qwen3.5-4b-code-forged-defragged.alloy.json
- Command line
-
hf download hf://continuum-ai/qwen3.5-4b-code-forged-defragged/qwen3.5-4b-code-forged-defragged.alloy.json
-
curl -L -o qwen3.5-4b-code-forged-defragged.alloy.json https://huggingface.co/continuum-ai/qwen3.5-4b-code-forged-defragged/resolve/main/qwen3.5-4b-code-forged-defragged.alloy.json
3.42 kB
| { | |
| "name": "qwen3.5-4b-code-forged-defragged", | |
| "version": "1.0.0", | |
| "description": "DEFRAGGED derivative of [`qwen3.5-4b-code-forged`](https://huggingface.co/continuum-ai/qwen3.5-4b-code-forged). Same forge journey as the parent (prune + train as published in the parent's alloy); this artifact adds a single 'defragged' transformation stage to produce a smaller / faster / more-portable variant of the same logical model. Inherits the parent's published benchmark results; per-variant evaluation samples will land in a follow-up release if/when per-variant benchmarks are run.", | |
| "author": "continuum-ai", | |
| "tags": [ | |
| "alloy-backfilled", | |
| "forge-alloy", | |
| "defragged", | |
| "delta-forge", | |
| "derivative" | |
| ], | |
| "license": "apache-2.0", | |
| "source": { | |
| "baseModel": "Qwen/Qwen3.5-4B", | |
| "architecture": "qwen3_5", | |
| "isMoE": false | |
| }, | |
| "stages": [ | |
| { | |
| "type": "train", | |
| "domain": "code", | |
| "steps": 1000, | |
| "learningRate": "2e-4" | |
| }, | |
| { | |
| "type": "quant", | |
| "format": "gguf", | |
| "quantTypes": [ | |
| "Q4_K_M" | |
| ], | |
| "deviceTargets": [] | |
| }, | |
| { | |
| "type": "eval", | |
| "benchmarks": [ | |
| { | |
| "name": "humaneval" | |
| } | |
| ], | |
| "compareToBase": true | |
| }, | |
| { | |
| "type": "package", | |
| "format": "safetensors-defragged", | |
| "validateOn": [], | |
| "includeTokenizer": true, | |
| "notes": "Defrag-only derivative of the parent forge. The parent's prune stage marks heads as dead via forward-hooks; this artifact reifies that pruning by physically reshaping the projection matrices to remove the dead heads' parameters. Behaviorally equivalent to the parent (same logits per surviving head); structurally smaller on disk and in VRAM." | |
| } | |
| ], | |
| "cycles": 3, | |
| "derivedFrom": { | |
| "repo": "continuum-ai/qwen3.5-4b-code-forged", | |
| "alloyHash": null, | |
| "kind": "defragged" | |
| }, | |
| "results": { | |
| "completedAt": "2026-03-31T12:13:43-0500", | |
| "baselinePerplexity": 3.0382, | |
| "finalPerplexity": 2.3487, | |
| "improvementPct": 22.7, | |
| "benchmarks": [ | |
| { | |
| "name": "perplexity", | |
| "metrics": { | |
| "baseline": 3.0382, | |
| "final": 2.3487, | |
| "improvement": 22.7 | |
| } | |
| }, | |
| { | |
| "name": "humaneval", | |
| "subset": null, | |
| "metrics": { | |
| "status": "pending" | |
| }, | |
| "submittedToLeaderboard": false | |
| } | |
| ], | |
| "hardwareVerified": [ | |
| { | |
| "device": "NVIDIA GeForce RTX 5090", | |
| "format": "fp16", | |
| "verified": true | |
| } | |
| ], | |
| "samples": [], | |
| "integrity": { | |
| "trustLevel": "self-attested", | |
| "code": { | |
| "runner": "sentinel-ai/derive_alloy_from_parent (defragged)", | |
| "version": "1.0", | |
| "binaryHash": "sha256:derivation-tool-only" | |
| }, | |
| "modelHash": "sha256:4d59fce78f3541375dbc1adf849fa0474426dd4036fe958bad4e9270d7fe776d", | |
| "fileHashes": [ | |
| { | |
| "filename": "model-00001-of-00002.safetensors", | |
| "sha256": "a1bd60ee8c791971867535382ac26a59165d04f6ba41a39bb7365bed37b41c07", | |
| "size": 5351237632 | |
| }, | |
| { | |
| "filename": "model-00002-of-00002.safetensors", | |
| "sha256": "a5529bb11406d9d72422c6728453e10b32628976cc6304f3c0ebe99fa6f3c16e", | |
| "size": 2913513560 | |
| } | |
| ], | |
| "datasets": [], | |
| "attestedAt": "2026-04-08", | |
| "parentAlloyHash": null | |
| } | |
| } | |
| } | |