Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128
Overview
Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.
At a glance
| Field | Details |
|---|---|
| Format | GPTQ |
| Source / base | Qwen/Qwen3.6-35B-A3B |
| Intended task | image-text-to-text |
| License | other |
What is included
*.safetensors(6 files)config.jsongeneration_config.jsontokenizer.jsontokenizer_config.jsonprocessor_config.jsonchat_template.jinjaquantize_config.json- Additional configuration, tokenizer, processor, or shard files (25 visible artifacts total)
Quick start
vLLM (documented configuration)
vllm serve groxaxo/Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128 \
--quantization gptq_marlin \
--dtype float16 \
--trust-remote-code
This command is taken from the repository documentation. Adjust tensor parallelism, context length, and cache settings to match your hardware and vLLM version.
Compatibility and responsible use
- Use a runtime that explicitly supports this format, architecture, and modality.
- Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
- Review the source model card and license before redistribution or deployment.
- Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
- Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.
Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.
Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.
A fast, production-minded GPTQ-Pro quantization of Qwen3.6-35B-A3B, built for local agentic coding, reasoning, tool use, and high-throughput vLLM serving on dual 24GB GPUs.
This checkpoint takes the official Qwen3.6-35B-A3B model — a sparse MoE with 35B total parameters and 3B active parameters per token — and packages it into a practical 4-bit GPTQ-Pro deployment profile for builders who want Qwen3.6-class capability without needing enterprise GPU memory.
The goal is simple: strong Qwen3.6 reasoning, much lower VRAM pressure, clean vLLM deployment, and real benchmark numbers instead of vibes.
Why this checkpoint exists
Official Qwen3.6-35B-A3B is designed for real-world developer workflows: agentic coding, frontend tasks, repository-level reasoning, tool use, and iterative problem solving. This quantized release keeps that use case in focus while making the model easier to run on accessible local hardware.
This release is tuned for:
- Local coding agents
- OpenAI-compatible inference servers
- Multi-agent workflows
- Tool-calling systems
- Long-context reasoning
- Repository analysis
- Private/local AI stacks
- Dual RTX 3090 / 4090-class deployments
Key features
| Feature | Details |
|---|---|
| Base model | Qwen/Qwen3.6-35B-A3B |
| Architecture | Sparse Mixture-of-Experts |
| Parameters | 35B total / 3B active |
| Quantization | GPTQ-Pro FOEM 4-bit |
| Group size | 128 |
| vLLM backend | gptq_marlin |
| Tested serving profile | TP2 across 2x RTX 3090 |
| Tested context | 150K max sequence length |
| API style | OpenAI-compatible /v1/chat/completions |
| Reasoning parser | qwen3 |
| Tool parser | qwen3_coder |
Why GPTQ-Pro FOEM?
This quantization is aimed at the ugly real world: limited VRAM, large prompts, coding-agent loops, and production serving where speed and stability matter.
GPTQ-Pro FOEM 4-bit gives this model a practical deployment shape:
- smaller memory footprint than full precision
- fast vLLM serving through
gptq_marlin - good throughput on consumer GPUs
- practical long-context operation
- clean OpenAI-compatible integration
- no exotic runtime ceremony
This is not a toy quant for screenshots. It is built to serve.
Benchmark snapshot
Validated on a TP2 vLLM deployment with 2x RTX 3090 GPUs.
| Metric | Value |
|---|---|
| Total tests | 109 |
| Latency stdev | 171.63 ms |
| Mean throughput | 170.90 tokens/sec |
| Median throughput | 171.50 tokens/sec |
| Min throughput | 150.54 tokens/sec |
| Max throughput | 174.13 tokens/sec |
Benchmark mix: reasoning, coding, math, general QA, and practical chat prompts.
These numbers are provided to give builders an actual deployment baseline. Your exact results will depend on GPUs, PCIe topology, CPU, RAM, CUDA stack, vLLM version, context length, batch size, and sampling settings.
Recommended deployment
Validated with:
- vLLM 0.19.0
- Tensor parallel size:
2 - Quantization backend:
gptq_marlin - Reasoning parser:
qwen3 - Tool call parser:
qwen3_coder - Tested max model length:
150000 - Hardware: 2x RTX 3090-class GPUs
The official Qwen3.6-35B-A3B model supports a larger native context window, but this quantized release was specifically validated at 150K in the tested setup. Treat 150K as the known-good production profile.
Download
hf download groxaxo/Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128 \
--local-dir ./qwen36-gptqpro
- Downloads last month
- 46
Model tree for groxaxo/Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128
Base model
Qwen/Qwen3.6-35B-A3B