Qwen3.6-35B-A3B GPTQ-Pro FOEM banner

Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128

Overview

Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo. It is intended for open-source evaluation, reproducible experimentation, and compatible local or hosted inference workflows. The wording below is deliberately limited to what can be verified from this repository's metadata and artifacts.

At a glance

Field Details
Format GPTQ
Source / base Qwen/Qwen3.6-35B-A3B
Intended task image-text-to-text
License other

What is included

  • *.safetensors (6 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • processor_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (25 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128 \
  --quantization gptq_marlin \
  --dtype float16 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

A fast, production-minded GPTQ-Pro quantization of Qwen3.6-35B-A3B, built for local agentic coding, reasoning, tool use, and high-throughput vLLM serving on dual 24GB GPUs.

This checkpoint takes the official Qwen3.6-35B-A3B model — a sparse MoE with 35B total parameters and 3B active parameters per token — and packages it into a practical 4-bit GPTQ-Pro deployment profile for builders who want Qwen3.6-class capability without needing enterprise GPU memory.

The goal is simple: strong Qwen3.6 reasoning, much lower VRAM pressure, clean vLLM deployment, and real benchmark numbers instead of vibes.


Why this checkpoint exists

Official Qwen3.6-35B-A3B is designed for real-world developer workflows: agentic coding, frontend tasks, repository-level reasoning, tool use, and iterative problem solving. This quantized release keeps that use case in focus while making the model easier to run on accessible local hardware.

This release is tuned for:

  • Local coding agents
  • OpenAI-compatible inference servers
  • Multi-agent workflows
  • Tool-calling systems
  • Long-context reasoning
  • Repository analysis
  • Private/local AI stacks
  • Dual RTX 3090 / 4090-class deployments

Key features

Feature Details
Base model Qwen/Qwen3.6-35B-A3B
Architecture Sparse Mixture-of-Experts
Parameters 35B total / 3B active
Quantization GPTQ-Pro FOEM 4-bit
Group size 128
vLLM backend gptq_marlin
Tested serving profile TP2 across 2x RTX 3090
Tested context 150K max sequence length
API style OpenAI-compatible /v1/chat/completions
Reasoning parser qwen3
Tool parser qwen3_coder

Why GPTQ-Pro FOEM?

This quantization is aimed at the ugly real world: limited VRAM, large prompts, coding-agent loops, and production serving where speed and stability matter.

GPTQ-Pro FOEM 4-bit gives this model a practical deployment shape:

  • smaller memory footprint than full precision
  • fast vLLM serving through gptq_marlin
  • good throughput on consumer GPUs
  • practical long-context operation
  • clean OpenAI-compatible integration
  • no exotic runtime ceremony

This is not a toy quant for screenshots. It is built to serve.


Benchmark snapshot

Validated on a TP2 vLLM deployment with 2x RTX 3090 GPUs.

Metric Value
Total tests 109
Latency stdev 171.63 ms
Mean throughput 170.90 tokens/sec
Median throughput 171.50 tokens/sec
Min throughput 150.54 tokens/sec
Max throughput 174.13 tokens/sec

Benchmark mix: reasoning, coding, math, general QA, and practical chat prompts.

These numbers are provided to give builders an actual deployment baseline. Your exact results will depend on GPUs, PCIe topology, CPU, RAM, CUDA stack, vLLM version, context length, batch size, and sampling settings.


Recommended deployment

Validated with:

  • vLLM 0.19.0
  • Tensor parallel size: 2
  • Quantization backend: gptq_marlin
  • Reasoning parser: qwen3
  • Tool call parser: qwen3_coder
  • Tested max model length: 150000
  • Hardware: 2x RTX 3090-class GPUs

The official Qwen3.6-35B-A3B model supports a larger native context window, but this quantized release was specifically validated at 150K in the tested setup. Treat 150K as the known-good production profile.


Download

hf download groxaxo/Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128 \
  --local-dir ./qwen36-gptqpro
Downloads last month
46
Safetensors
Model size
7B params
Tensor type
BF16
·
I32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for groxaxo/Qwen3.6-35B-A3B-GPTQ-Pro-FOEM-4bit-g128

Finetuned
(350)
this model