Instructions to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l") model = AutoModelForMultimodalLM.from_pretrained("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l
- SGLang
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l with Docker Model Runner:
docker model run hf.co/littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l
Prismyra decision FP8, 40 layers, for Qwen3.6-35B-A3B-FP8
A full-weight FP8 checkpoint of Qwen/Qwen3.6-35B-A3B-FP8, every one of its 40 layers, with the Prismyra decision LoRA folded into the FP8 weights and requantized per 128x128 block. Prismyra is an engine that answers typed questions (booleans, choices) about a document by reading the probability of each declared option's token in a single forward pass, rather than generating an answer -- see the Prismyra repository for what that trades off and where it does not apply.
This is the checkpoint the adapter was trained against, and the teacher whose output distribution the two shorter checkpoints in this family were distilled from.
Use
pip install "prismyra[server,fast] @ git+https://github.com/littlemex/Prismyra@v0.2.2"
from prismyra import Prismyra, Boolean
engine = Prismyra("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l")
result = engine.ask(document, [Boolean(id="q", prompt="Is this a two-sided agreement?")])
prismyra-serve --model littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l --require-kernels
--require-kernels refuses to start rather than serve at a fraction of the speed. Prismyra's fused kernels are
matched against the checkpoint by per-layer tensor shape, not by layer count, so they apply to this checkpoint the
same way they apply to the base model.
Evaluation
Five sets, each checkpoint served by Prismyra on one L40S: RACE (150 articles, 579 short-passage questions), BoolQ (400), bury7k and bury10k (the same RACE questions buried in about 7,000 and 10,000 tokens of unrelated articles, 157 questions each), and Kev (a 764-question transfer test built from five other datasets, of which 228 questions come from three families this model never trained on in any form). The bar for each set is the better-scoring of two other decision models, Decider-4B and Lux-9B; which one wins changes by set.
| set | this checkpoint | bar | which competitor |
|---|---|---|---|
| RACE | 95.85 | 93.78 | Decider-4B |
| BoolQ | 90.25 | 89.50 | Decider-4B |
| bury7k | 93.63 | 91.08 | Decider-4B |
| bury10k | 94.27 | 91.08 | Decider-4B |
| Kev | 85.47 | 82.85 | Lux-9B |
The BoolQ margin is 3 questions out of 400 (361 against 358) and is within a single evaluation run's noise; read it
as a tie with the bar, not a clear win. The other four sets clear their bar by a wider margin. Options in every
Choice question here are read in the order the source data gives them (perm 1); a companion checkpoint in this
family (36 layers) was additionally scored with the four cyclic rotations of each Choice averaged (perm 4) and
moved by 0.2 to 1.9 points per set under that averaging, which is the size of run-to-run noise on these sets, not
evidence that one ordering is correct.
This checkpoint is also measured against JevBench (231 questions none of the checkpoints in this family trained on) and against four general-purpose LLMs on the same five sets above, in the training recipe's README.
Latency
Two conditions, both measured directly from Python (no network), at the engine's default request width
(group=32). Against the untrained base and the other two published checkpoints, same script and GPU, on the same
~5,300-token document repeated across calls:
| checkpoint | 1 question | 16 questions | 64 questions |
|---|---|---|---|
| base, untrained, 40 layers | 276.8 ms | 408.7 ms | 733.4 ms |
| this checkpoint (40-layer) | 269.9 ms | 399.4 ms | 716.8 ms |
| 36-layer | 244.3 ms | 362.1 ms | 650.3 ms |
| 32-layer | 218.0 ms | 324.5 ms | 584.7 ms |
Folding the adapter in costs nothing measurable against the untrained base at the same depth. The 36-layer checkpoint is also measured, same day, against two other decision models over their own HTTP servers -- see its model card for that table.
How it relates to the other checkpoints
This project publishes four repositories:
littlemex/prismyra-decision-lora-qwen3.6-35b-a3b-- the LoRA adapter alone, trained against this checkpoint's base weights.littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l-- this repository. All 40 layers, that adapter folded into the FP8 weights.littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l-- 36 layers, trained with a fresh LoRA distilled from this checkpoint's own output distribution. Faster (above) and scores 2 questions above this checkpoint on Kev (655 against 653 of 764) -- not more accurate, evidence that the four fewer layers did not cost accuracy on that set.littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-32l-- 32 layers, trained the same way. Faster again, and ahead of both longer checkpoints on the two long-context sets above, but the one of the three that misses the BoolQ bar (by a single question).
The 36-layer checkpoint is the one recommended by default; this one is the reference it was distilled from.
Licence and data terms
Apache-2.0, the base model's licence, and this repository carries the base model's LICENSE file unchanged. Some of
the adapter's training sources carry their own terms -- RACE and SciQ are distributed for non-commercial research use
-- so check those terms before commercial use of this checkpoint.
- Downloads last month
- 64