Reinforcement Learning
Safetensors
English
qwen3_5_text
tmax
terminal-agent

TMax

💻 Code · 🤗 Models & Data · 📜 Paper · 📓 Blog

Qwen 3.5 9B — OpenThoughts-Agent (TMax v1.1)

This terminal-agent model was trained using DPPO on Qwen/Qwen3.5-9B with the OpenThoughts-Agent task set.

This model is part of the TMax v1.1 collection. The original TMax recipe is described in our paper.

Checkpoints

This branch contains step 500. The main branch contains step 500, the highest-scoring available checkpoint in the completed TB-Lite sweep. Checkpoint selection uses TB-Lite, not Terminal-Bench 2.1.

Checkpoints 100, 200, 300, 400, 500 are provided as branches named step-0100, step-0200, etc.

Evaluation Results

Training step TB-Lite (%) ± SE (pp) Selection
100 52.24 ± 2.12
200 55.60 ± 2.03
300 59.30 ± 2.00
400 52.60 ± 2.00
500 59.40 ± 2.12 main

The selected checkpoint scores 22.73 ± 1.56% on Terminal-Bench 2.1 (88 tasks, three trials per task; configure-git-webserver excluded).

These v1.1 scores use the fixed Vanillux0.2.3 harness with Sandfleet through Harbor, three trials per task (96 TB-Lite tasks). Scores and standard errors are taken from the completed evaluation sweep updated on October 5, 2026. The harness differs from the original release; use the original model cards for v1.0 results. The per-checkpoint results and pinned source revisions are included in release_manifest.json.

Model Details

Model Description

Use

Serve the model with vLLM and use the terminal-agent harness in our codebase:

uvx vllm==0.19.1 serve TMaxxx/qwen35-9b-dppo-openthoughts \
  --served-model-name tmax-v1.1 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --port 8008 \
  --max-model-len 65536 \
  --tensor-parallel-size 8

The Qwen3.5/3.6 checkpoints contain the text-only causal language model; use the Qwen3 XML tool parser.

Training Details

  • Released training step: 500
  • Main checkpoint: 500, selected on TB-Lite
  • Available milestone checkpoints: 100, 200, 300, 400, 500

The source checkpoint is TMaxxx/qwen35-9b-dppo-openthoughts at step 500. Original checkpoint files are preserved, including campaign_provenance.json where supplied. See release_manifest.json for the immutable source commits and evaluation results. See the training launch scripts and original provenance for run-specific settings.

License

This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.

Citation

If you use our model or data, please cite our paper:

@misc{ivison2026tmaxsimplerecipeterminal,
      title={Tmax: A simple recipe for terminal agents}, 
      author={Hamish Ivison and Junjie Oscar Yin and Rulin Shao and Teng Xiao and Nathan Lambert and Hannaneh Hajishirzi},
      year={2026},
      eprint={2606.23321},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.23321}, 
}
Downloads last month
25
Safetensors
Model size
9B params
Tensor type
BF16
·
Video Preview
loading

Model tree for TMaxxx/qwen35-9b-dppo-openthoughts

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(999)
this model

Dataset used to train TMaxxx/qwen35-9b-dppo-openthoughts

Paper for TMaxxx/qwen35-9b-dppo-openthoughts