NeoHorse-1-9B

Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

GitHub ModelScope Hugging Face Company Twitter / X License: Apache-2.0

Technical Report

NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-9B for text-based agent harnesses, tool use, coding, and instruction following.

Derived from Qwen/Qwen3.5-9B and fine-tuned by TokenRhythm. This release contains language-model weights only and is repackaged for text-only inference. Vision weights are not included. Repackaging changes configuration and tensor key names, without changing the fine-tuned tensor values.

NeoHorse-1-9B evaluation results

Highlights

  • Path toward RSI: the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation–selection–update loop; extending this loop across successive iterations is the next step toward RSI.
  • Agentic post-training framework: the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
  • Data quality: exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
  • Broad gains: 69.04 macro average across ten benchmarks versus 65.60 for Qwen3.5-9B (+3.44).

Model Details

Property Value
Model family NeoHorse Agent-Native Causal Language Model
Parameters Approximately 9B
Base model Qwen3.5-9B
Post-training Routing-guided agentic post-training
Interface Text input and text output
Context length 262,144 natively and extensible up to 1,010,000 tokens.
Weight format / precision Safetensors / BF16

Evaluation

The 9B track compares NeoHorse-1-9B with five representative open-weight baselines: Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and Muse-Glimmer-30B. Results cover ten benchmarks and are grouped by capability. Higher is better; Ξ” is NeoHorse-1-9B minus Qwen3.5-9B. Bold and underline mark the best and second-best results in each benchmark row, respectively; ties share the same formatting.

Benchmark Granite-4.2-8B Qwen3.5-9B Ornith-1.5-9B Gemma-4-12B-it Muse-Glimmer-30B NeoHorse-1-9B Ξ” vs Qwen3.5-9B
πŸ€– Agentic
QwenClawBench
37.01
44.04
47.27
43.53
46.11
48.73
+4.69
WorkBuddy Bench
35.07
39.60
29.29
29.65
45.85
40.15
+0.55
PinchBench
56.93
74.55
68.22
58.89
71.35
82.25
+7.70
VitaBench
23.00
31.25
26.75
36.50
48.50
42.25
+11.00
BFCL v4
52.06
64.88
65.03
62.06
53.74
67.43
+2.55
tau2-Bench
62.28
88.04
83.68
59.37
76.64
90.82
+2.78
πŸ’» Coding
HumanEval
96.34
92.68
93.90
100.00
98.17
98.17
+5.49
LiveCodeBench v6
72.00
65.14
47.43
73.14
65.71
65.14
+0.00
πŸ“š Instruction Following
IFBench
78.00
66.33
40.00
77.67
78.67
66.33
+0.00
IFEval
92.98
89.46
71.35
94.27
93.90
89.09
-0.37
πŸ“Š Overall
Ten-benchmark average
60.57
65.60
57.29
63.51
67.86
69.04
+3.44

Reported protocol: SGLang v0.5.17 Β· temperature=1.0 Β· top_p=0.95 Β· top_k=20 Β· min_p=0.0 Β· presence_penalty=1.5 Β· repetition_penalty=1.0 Β· thinking mode enabled with enable_thinking=true and force_nonempty_content=true. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.

Deployment

The examples below are for self-hosted deployment from a downloaded local checkpoint.

Local checkpoint path

The examples below assume the checkpoint has already been downloaded to local disk. Set MODEL_PATH to the directory containing config.json, tokenizer files, and model weights.

MODEL_PATH="/path/to/NeoHorse-1-9B"

The OpenAI-compatible requests below use the server's --served-model-name (for example, neohorse-1-9b), not the filesystem path.

SGLang

The technical report uses SGLang v0.5.17.

pip install "sglang==0.5.17"
MODEL_PATH="/path/to/NeoHorse-1-9B"
python3 -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --served-model-name neohorse-1-9b \
  --host 0.0.0.0 \
  --port 30000 \
  --context-length 262144 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder

Send an OpenAI-compatible request after the server starts:

curl http://localhost:30000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'

vLLM

pip install -U vllm
MODEL_PATH="/path/to/NeoHorse-1-9B"
vllm serve "$MODEL_PATH" \
  --served-model-name neohorse-1-9b \
  --host 0.0.0.0 \
  --port 8000 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

The server exposes an OpenAI-compatible /v1/chat/completions endpoint. Send a request after the server starts:

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'

The example uses the configured 262,144-token context limit. Actual capacity depends on GPU memory and serving settings; reduce the context limit if needed.

License

NeoHorse-1-9B is released under the Apache License 2.0.

The upstream model is Qwen/Qwen3.5-9B. Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm has modified the model through fine-tuning and repackaging for text-only inference. Modification notices are included in this model card and the released configuration, weight index, and Safetensors metadata.

Citation

@misc{neohorse2026,
  title        = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
  author       = {NeoHorse Team},
  year         = {2026},
  howpublished = {arXiv preprint},
  eprint       = {2609.08183},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url          = {https://arxiv.org/abs/2609.08183}
}

For questions or issue reports, use the NeoHorse project repository.

Downloads last month
14,901
Safetensors
Model size
9B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ 2 Ask for provider support

Model tree for TokenRhythm/NeoHorse-1-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(976)
this model
Finetunes
5 models
Merges
5 models
Quantizations
13 models

Spaces using TokenRhythm/NeoHorse-1-9B 4

Collection including TokenRhythm/NeoHorse-1-9B

Paper for TokenRhythm/NeoHorse-1-9B