ClarifyRL — Run 2 — Qwen3-1.7B GRPO (β=0, no KL anchor) — the regression checkpoint

This is the negative-result checkpoint. Run 2 is the same recipe as Run 4 with the only difference being beta=0 (no KL anchor against the frozen base). It demonstrates capability collapse: the model's mean score dropped on 3 of 4 families while it found one peak in meeting_scheduling. We publish this checkpoint deliberately as the "before" half of the hackathon ablation.

If you want the working model, use anurag203/clarify-rl-run4-qwen3-1.7b-beta0.2 instead.

Why publish a regression?

Hackathon evidence works both ways. Run 2 is half of a counter-factual: same base, same data, same step count, only β changes. By publishing both the regression (β=0, this card) and the recovery (β=0.2, Run 4), we let judges and other researchers verify the central thesis end-to-end on the actual weights, not just on plots.

Eval metric (n=50, held-out) 1.7B base Run 2 (β=0) ← this checkpoint Run 4 (β=0.2)
avg_score (μ across 4 families) 0.067 0.029 ↓ 0.056 ✅
completion_rate 18% 6% ↓ 14%
event_planning μ 0.138 0.000 ❌ 0.175 ✅
meeting_scheduling μ 0.153 0.130 0.064
meeting_scheduling max 0.500 0.725 ↑ 0.350

Read the table this way: without a KL anchor, GRPO traded broad competence for one extreme peak in meeting_scheduling. The peak is real (max 0.725 is the highest across all evaluated 1.7B variants), but the mean across families collapsed.

Model summary

Field Value
Base model Qwen/Qwen3-1.7B
Algorithm TRL GRPO (Group Relative Policy Optimization)
KL anchor (β) 0.0 (intentionally — this is the "no anchor" arm of the ablation)
Learning rate 1e-6 (vs Run 4's 5e-7)
Steps 300
Wall time ~70 min on a single A100 (HF Jobs a100-large)
Generations / step 4
Reward stack OutputCorrectnessRubric (0.6) + EfficiencyRubric (0.2) + FormatCheckRubric (0.2)
Cost ~$2.21 of HF Jobs credit

Where the weights live

This anurag203/* repo hosts the rich card / metadata only. The actual 300-step Run 2 weights are checkpointed at agarwalanu3103/clarify-rl-grpo-qwen3-1-7b on the training account. A unified-namespace mirror is in flight; in the meantime download from the upstream repo directly:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

repo = "agarwalanu3103/clarify-rl-grpo-qwen3-1-7b"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
mdl = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype=torch.bfloat16, device_map="auto",
    trust_remote_code=True,
)
# … same agent loop as Run 4. See scripts/eval_agent.py for the full driver.

Intended use

  • Ablation comparison only. The model exists to be compared against Run 4 (β=0.2) on the same eval; that's the entire reason it is public.
  • Reproducing the docs/blog.md KL-anchor finding from raw weights.
  • Diagnosing what capability collapse looks like in a small RL'd reasoner.

Out-of-scope use

  • Anything resembling production. This is a deliberately broken checkpoint, kept for science.
  • Anything that depends on broad multi-family competence — see the per-family table.

Training procedure

Same env / reward / scaffolding as Run 4. The reproducible command is:

BETA=0 LEARNING_RATE=1e-6 NUM_STEPS=300 NUM_GENERATIONS=4 \
MAX_COMPLETION_LEN=768 \
python training/train_grpo.py

See training/train_grpo.py. The 300-step log_history.json is committed at outputs/run_artifacts/1.7B-noKL/.

Evaluation

Family 1.7B base μ Run 2 (β=0) μ Run 4 (β=0.2) μ
event_planning 0.138 0.000 ❌ 0.175 ✅
meeting_scheduling 0.153 0.130 0.064
medical_intake 0.000 0.000 0.000
support_triage 0.000 0.000 0.000
avg_score (μ) 0.067 0.029 0.056
completion_rate 18% 6% 14%
format_pass_rate 0% 0% 0%

Limitations

  • This model is intentionally worse than the base on average. Don't deploy it.
  • See the Run 4 model card for the full discussion; everything in the "Limitations" section there applies more strongly here.

Citation

@misc{agarwal2026clarifyrl,
  author       = {Agarwal, Anurag},
  title        = {ClarifyRL: Teaching small LLMs to ask before they act,
                  with KL-anchored GRPO},
  year         = {2026},
  howpublished = {\url{https://github.com/anurag203/clarify-rl}},
  note         = {Hackathon submission, Apr 26 2026.}
}

License

Apache-2.0 — same as the upstream Qwen3-1.7B base.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for anurag203/clarify-rl-run2-qwen3-1.7b-no-kl

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1264)
this model

Evaluation results

  • avg_score (μ over 4 families × 50 scenarios) on ClarifyRL eval (50 held-out scenarios, n=50 v4)
    self-reported
    0.029
  • completion_rate on ClarifyRL eval (50 held-out scenarios, n=50 v4)
    self-reported
    0.060
  • event_planning μ (regression — was 0.138 in base) on ClarifyRL eval (50 held-out scenarios, n=50 v4)
    self-reported
    0.000
  • meeting_scheduling max (the lone peak) on ClarifyRL eval (50 held-out scenarios, n=50 v4)
    self-reported
    0.725