Applied Results

SFT & RLFT

The strongest proof page in the repo: Qwen2.5-Math-1.5B on gsm8k with measured gains.

1.56% Zero-shot baseline
62.9% After SFT
100% Format compliance
  1. Load gsm8k train and test sets.
  2. Evaluate Qwen2.5-Math-1.5B zero-shot behavior.
  3. Run supervised fine-tuning for answer-format alignment.
  4. Run reinforcement fine-tuning to improve reward-shaped reasoning behavior.