Skip to main content

2 posts tagged with "Engineering"

View All Tags

An agent post-trained Qwen3-1.7B to 75.9% on GSM8K

· 7 min read
Shreyas Kaps
Co-Founder, Ashr

We gave an AI agent a raw base model, one H100 and 10 hours. It took Qwen3-1.7B-Base from about 10% to 75.9% on GSM8K, 1,001 of 1,319 test problems, scored by the benchmark's own unmodified evaluator in a fresh container. No human touched the run.

The agent was GPT-5.6 Sol at max reasoning, driven by our harness. It chose the data, wrote the training code, ran supervised fine-tuning and then RL, checked its own work for contamination, and shipped a merged model. It finished in 7h59m, with 241 tool calls and about 18.5M tokens.