PhD-Zero lifts Qwen3-1.7B to 20% AIME25

// 71d agoBENCHMARK RESULT

PhD-Zero lifts Qwen3-1.7B to 20% AIME25

A LocalLLaMA post reports that a PhD student used PhD-Zero, an autonomous R&D agent workflow, to tune Qwen3-1.7B Base from 0.0% to 20.0% on AIME25 in 48 hours across 11 mostly hands-off iterations. The author attributes the jump to thinking compression and an agent-detected training bug fix (loss_mask mismatch) that unlocked learning.

// ANALYSIS

Interesting signal for agentic model optimization, but it reads as an early proof-of-concept rather than a settled breakthrough.

–The result suggests small models may benefit from shorter, cleaner reasoning traces instead of longer CoT.
–The autonomous debugging claim is notable: finding and fixing a `qwen` vs `qwen3` masking issue is exactly the kind of tedious bottleneck agents can remove.
–Even with the gain, 20.0% is still well below the cited reproduced Qwen-Thinking baseline (33.3%), so there is headroom before parity.
–Evidence is currently a Reddit discussion plus project repo, so independent reproduction will matter before treating this as a robust benchmark shift.

// TAGS

phd-zeroqwen3-1.7bbenchmarkfine-tuningagentreasoningopen-source

DISCOVERED

71d ago

2026-03-17

PUBLISHED

71d ago

2026-03-17

RELEVANCE

8/ 10

AUTHOR

Rare-Salt2588

// KEEP READING

More AI developer news from the feed

EXPLORE FULL FEED

UPDATE1h ago

Cursor adds dedicated subagents for skills

Cursor now allows developers to execute tool-heavy or research-intensive agent skills within dedicated subagents. This architectural shift isolates noisy background tasks, keeping the main chat context clean and focused.

UPDATE2h ago

YouTube moves AI labels to video player

YouTube is moving its AI content disclosures from video descriptions to more prominent placements beneath the player and on Shorts overlays. Starting in May, the platform will use internal signals to automatically label photorealistic AI content that creators fail to disclose.

OPEN SOURCE5h ago

Taste Skill kills AI "frontend slop"

Taste-Skill is an open-source framework that provides portable "agent skills" to enforce high-end design principles in AI-generated code. By injecting specific design directives and "anti-slop" rules, it enables LLMs to produce editorial-grade UIs that bypass generic, boilerplate-heavy AI templates.