Alibaba tops AIME25 with agentic data synthesis

// 87d agoRESEARCH PAPER

Alibaba tops AIME25 with agentic data synthesis

Alibaba and SJTU's Agentic Proposing framework trains a 4B proposer model to synthesize high-difficulty reasoning problems by composing modular skills as a sequential agentic decision process. A 30B downstream solver trained on just 11,000 agent-synthesized trajectories hits 91.6% on AIME 2025 — rivaling frontier proprietary models.

// ANALYSIS

The real bottleneck for reasoning model training has always been data quality, not model scale — and this paper makes the case more forcefully than most with a surprisingly small 11K trajectory count.

–The core insight: treat problem synthesis as a multi-step agentic task (Draft → Check → Refine → Finalize) with a library of composable atomic skills, so the proposer can build genuinely hard, verifiable problems rather than trivially easy or unsolvable ones
–Multi-Granularity Policy Optimization (MGPO) is the RL secret sauce — it combines trajectory-level and stage-level advantage estimates to handle sparse rewards during proposer training, outperforming standard GRPO by 6.5 points
–Cross-domain generalization is notable: gains aren't just on math (AIME, HMMT) but extend to coding (LiveCodeBench +5pts), science (OlympicArena +4.4%), and general reasoning (GPQA +6.3%)
–The 91.6% AIME25 figure needs scrutiny — no independent replication yet, and the GitHub repo is an empty placeholder; the results hinge on their proprietary verifier ensemble (gpt-oss-120b, DeepSeek-V3.2-Special, Qwen3-235B-Thinking)
–If the code/weights release materializes, this could meaningfully shift how labs approach cold-start synthetic data pipelines for reasoning models

// TAGS

agentic-proposingllmreasoningfine-tuningbenchmarkagentresearch

DISCOVERED

87d ago

2026-03-15

PUBLISHED

87d ago

2026-03-15

RELEVANCE

8/ 10

AUTHOR

Discover AI

// KEEP READING

More AI developer news from the feed

EXPLORE FULL FEED

UPDATE13m ago

OpenCode 1.17.3 references external Git repositories

OpenCode, the open-source and model-agnostic AI coding agent, has released version 1.17.3. This update introduces a references block to grant the agent direct context and read/write access to external Git repositories or local folders.

TUTORIAL27m ago

Nadia Zueva automates TikTok marketing for Aesty

Deep Learning engineer Nadia Zueva has automated her startup's organic marketing on TikTok by building an autonomous content generation and scheduling pipeline. By using a custom Claude Code skill powered by Anthropic's Claude Fable 5 to analyze and reverse-engineer successful competitor video formats, generating AI-based face avatars, and scheduling the posts using the open-source social media management tool Postiz, her self-running machine generates over 100,000 weekly views at zero ad spend to drive traffic to her fashion app Aesty.

MODEL46m ago

Claude Fable 5 enables one-shot game generation

Anthropic has launched Claude Fable 5, a Mythos-class AI model featuring a 1-million-token context window and safety guardrails that route sensitive queries to Claude Opus 4.8. The model is generating buzz as developers successfully build playable games and 3D worlds in a single prompt.