Qwen 3.8 prefill reveals GPT-5.5 Pro distillation
In a follow-up experiment to the "Stolen Thoughts" research, independent researcher wsxiaoys evaluated four open-weight models—DeepSeek V4 Flash, Inkling, Kimi K3, and Qwen 3.8 A95B—by injecting the first 1% of GPT-5.5 Pro's reasoning trace into each target model's reasoning channel across 45 STEM, non-STEM, and synthetic puzzle problems. Measuring the n-gram recall of the teacher's visible response within the target model's freely generated answer showed that while DeepSeek V4 Flash and Inkling were unaffected (-1.17 pp and +0.46 pp), Qwen 3.8 surged by +18.18 percentage points (rising from 16.79% to 34.97%), spearheaded by a +26.99 pp jump on STEM problems and a +14.75 pp gain on synthetic puzzles. Given that Qwen showed virtually no movement toward Claude Opus 4.8 in prior testing, the findings strongly suggest that Qwen was trained or distilled on reasoning traces derived from GPT-5.5 Pro or related OpenAI models.
Reasoning trace prefilling serves as an empirical lie detector for model provenance, demonstrating that distillation leaves indelible activation signatures that surface-level fine-tuning cannot conceal.
- –Prefill as a mechanistic fingerprint: Injecting merely 1% of a teacher's reasoning trace triggered an +18.18 pp alignment surge in Qwen 3.8, proving that target models can be steered along the exact problem-solving trajectories of their distillation sources.
- –Stark architectural divergence: While Qwen demonstrated heavy sensitivity, DeepSeek V4 Flash (-1.17 pp) and Inkling (+0.46 pp) remained agnostic, underscoring fundamentally different training corpora and independence from GPT-5.5 reasoning traces.
- –Synthetic puzzle gains rule out data contamination: The +14.75 pp increase on private, synthetic puzzle tasks demonstrates true heuristic mimicry rather than rote pretraining memorization of public benchmark problems.
- –Heightened pressure on open-weights provenance: As frontier labs increasingly prohibit using outputs to train competing models, reproducible prefill benchmarks provide clear evidentiary trails of cross-lab synthetic distillation.
DISCOVERED
1h ago
2026-09-09
PUBLISHED
4h ago
2026-09-09
RELEVANCE
AUTHOR
wsxiaoys