Dust Challenges Backprop With Activation-Space Search
QLabs presents Dust, a zeroth-order method that pretrains transformers by perturbing activations and estimating credit from forward passes instead of backpropagation. It approaches or sometimes exceeds backpropagation in experiments, but currently requires substantially more compute.
Dust is a compelling research result, not yet a practical replacement for backpropagation—the real breakthrough is making forward-only learning competitive enough to expand the space of trainable architectures.
- –Activation perturbations create a virtual population evaluated in parallel across tokens, avoiding the cost of materializing weight-space populations
- –Dust reportedly outperforms weight-space EGGROLL by roughly 1,000–10,000× in efficiency based on extrapolations
- –Larger models become more population-efficient in the experiments, challenging the assumption that zeroth-order methods collapse in high dimensions
- –Gradient estimates increasingly align with backpropagation as population size grows, while occasionally finding better optimization trajectories
- –The method remains compute-heavy, but could matter for recurrent, externally programmable, or otherwise nondifferentiable networks
DISCOVERED
1h ago
2026-10-06
PUBLISHED
4h ago
2026-10-05
RELEVANCE
AUTHOR
E-Reverance