NVIDIA LoGRA Cuts LLM RL Memory 45.7%
NVIDIA researchers introduce LoGRA, a reinforcement-learning post-training method that compresses gradients into low-rank sketches and reduces average training memory by up to 45.7% without sacrificing tested reasoning performance. It also trains a 27B model for over 1,100 steps on one eight-GPU node. [Paper](https://arxiv.org/abs/2610.06647)
LoGRA attacks one of RL post-training’s nastiest constraints: optimizer and gradient memory, not just model weights. The results are promising, but broader validation beyond reasoning benchmarks will determine whether this becomes a practical training default.
- –Low-rank gradient sketches reduce storage while preserving useful update signals
- –Predicted-KL step control limits destabilizing policy updates caused by approximation
- –The 27B single-node result could make larger RL experiments accessible to smaller teams
- –The reported gains are benchmark-specific, and dense Adam could not run at 27B for a direct quality comparison
- –Code is available through NVIDIA’s Molt library
DISCOVERED
1h ago
2026-10-07
PUBLISHED
2h ago
2026-10-07
RELEVANCE
AUTHOR
mark_k
