USTC debuts PUMA to curb LLM overthinking
PUMA is a training-free diagnostic framework developed by USTC researchers to detect and mitigate cognitive stagnation in Large Reasoning Models. By analyzing geometric momentum and entropic uncertainty in latent space, PUMA enables inference engines to adaptively truncate redundant reasoning trajectories and lower token consumption.
As extended Chain-of-Thought reasoning becomes the default pattern for complex LLM workloads, dynamic diagnostic tools like PUMA will be essential for reducing compute costs and preventing inference deadlocks.
- –Differentiates productive deep reasoning from redundant, repetitive overthinking loops in real time.
- –Uses latent-space geometric momentum and entropic uncertainty alignment without requiring extra model training.
- –Enables automatic adaptive truncation of wasteful reasoning paths to dramatically lower token bills for enterprise LLM deployments.
DISCOVERED
2h ago
2026-07-22
PUBLISHED
2h ago
2026-07-22
RELEVANCE
AUTHOR
Discover AI