SLAI T-Rex enables DeepSeek-V4 post-training on Ascend SuperPODs
SLAI T-Rex is an open training framework enabling full-parameter post-training of DeepSeek-V4 MoE models on Huawei Ascend SuperPOD clusters, achieving 34.22% MFU (a 2.93x speedup). Paired with a 10K solver-verified synthetic dataset, fine-tuned DeepSeek-V4-Flash achieved a 71.81% zero-shot Pass@1 score on complex Operations Research tasks, outperforming GPT-5.4-Mini.
SLAI T-Rex demonstrates that full-parameter post-training of trillion-parameter scale MoE models on alternative NPU hardware like Huawei's Ascend SuperPOD can achieve competitive throughput and stability through full-stack infrastructure co-design.
- –Infrastructure Efficiency: Achieves a 2.93x MFU improvement (reaching 34.22% MFU) through hierarchical parallelism and computation-communication orchestration on Ascend SuperPODs.
- –Domain Specialization: Combines domain literature with 10K solver-verified synthetic optimization samples to tailor reasoning capabilities specifically for mathematical modeling and Operations Research.
- –Benchmark Performance: Reaches a 71.81% zero-shot Pass@1 accuracy on complex OR problems, outperforming GPT-5.4-Mini by 3.98 percentage points.
DISCOVERED
3h ago
2026-07-23
PUBLISHED
3h ago
2026-07-23
RELEVANCE
AUTHOR
_akhaliq