Ray 2.58.0 Posts 60% Throughput Gain
Anyscale stress-tested Ray 2.58.0 across 1,600 NVIDIA Blackwell GPUs to caption 600 TB of video with CoreWeave. The reported workload delivered 60% higher streaming throughput, 23% faster batch inference, and 24% faster Ray Data shuffle.
This is a meaningful infrastructure result, but it is a workload benchmark—not a universal speedup guarantee.
- –The test produced 70 million captions in 95 minutes, highlighting Ray Data’s fit for large-scale video and physical-AI data curation.
- –Gains across streaming, inference, and shuffle suggest improvements throughout the data plane, not just in one optimized stage.
- –Storage and caching were critical: CoreWeave’s LOTA cache loaded a 17 GB Qwen3-VL model in roughly two seconds.
- –The results may not generalize beyond Blackwell GPUs, CoreWeave storage, and the tested pipeline configuration.
- –Teams should benchmark their own GPU, storage, and shuffle topology before planning around these percentages.
DISCOVERED
2h ago
2026-09-02
PUBLISHED
1d ago
2026-08-31
RELEVANCE
AUTHOR
xinyzng