llama.cpp Adds RPC Tensor Parallelism
llama.cpp merged experimental `-sm tensor` support, letting RPC-connected machines split model weights and KV work across devices instead of assigning each host separate layers. Tests on two DGX Sparks over RDMA reached 19.75 tok/s generation and 619.36 prompt tok/s on DeepSeek4 MXFP4. [PR #26610](https://github.com/ggml-org/llama.cpp/pull/26610)
This is a major step toward practical multi-box local inference, bringing tensor-parallel execution to llama.cpp without a fork. The catch is that it remains highly dependent on fast interconnects and model support.
- –Adds asynchronous graph execution, cross-node all-reduce, graph caching, and 2D tensor transfers for RPC workers.
- –Pools memory across machines, making models possible that cannot fit on either box individually.
- –RDMA is the sweet spot; tests over gigabit networking showed tensor mode far behind layer splitting.
- –The feature is experimental, and users must keep RPC clients and servers on compatible builds.
- –llama.cpp documents tensor mode as parallelized but experimental, with flash attention typically required. [Multi-GPU documentation](https://github.com/ggml-org/llama.cpp/blob/master/docs/multi-gpu.md)
DISCOVERED
1h ago
2026-10-06
PUBLISHED
1h ago
2026-10-06
RELEVANCE
AUTHOR
BTC_Cracker