llama.cpp Vulkan Posts 90.5% on Ornith 9B
A DIY Smart Code benchmark compares llama.cpp’s Vulkan backend with Ollama on an AMD Radeon RX 7900 XTX running Ornith 9B. The direct llama.cpp setup reaches 90.5% accuracy and 114.6 generation tokens per second.
This is a strong reminder that local inference performance depends heavily on runtime, backend, drivers, and configuration—not just the model or GPU.
- –llama.cpp’s [Vulkan support](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md) gives AMD users a flexible alternative to ROCm-based deployments.
- –The result is compelling for interactive local use, but it reflects one model, GPU, quantization, and benchmark setup.
- –Community testing has also found performance gaps between standalone llama-server and Ollama’s Vulkan path on RX 7000 cards, suggesting integration and version differences matter.
- –Developers should benchmark llama.cpp, Ollama, and ROCm/Vulkan on their own workloads before choosing a runtime.
DISCOVERED
46d ago
2026-08-24
PUBLISHED
46d ago
2026-08-24
RELEVANCE
AUTHOR
DIY Smart Code