David Ha runs K3 locally on M5 Max
AI researcher David Ha (@hardmaru) shared an experiment running the massive K3 language model locally on Apple's M5 Max chip. Operating at a throughput of approximately 0.3 tokens per second, the test demonstrates the capability of high-capacity Apple Silicon unified memory to host huge models, even if the current performance is exceptionally slow.
Local execution of massive frontier models on workstation hardware is technically impressive but currently constrained by severe compute and memory bandwidth bottlenecks.
- –Unified memory architecture allows consumer-tier workstations to fit multi-hundred-billion parameter models that previously required enterprise GPU clusters.
- –An inference speed of 0.3 tokens per second is impractical for interactive workflows, serving primarily as an early proof-of-concept.
- –Future optimizations in quantization, MoE sparse execution, and specialized matrix kernels will likely improve local token generation speeds substantially.
DISCOVERED
1h ago
2026-07-28
PUBLISHED
1h ago
2026-07-28
RELEVANCE
AUTHOR
hardmaru